Hyperspectral image classification model training method, system, equipment and medium
By combining multi-scale feature extraction and attention mechanisms in hyperspectral image classification, target images are generated and binding training are carried out, the problem of insufficient accuracy and generalization ability of the existing hyperspectral image classification methods is solved, and efficient and accurate hyperspectral image classification is achieved.
Patent Information
- Application Number
- CN202510196834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing hyperspectral image classification methods have problems such as low classification accuracy and poor generalization ability when processing complex hyperspectral data. The application of deep learning models in hyperspectral image classification has problems such as complex model structure, large computing resource consumption, and insufficient feature extraction.
By acquiring the hyperspectral image dataset, a hyperspectral image classification model is constructed, and the multi-scale feature extraction, feature fusion, channel attention module and spatial attention module weighting processing is finally generated through depth separation convolution, and binding training is performed to obtain the target hyperspectral image classification model.
The high accuracy and strong generalization capability of the hyperspectral image classification model are achieved, the consumption of computing resources is reduced, feature extraction is optimized, and the operation efficiency of the model is improved.
Smart Images

Figure CN120125847A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of hyperspectral image processing, and in particular, to a method, system, device and medium for training a hyperspectral image classification model. Background Art
[0002] Hyperspectral image classification refers to classifying each pixel in a hyperspectral image into different ground object categories according to its semantics; a hyperspectral image not only contains rich spectral information, but also has spatial information corresponding to the spectral information. Each pixel has a continuous spectral curve, which can simultaneously reflect the spectral characteristics and spatial distribution of the ground object; it usually also contains dozens to hundreds of continuous spectral bands, which can provide more fine-grained ground object spectral characteristics.
[0003] Hyperspectral images contain rich spectral information and can perform fine classification and recognition of different ground objects, and have broad application prospects in the fields of agriculture, geology, environmental monitoring, etc.; however, hyperspectral images have the characteristics of large data volume, strong band correlation, and noise interference, which bring many challenges to the classification task; traditional hyperspectral image classification methods, such as support vector machines, decision trees, etc., often have problems such as low classification accuracy and poor generalization ability when dealing with complex hyperspectral data; in recent years, deep learning technology has achieved remarkable results in the field of image classification, but its application in hyperspectral image classification still faces some problems, such as complex model structure, large consumption of computing resources, insufficient feature extraction, etc. Therefore, a method, system, device and medium for training a hyperspectral image classification model are proposed to solve the above problems. Summary of the Invention
[0004] The present application provides a method, system, device and medium for training a hyperspectral image classification model to solve one or more technical problems existing in the prior art, and at least provide a beneficial choice or creation condition.
[0005] Other features and advantages of the present application will become apparent through the following detailed description, or be partially learned through the practice of the present application.
[0006] According to one aspect of the embodiments of the present application, a method for training a hyperspectral image classification model is provided, and the method includes:
[0007] Obtain a hyperspectral image data set, where the hyperspectral image data set includes multiple hyperspectral images;
[0008] Construct a hyperspectral image classification model, and generate a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model;
[0009] For each hyperspectral image, use the hyperspectral image and the target image as a bound training sample;
[0010] Train the hyperspectral image classification model according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model;
[0011] Among them, generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model includes:
[0012] Perform multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale;
[0013] Perform feature fusion on the first target feature and the second target feature to obtain a target fusion feature;
[0014] Perform weighted processing on the target fusion feature through a channel attention module and a spatial attention module to obtain a target key feature;
[0015] Perform depthwise separable convolution on the target key feature to generate the target image.
[0016] In an embodiment of the present application, based on the foregoing solution, the hyperspectral image is composed of a plurality of image channels, and the first target feature is obtained through the following steps:
[0017] Perform local feature extraction at a shallow scale on each of the image channels of the hyperspectral image according to preset first convolution parameters to obtain the first target feature corresponding to each of the image channels one by one.
[0018] In an embodiment of the present application, based on the foregoing solution, the second target feature is obtained through the following steps:
[0019] Perform global feature extraction at a deep scale on the hyperspectral image according to preset second convolution parameters to obtain a second target feature corresponding to the hyperspectral image.
[0020] In an embodiment of the present application, based on the foregoing solution, the target fusion feature includes channel features and spatial features, and performing weighted processing on the target fusion feature through a channel attention module and a spatial attention module includes:
[0021] Perform global average pooling and max pooling operations on the channel dimensions of each of the image channels through the channel attention module, and perform first weighted processing on the channel features corresponding to each of the image channels;
[0022] Performing a convolution operation on the spatial dimension of each of the image channels through the spatial attention module to generate a spatial attention weight table, and performing a second weighting process on the spatial features corresponding to each of the image channels according to the spatial attention weight table.
[0023] In one embodiment of the present application, based on the foregoing solution, the target key feature is composed of the key features of each of the image channels. For each image channel, the key feature of the image channel is obtained through the following steps:
[0024] Performing a first weighting process on the channel features corresponding to the image channel to obtain a first feature;
[0025] Performing a second weighting process on the spatial features corresponding to the image channel to obtain a second feature;
[0026] Determining the key feature of the image channel according to the first feature and the second feature.
[0027] In one embodiment of the present application, based on the foregoing solution, the performing depthwise separable convolution on the target key feature to generate the target image includes:
[0028] For the key feature of each image channel, performing depthwise convolution on the key feature of the image channel to obtain a depthwise convolution feature;
[0029] Performing pointwise convolution on each of the depthwise convolution features to obtain convolution fusion features of each of the image channels;
[0030] Generating the target image according to the convolution fusion features.
[0031] According to one aspect of the embodiments of the present application, there is provided a hyperspectral image classification model training system, the system including:
[0032] An acquisition unit, configured to acquire a hyperspectral image data set, where the hyperspectral image data set includes multiple hyperspectral images;
[0033] A model construction unit, configured to construct a hyperspectral image classification model, and generate a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model;
[0034] A binding unit, configured to use the hyperspectral image and the target image as a bound training sample for each hyperspectral image;
[0035] A model training unit, configured to train the hyperspectral image classification model according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model;
[0036] Among them, generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model includes:
[0037] Performing multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale;
[0038] Performing feature fusion on the first target feature and the second target feature to obtain a target fusion feature;
[0039] Performing weighted processing on the target fusion feature through a channel attention module and a spatial attention module to obtain a target key feature;
[0040] Performing depthwise separable convolution on the target key feature to generate the target image.
[0041] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium on which a computer program is stored. The computer program includes executable instructions, and when the executable instructions are executed by a processor, the method described in the above embodiments is implemented.
[0042] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a memory for storing executable instructions of the processor, and when the executable instructions are executed by the one or more processors, the one or more processors implement the method described in the above embodiments.
[0043] Advantages of the present application: Generally speaking, the present application generates corresponding target images through multiple hyperspectral images in the hyperspectral image dataset. By binding each target image to its corresponding hyperspectral image and using them as bound training samples to train the hyperspectral image classification model, the finally obtained target hyperspectral image classification model has the characteristics of high classification accuracy, strong generalization ability, and sufficient feature extraction.
[0044] Specifically, the target image is obtained by performing multi-scale feature extraction on the hyperspectral image, obtaining a first target feature with a shallow scale and a second target feature with a deep scale, and obtaining a target fusion feature through the first target feature and the second target feature. At this time, the target fusion feature is weighted by a channel attention module and a spatial attention module to obtain a target key feature. The obtained target key feature greatly reduces the data complexity of the generated image and can fully extract various key features of the hyperspectral image. Finally, the target key feature is subjected to depthwise separable convolution to generate the target image. Therefore, the obtained target image has a high similarity with the hyperspectral image. Therefore, in subsequent hyperspectral image classification, the hyperspectral image can be directly and quickly recognized through the target hyperspectral image classification model, which can reduce the large consumption of computing resources.
[0045] Filter the key information of the hyperspectral image from the channel dimension and the spatial dimension, enable the Zhen Gege model to focus on important features, enhance the adaptability to complex scenes, optimize feature extraction, and further improve the classification accuracy; reduce the amount of calculation and the number of parameters, build a lightweight model, reduce the requirements for hardware resources, improve the operation efficiency, enable the model to run on various devices, especially resource-constrained devices, and broaden the application scenarios.
[0046] In summary, through methods such as multi-scale feature extraction, the introduction of attention mechanisms, and feature enhancement, the features of hyperspectral images can be extracted more comprehensively and accurately, improving the model's classification ability for different ground objects. Compared with traditional methods, the classification accuracy is greatly improved; at the same time, the lightweight model design reduces the number of parameters and computational complexity of the model, improves the operation efficiency of the model, and reduces the computational resource requirements of the model on the premise of ensuring classification performance, enabling it to run on resource-constrained devices.
[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0049] Figure 1 It is a flowchart of a method for training a hyperspectral image classification model according to an embodiment of this application;
[0050] Figure 2A schematic plan view of a hyperspectral image shown according to an embodiment of the present application;
[0051] Figure 3 A block diagram of a hyperspectral image classification model training system shown according to an embodiment of the present application;
[0052] Figure 4 A structural diagram of an electronic device shown according to the present application;
[0053] Figure 5 A training logic diagram of a hyperspectral image classification model shown according to an embodiment of the present application. Detailed implementation manners
[0054] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0055] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or can be implemented using other methods, components, devices, steps, etc. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0056] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontrol node devices.
[0057] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0058] It should be noted that: "a plurality of" as mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0059] The implementation details of the technical solution of the embodiments of the present application are elaborated in detail as follows:
[0060] According to one aspect of the present application, a method for training a hyperspectral image classification model is provided. Figure 1 FIG. is a flowchart of the method for training a hyperspectral image classification model shown in the embodiments of the present application. The method for training a hyperspectral image classification model includes at least steps S1 to S4, which are introduced in detail as follows:
[0061] In step S1, a hyperspectral image dataset is obtained, and the hyperspectral image dataset includes multiple hyperspectral images.
[0062] Specifically, the hyperspectral image dataset contains multiple hyperspectral images for training a hyperspectral image classification model. Each hyperspectral image is obtained after adaptive data preprocessing, where the preprocessing includes using an adaptive denoising algorithm based on wavelet transform to automatically adjust the wavelet threshold according to the noise level of the image; and using an adaptive normalization method to automatically select appropriate normalization parameters according to the data distribution.
[0063] In step S2, a hyperspectral image classification model is constructed, and target images corresponding to each of the hyperspectral images are generated through the hyperspectral image classification model.
[0064] Specifically, the hyperspectral image classification model is constructed in the following manner:
[0065] First, a multi-scale feature extraction network based on the Mamba architecture is constructed. By setting convolutional kernels of different sizes (i.e., the first convolutional parameter and the second convolutional parameter described in the present application), multi-scale features of the hyperspectral image are extracted, and the multi-scale features are fused.
[0066] Second, an attention module is introduced into the Mamba architecture, including a channel attention module and a spatial attention module, to adaptively weight the fused features and highlight the target key features of the hyperspectral image.
[0067] Third, group convolution, depthwise separable convolution techniques, as well as pruning and quantization processing are used to construct a lightweight hyperspectral image classification model based on the Mamba architecture.
[0068] In an embodiment of the present application, generating target images corresponding to each of the hyperspectral images through the hyperspectral image classification model includes:
[0069] Performing multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale.
[0070] Perform feature fusion on the first target feature and the second target feature to obtain a target fusion feature;
[0071] Perform weighted processing on the target fusion feature through a channel attention module and a spatial attention module to obtain a target key feature;
[0072] Perform depthwise separable convolution on the target key feature to generate the target image.
[0073] Specifically, perform local feature extraction at a shallow scale on each image channel of the hyperspectral image according to a preset first convolution parameter to obtain the first target feature corresponding to each image channel one by one. The preset first convolution parameter can be set as needed. In the embodiments of the present application, the preset first convolution parameter may include a plurality of shallow small convolution kernels corresponding to each image channel one by one. The preset second convolution parameter is a deep large convolution kernel corresponding to the entire hyperspectral image. Shallow small convolution kernels are good at capturing local detail features, such as texture and edges; deep large convolution kernels are good at capturing macroscopic structural features, such as shape and contour.
[0074] For example, to distinguish grassland and sandy land. They may both be relatively flat in macroscopic structure, but their texture details are very different. The grassland has fine grass leaf textures, while the sandy land is relatively rough. Shallow small convolution kernels can effectively capture this texture difference and help distinguish the grassland and the sandy land.
[0075] For example, to distinguish buildings and forests. Their texture details may be relatively complex, but there are great differences in macroscopic structure. Buildings usually have regular rectangular or square contours, while forests present irregular crown shapes. Deep large convolution kernels can effectively capture this shape difference and help distinguish buildings and forests.
[0076] Here, it should be noted that a single image channel corresponds to the pixel points divided in a single hyperspectral image. For example, if the hyperspectral image is divided into pixel points of 200×200, then any pixel point corresponds to an image channel, and the image channel is the spectral band corresponding to this pixel point, which can be regarded as the information of the ground object (i.e., the ground object) within a specific spectral range.
[0077] In an embodiment of the present application, the target fusion feature includes a channel feature and a spatial feature. The performing weighted processing on the target fusion feature through a channel attention module and a spatial attention module includes:
[0078] Perform global average pooling and max pooling operations on the channel dimension of each image channel through the channel attention module, and perform first weighted processing on the channel features corresponding to each image channel;
[0079] Perform a convolution operation on the spatial dimension of each of the image channels through the spatial attention module to generate a spatial attention weight table, and perform a second weighting process on the spatial features corresponding to each of the image channels according to the spatial attention weight table.
[0080] The target key features are composed of the key features of each of the image channels. For each image channel, the key feature of the image channel is obtained through the following steps:
[0081] Perform a first weighting process on the channel features corresponding to the image channel to obtain a first feature;
[0082] Perform a second weighting process on the spatial features corresponding to the image channel to obtain a second feature;
[0083] Determine the key feature of the image channel according to the first feature and the second feature.
[0084] Specifically, whether it is the first weighting process or the second weighting process, the process is to multiply the features of each image channel by a weighting coefficient representing its importance. This weighting coefficient is automatically learned by the fully connected layer according to the data, and the larger the value, the more important the channel.
[0085] In an embodiment of the present application, the performing a depthwise separable convolution on the target key features to generate the target image includes:
[0086] For the key feature of each image channel, perform a depth convolution on the key feature of the image channel to obtain a depth convolution feature;
[0087] Perform a pointwise convolution on each of the depth convolution features to obtain a convolution fusion feature for each of the image channels;
[0088] Generate the target image according to the convolution fusion feature.
[0089] Specifically, the core feature of deep convolution is the first step of "depth separability". It performs convolution operations on each image channel independently. That is, each input image channel is convolved with only one convolution kernel, and this convolution kernel is only responsible for processing the information of this image channel. Unlike standard convolution, deep convolution does not perform cross-channel feature fusion. In hyperspectral images, each class channel (spectral band) can be regarded as capturing the information of the ground object within a specific spectral range. Deep convolution extracts spatial features independently for the image channel where each spectral band is located, and can learn the spatial texture, edge, shape and other features (i.e., deep convolution features) within each spectral band without mixing the information of other spectral bands. Compared with standard convolution, deep convolution greatly reduces the number of parameters and calculations because it avoids cross-channel convolution operations in the first step. This is especially important for processing hyperspectral images with a large number of channels and a large amount of data, and can build a lighter and more efficient model. Deep convolution is mainly responsible for extracting spatial local patterns and structures within a single spectral channel. It can be understood that it performs spatial feature analysis on each spectral "layer" separately.
[0090] Point-wise convolution, also known as 1x1 convolution, uses a 1x1 convolution kernel to perform sliding convolution on the input feature map (i.e., the hyperspectral image after feature fusion). The key point is that although the spatial size of the convolution kernel is 1x1, its depth is the same as the number of channels of the input feature map. Therefore, the 1x1 convolution is actually a linear combination of the eigenvalues of all input channels at each spatial position. The core function of point-wise convolution is to fuse the feature information between channels. After extracting the spatial features in each spectral channel through deep convolution, point-wise convolution can effectively combine and cross-fuse the spatial features of different spectral channels. It can learn the correlation and complementarity between the features of different spectral channels. The convolution kernel parameters of point-by-point convolution actually define a weight matrix between channels. By learning these weights, the network can automatically determine which channel features should be more effectively combined to extract higher-level and more abstract features.
[0091] Depthwise convolution and pointwise convolution are usually used together to form depthwise separable convolution. A complete depthwise separable convolution layer usually includes two steps: depthwise convolution and pointwise convolution.
[0092] Spatial feature extraction (deep convolution): First, deep convolution is responsible for efficiently extracting spatial features within each spectral channel, maintaining the independence of the spectral channels.
[0093] Spectral feature fusion (point-wise convolution): Then, point-by-point convolution is responsible for effectively fusing the spatial features of different spectral channels, mining the correlation between spectral channels, and forming the final multi-spectral fusion feature representation.
[0094] In step S3, for each hyperspectral image, the hyperspectral image and the target image are used as binding training samples.
[0095] Specifically, the processing step of binding training samples is performed on each hyperspectral image, thereby using the generative adversarial network to generate a synthetic image similar to the real hyperspectral image (i.e., the target image described in this application),
[0096] In step S4, the hyperspectral image classification model is trained according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model.
[0097] The preprocessed original image (i.e., the hyperspectral image described in this application) and the target image (synthetic image) are used as training data (i.e., bound training samples) and input into a lightweight hyperspectral image classification model based on the Mamba architecture for training to obtain the final target hyperspectral image classification model.
[0098] Furthermore, after obtaining the target hyperspectral image classification model, a generative adversarial network (i.e., the generator in the target hyperspectral image classification model) can be used to generate a synthetic image. The generator adopts a network structure based on the Mamba architecture, takes random noise as input, and generates a synthetic image (i.e., the target image) through multi-layer convolution and deconvolution operations; the discriminator in the target hyperspectral image classification model is used to judge the authenticity of the input image, and the generator parameters are optimized through adversarial training between the generator and the discriminator, and finally each pixel is labeled. After obtaining the synthetic image, it can be classified as a hyperspectral image according to the classification of the label.
[0099] The input of the generator is usually a random noise vector. This random noise vector can be regarded as a point in the "latent space". The goal of the generator is to learn a mapping from the latent space to the image space, converting the random noise vector into a meaningful image similar to the real hyperspectral image. The generator network usually uses multiple layers of convolution and deconvolution (or transposed convolution) operations to gradually convert the low-dimensional random noise vector into high-dimensional image data.
[0100] The core idea of the algorithm is adversarial training. The generator and the discriminator compete and learn with each other in a zero-sum game. The goal of the generator is to deceive the discriminator so that it misjudges the synthetic image as a real image. The generator constantly strives to improve its generation ability and generate more realistic synthetic images; the discriminator's goal is to distinguish between real images and synthetic images as accurately as possible to avoid being deceived by the generator. The discriminator constantly strives to improve its discrimination ability and learns more effective features to distinguish between real and fake images.
[0101] Among them, hyperspectral images can be Figure 2As shown, Figure 2 What is shown is a plane of a hyperspectral image, but in fact, a hyperspectral image is a three-dimensional image, and each pixel corresponds to an image channel described in this application. Figure 5 As shown, Figure 5 This is the logical process of the entire model training. The target hyperspectral image classification model finally generated can effectively generate synthetic images similar to hyperspectral images and perform effective pixel classification on the synthetic images.
[0102] According to one aspect of the embodiments of the present application, a hyperspectral image classification model training system 300 is proposed. Figure 3 This is a schematic diagram of a hyperspectral image classification model training system 300 proposed in an embodiment of the present application. The system 300 includes: an acquisition unit 301, a model construction unit 302, a binding unit 303, and a model training unit 304.
[0103] An acquisition unit 301 is used to acquire a hyperspectral image dataset, where the hyperspectral image dataset includes a plurality of hyperspectral images;
[0104] A model building unit 302 is used to build a hyperspectral image classification model, and generate a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model;
[0105] A binding unit 303 is used for taking, for each hyperspectral image, the hyperspectral image and the target image as binding training samples;
[0106] A model training unit 304 is used to train the hyperspectral image classification model according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model;
[0107] Wherein, generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model includes:
[0108] Performing multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale;
[0109] Performing feature fusion on the first target feature and the second target feature to obtain a target fusion feature;
[0110] The target fusion features are weighted by a channel attention module and a spatial attention module to obtain target key features;
[0111] Performing depth-wise separable convolution on the target key features to generate the target image.
[0112] As another aspect, the present application further provides a computer-readable storage medium on which a program product capable of implementing the method provided above in this specification is stored. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary implementations of the present application described in the above "Embodiment Method" section of this specification.
[0113] According to the embodiment of the present application, the program product for implementing the above method can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0114] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0115] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0116] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0117] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0118] Refer to the following Figure 4 The electronic device 400 according to this embodiment of the present application is described. Figure 4 The electronic device 400 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0119] like Figure 4 As shown, the electronic device 400 is in the form of a general computing device. The components of the electronic device 400 may include but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410).
[0120] The storage unit stores program codes, which can be executed by the processing unit 410, so that the processing unit 410 executes the steps described in the above “Example Method” section of this specification according to various exemplary implementations of the present application.
[0121] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 421 and / or a cache memory unit 422 , and may further include a read-only memory unit (ROM) 423 .
[0122] The storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, such program modules 425 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0123] The bus 430 may represent one or more of several types of bus architectures, including a memory bus, or a memory controller node, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures.
[0124] The electronic device 400 may also communicate with one or more external devices 1200 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 450. Also, the electronic device 400 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 460. As shown in the figure, the network adapter 460 communicates with other modules of the electronic device 400 through the bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0125] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0126] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0127] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A hyperspectral image classification model training method, characterized in that: The method comprises: Acquire a hyperspectral image dataset, wherein the hyperspectral image dataset includes a plurality of hyperspectral images; Constructing a hyperspectral image classification model, and generating target images corresponding to each of the hyperspectral images through the hyperspectral image classification model; For each hyperspectral image, the hyperspectral image and the target image are used as binding training samples; The hyperspectral image classification model is trained according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model; Wherein, generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model includes: Performing multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale; Performing feature fusion on the first target feature and the second target feature to obtain a target fusion feature; The target fusion features are weighted by a channel attention module and a spatial attention module to obtain target key features; Performing depth-wise separable convolution on the target key features to generate the target image.
2. The hyperspectral image classification model training method according to claim 1, characterized in that: The hyperspectral image is composed of a plurality of image channels, and the first target feature is obtained by the following steps: According to the preset first convolution parameters, shallow-scale local feature extraction is performed on each image channel of the hyperspectral image to obtain the first target features corresponding to each image channel one by one.
3. The hyperspectral image classification model training method according to claim 2, characterized in that: The second target feature is obtained by the following steps: The deep-scale overall feature extraction is performed on the hyperspectral image according to the preset second convolution parameter to obtain a second target feature corresponding to the hyperspectral image.
4. The hyperspectral image classification model training method according to claim 3, characterized in that: The target fusion feature includes a channel feature and a spatial feature, and the weighted processing of the target fusion feature by the channel attention module and the spatial attention module includes: Performing global average pooling and maximum pooling operations on the channel dimensions of each of the image channels through the channel attention module, and performing a first weighted processing on the channel features corresponding to each of the image channels; The spatial attention module performs a convolution operation on the spatial dimension of each of the image channels to generate a spatial attention weight table, and performs a second weighted processing on the spatial features corresponding to each of the image channels according to the spatial attention weight table.
5. The hyperspectral image classification model training method according to claim 4, characterized in that: The target key features are composed of the key features of each of the image channels. For each image channel, the key features of the image channel are obtained by the following steps: Performing a first weighted processing on the channel feature corresponding to the image channel to obtain a first feature; Performing a second weighted processing on the spatial feature corresponding to the image channel to obtain a second feature; A key feature of the image channel is determined according to the first feature and the second feature.
6. The hyperspectral image classification model training method according to claim 5, characterized in that: The step of performing a depth-separable convolution on the target key features to generate the target image includes: For each key feature of the image channel, performing a deep convolution on the key feature of the image channel to obtain a deep convolution feature; Performing point-by-point convolution on each of the deep convolution features to obtain convolution fusion features of each of the image channels; The target image is generated according to the convolutional fusion features.
7. A hyperspectral image classification model training system, characterized in that: The system comprises: An acquisition unit, configured to acquire a hyperspectral image dataset, wherein the hyperspectral image dataset includes a plurality of hyperspectral images; A model building unit, used for building a hyperspectral image classification model, and generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model; A binding unit, for each hyperspectral image, using the hyperspectral image and the target image as binding training samples; A model training unit, used for training the hyperspectral image classification model according to the bound training samples corresponding to each of the hyperspectral images to obtain a target hyperspectral image classification model; Wherein, generating a target image corresponding to each of the hyperspectral images through the hyperspectral image classification model includes: Performing multi-scale feature extraction on the hyperspectral image to obtain a first target feature with a shallow scale and a second target feature with a deep scale; Performing feature fusion on the first target feature and the second target feature to obtain a target fusion feature; The target fusion features are weighted by a channel attention module and a spatial attention module to obtain target key features; Performing depth-wise separable convolution on the target key features to generate the target image.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the hyperspectral image classification model training method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the hyperspectral image classification model training method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Hyperspectral multispectral image fusion method and system based on spectral state fusion tree Mamba and storage medium
CN120387929A
A hyperspectral and multispectral image fusion method, system and storage medium based on spectral state fusion tree Mamba
CN120387929B