Large model guided hyperspectral image understanding method, system, medium and product
Patent Information
- Application Number
- CN202610925105.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-25
AI Technical Summary
然而,现有视觉问答方法大多针对自然图像或RGB遥感图像设计,难以有效处理高光谱图像中高维光谱信息与复杂的光谱-空间结构特征,同时缺乏对高光谱语义信息与语言语义信息的有效融合,从而导致模型在复杂遥感场景下的语义理解能力和推理能力受到限制
[0016]Compared with existing technologies, the present invention mainly achieves the following beneficial effects: The large-model-guided hyperspectral image understanding method of the present invention includes: using a three-dimensional U-shaped Mamba visual encoder-decoder network to perform hyperspectral visual encoding on the input hyperspectral image to obtain object pseudomasks and visual features; using a large model to perform language encoding on the input question text to obtain contextual semantic features; aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features; and performing fine-grained visual-linguistic fusion and answer selection on the object pseudomask, visual features, and cross-modal text features. The present invention, by utilizing a three-dimensional U-shaped Mamba visual encoder-decoder network, maintains high computational efficiency while taking into account both local spatial structure modeling and long-range dependency modeling across bands in hyperspectral images; guiding visual feature aggregation through object pseudomasks provides pixel-level semantic cues for visual question answering tasks, improving complex relationship reasoning and object analysis capabilities; and enhancing question text understanding and cross-modal semantic reasoning capabilities by introducing a large language model and an efficient LoRA parameter fine-tuning strategy. Compared with existing methods, the large-model-guided hyperspectral image understanding method of the present invention achieves better results in terms of time and reconstructed image quality.
Smart Images

Figure CN122454423B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of remote sensing image processing and artificial intelligence technology, specifically to a large model-guided hyperspectral image understanding method, system, medium, and product. Background Technology
[0002] With the rapid development of Earth observation technology, remote sensing data acquisition capabilities are moving from traditional two-dimensional spatial observation to high-dimensional information perception. Among various remote sensing data acquisition methods, hyperspectral imaging refers to the dense sampling of spectral features with many narrow bands. Unlike traditional RGB images, each pixel in a hyperspectral image contains a continuous spectral curve used to identify the material of the corresponding object. It is considered one of the most powerful data formats in the current remote sensing field in terms of information expression. Compared to traditional RGB or multispectral images, hyperspectral data can provide fine material-level "spectral fingerprints," which are of great value in distinguishing materials and have shown important application value in fields such as remote sensing, agriculture, geology, astronomy, medical imaging, food processing, and pollution monitoring. One characteristic of hyperspectral images (HIS) is their large data volume and high vector dimension. This high-dimensional data provides greater classification and discrimination capabilities and data redundancy, but it also brings high computational costs and more complex data modeling. On the one hand, the exponential growth of spectral dimensions leads to the "curse of dimensionality," significantly reducing the generalization ability of models under limited sample conditions. On the other hand, hyperspectral data exhibits strong spatial-spectral coupling, making it difficult for traditional two-dimensional vision models to simultaneously model local spatial structures and long-range dependencies across spectral bands. Therefore, early research mainly relied on shallow methods such as support vector machines or spectral angle matching, while in recent years, deep learning has gradually become the mainstream technical approach for hyperspectral understanding.
[0003] Convolutional neural networks (CNNs) have made significant progress in hyperspectral classification tasks through their local receptive field mechanism, but they are inherently biased towards local feature extraction and struggle to capture long-range spectral dependencies. Subsequently, the Transformer model, utilizing self-attention mechanism for global modeling, has been widely applied to visual tasks and achieved breakthroughs in hyperspectral classification and segmentation. However, the computational complexity of self-attention increases quadratically with sequence length, leading to significant computational and storage overhead in high-dimensional hyperspectral data. In recent years, the Structured State Space Model (SSM) and its representative structure, Mamba, have been proposed. Through a selective scanning mechanism, they achieve linear complexity sequence modeling, significantly improving computational efficiency while maintaining the global receptive field, providing a new technical path for modeling long hyperspectral sequences. Meanwhile, the research goal of remote sensing intelligent analysis is gradually shifting from "recognition" to "understanding." Traditional hyperspectral research mainly focuses on discrimination tasks such as classification, detection, or segmentation, with model outputs typically being discrete labels or probability maps, which are insufficient to meet the needs of humans for interactive queries and knowledge acquisition using natural language. Visual question answering (VQA), an important research direction at the intersection of visual understanding and natural language processing, achieves question-based semantic reasoning and knowledge representation by jointly modeling images and text, and is considered an important approach to achieving explainable artificial intelligence. Currently, VQA methods and datasets cover natural imagery and remote sensing, typically involving classification and general counting tasks. Although VQA methods have made significant progress in natural imagery, research on remote sensing scenarios is still under development. Works such as RSVQA and EarthVQA have extended visual question answering to the remote sensing domain for the first time, demonstrating that semantic information can effectively support reasoning about complex spatial relationships. However, these methods are mainly geared towards RGB or high-resolution remote sensing imagery and have not fully considered the spectral dimensional information unique to hyperspectral data. Because hyperspectral images contain continuous spectral variations and material property expressions, relying solely on two-dimensional visual features is insufficient for achieving fine semantic understanding and cross-object relationship reasoning; therefore, directly transferring existing VQA methods often yields limited results.
[0004] In recent years, the rise of Large Language Models (LLMs) has provided a new research paradigm for multimodal understanding. Based on large-scale pre-trained corpora, LLMs have demonstrated unprecedented capabilities in knowledge reasoning, semantic understanding, and text generation. Recent research shows that using visual perception results as structured prior input to language models can significantly improve the ability to understand complex scenes and generate decisions. For example, the EarthVL method proposes a progressive visual language learning framework of "segmentation-understanding-generation," enabling interactive analysis of remote sensing images in natural language form. This paradigm indicates that visual models are responsible for perceiving the world, while language models are responsible for reasoning and expression; their collaboration has become an important direction for the development of multimodal intelligence. However, most existing visual question answering methods are designed for natural images or RGB remote sensing images, making it difficult to effectively handle the high-dimensional spectral information and complex spectral-spatial structural features in hyperspectral images. Furthermore, they lack effective fusion of hyperspectral semantic information and linguistic semantic information, thus limiting the semantic understanding and reasoning capabilities of models in complex remote sensing scenarios. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a method, system, medium, and product for understanding hyperspectral images guided by a large model, in response to the above-mentioned problems in the prior art. This invention aims to make full use of the spectral and spatial information in hyperspectral images and combine the semantic reasoning ability of large models to enhance the ability of question text understanding and cross-modal semantic reasoning in visual question answering tasks of hyperspectral images.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A large-model-guided hyperspectral image understanding method includes the following steps: The input hyperspectral image is hyperspectral visually encoded using a 3D U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features; the input question text is linguistically encoded using a large model to obtain contextual semantic features, and the contextual semantic features are aligned with the dimensions of the visual features to obtain cross-modal text features; fine-grained visual-linguistic fusion and answer selection are performed on the object pseudomask, visual features, and cross-modal text features: S101, the object pseudomask is scale-matched to the same spatial resolution as the visual features, and then the scale-matched object pseudomask is coupled with the visual features to form coupled features through an object-guided attention module to highlight object regions and pixel-level semantic cues; S102, the coupled features are processed through a multi-head self-attention operation and then concatenated with the original coupled features, and the concatenated features are processed through a feedforward neural network to obtain visual features after visual object context modeling. S103, Visual features after modeling the context of a visual object. Cross-modal text features are obtained by using a bidirectional cross-attention mechanism to obtain cross-modal enhanced visual features. and text features S104, enhancing visual features across modalities. and text features After concatenation and pooling, the predicted probability distribution or predicted count result of the answer category is obtained through a classification head or a counting regression head: ; ; in, The predicted probability distribution of the answer category obtained from the classification head. For normalized activation functions, and These are the weight matrix and bias term of the classification head, respectively; The predicted count results obtained from the count regression head. and These represent the weight matrix and bias term of the counting regression head, respectively. This refers to the output features obtained from the pooling aggregation operation.
[0007] Optionally, the step of using a three-dimensional U-shaped Mamba visual encoder-decoder network to perform hyperspectral visual encoding on the input hyperspectral image to obtain object pseudomasks and visual features includes: performing block preprocessing on the hyperspectral image and performing patch embedding to obtain initial visual features. ; in, As initial visual features, For patch embedding operations, For convolutional mappings used to unify spectral dimensions, To perform block operations based on preset size pairs, Given a hyperspectral image as input, the initial visual features are encoded and decoded using a 3D U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features.
[0008] Optionally, the three-dimensional U-shaped Mamba visual encoder-decoder network is an encoder-decoder structure composed of stacked multi-level encoders and stacked multi-level decoders, and the functional expression for the output features acquired by any i-th level encoder is: ; in, For the output features of the i-th level encoder, For the patch fusion module of the i-th level encoder, This is the visual state space layer of the i-th level encoder. This is the feature extraction layer of the i-th level encoder. For the input features of the i-th level encoder, , Let be the number of encoder layers; the function expression for obtaining the output features of any j-th level decoder is: ; in, For the output features of the j-th level decoder, For the j-th level decoder, the convolutional fusion module is used to map and optimize the fusion result of the upsampled features and skip connection features obtained by the upsampled module; For upsampling module, For the input features of the j-th level decoder, Indicates splicing or residual fusion. For the skip connection features of the j-th level decoder, , The number of decoder layers; the functional expression for the three-dimensional U-shaped Mamba visual encoder-decoder network to obtain the object pseudomask and visual features is: ; ; in, For object pseudomask, For convolutional mapping operations used to generate object pseudomasks, For the first The output features obtained by passing the output features of the level decoder through a 1×1 convolution are as follows: As a visual feature, This is a linear projection operation used to project decoded features into a visual feature space. This represents the output feature of the first-level decoder.
[0009] Optionally, the processing of the input features of the i-th level encoder by the feature extraction layer for the i-th level encoder includes: processing the input features of the i-th level encoder... Features are extracted using parallel shallow and deep branches respectively: ; ; in, The shallow spectral features are extracted from the shallow branches. Spatial-spectral context features extracted for deep branches, It is the ReLU activation function. This indicates the instance normalization operation. It is a 3×3×3 convolution. The convolution is 1×1×1; the shallow spectral features extracted from the shallow branches and the spatial-spectral context features extracted from the deep branches are aggregated into aggregated features: ; in, The aggregated features are used to generate the output features of the feature extraction layer of the i-th level encoder. ; in, represents the output features of the feature extraction layer of the i-th level encoder.
[0010] Optionally, the visual state space layer of the i-th level encoder processes the input features by dividing the input features into multiple local windows: ; in, For the nth local window, For input features Window partitioning operation, dividing the window into 7×7 local windows, n=1,2,…,N w N w The number of local windows; each local window is input into multiple serial vision modules VSSBlock for processing: ; in, The local window features are obtained after processing by multiple vision modules VSSBlock in a serial manner; ~ These represent the VSSBlocks from the 1st to the Lth visual modules, where L is the total number of VSSBlocks in the sequence; the local window features of all local windows are merged to generate the merged window feature: ; in, To merge window features, This is a local window feature merging operation; among the multiple serial vision modules VSSBlock, the processing of input features by any vision module VSSBlock includes: linearly projecting the input features and then dividing them into main branch features and gated branch features: ; in, Main branch features, This is a gated branch feature. For the segmentation operation, The weight matrix is the linear projection. The main branch features are used as input features; local spatial-spectral context features are extracted from the main branch features through 3D deep convolution. ; in, For local spatial-spectral context features, This is a 3D depthwise convolution; the local spatial-spectral context features are transformed in multiple directions to obtain multiple directional sequences: ; in, For the m-th direction sequence, For feature flattening operation, For the m-th direction transformation of local spatial-spectral context features, the direction transformation includes some or all of the original, flipped, and rotated transformations; each direction sequence is fed into a Mamba module to perform a selective scan using the input-related state-space parameters to obtain the corresponding output features. ; Output features corresponding to each directional sequence After fusion, the gated branch features are modulated and projected back to the original feature dimensions: ; in, The output features of the vision module VSSBlock, To output the weight matrix, Representation layer normalization, To extract the output features corresponding to each directional sequence Summation operation after shaping This represents element-wise multiplication. The SiLU activation function is used. This is a gated branch feature.
[0011] Optionally, in any j-th level decoder, the upsampling module processes the input features by: adjusting the number of feature channels through convolution operations and performing preliminary integration of deep semantic information, and then using interpolation upsampling to restore the spatial size of the input feature map so that it is consistent with the skip connection features output by the corresponding coding layer in terms of spatial scale. ; in, This refers to the scale-aligned features obtained after interpolation and upsampling. For interpolation upsampling operation, For convolution operations, The input features of the upsampling module are used; the scale-aligned features obtained after interpolation and upsampling are fused with the skip connection features passed from the encoder: ; in, As a feature of fusion, For feature splicing operations, For the first The skip connection features from the layer encoder are processed by a residual block to obtain the features; the fused features can be processed through two residual blocks to enhance the feature representation capability and obtain the output features of the upsampling module. ; in, For the first The output features of the upsampling module in the layer decoder, and It is a residual block.
[0012] Optionally, the step of using a large model to perform language encoding on the input question text to obtain contextual semantic features includes: segmenting, truncating, and padding the question text Q to obtain a word index sequence I and an attention mask A; inputting the word index sequence I and the attention mask A into a pre-trained large model to obtain contextual semantic features; wherein the pre-trained large model is a large model fine-tuned and optimized by introducing a LoRA parameter fine-tuning strategy in the attention projection layer; the step of aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features refers to using linear projection and layer normalization on the contextual semantic features to obtain cross-modal text features aligned with the visual feature dimensions. ; in, For cross-modal text features, For layer normalization, For contextual semantic features, and These represent the learnable projection parameters.
[0013] The present invention also provides a large model-guided hyperspectral image understanding system, including an interconnected microprocessor and a memory, wherein the microprocessor is programmed or configured to execute the large model-guided hyperspectral image understanding method.
[0014] The present invention also provides a computer-readable storage medium storing a computer program or instructions that are programmed or configured to execute the large model-guided hyperspectral image understanding method via a processor.
[0015] The present invention also provides a computer program product, including a computer program or instructions that are programmed or configured to execute the large model-guided hyperspectral image understanding method via a processor.
[0016] Compared with existing technologies, the present invention mainly achieves the following beneficial effects: The large-model-guided hyperspectral image understanding method of the present invention includes: using a three-dimensional U-shaped Mamba visual encoder-decoder network to perform hyperspectral visual encoding on the input hyperspectral image to obtain object pseudomasks and visual features; using a large model to perform language encoding on the input question text to obtain contextual semantic features; aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features; and performing fine-grained visual-linguistic fusion and answer selection on the object pseudomask, visual features, and cross-modal text features. The present invention, by utilizing a three-dimensional U-shaped Mamba visual encoder-decoder network, maintains high computational efficiency while taking into account both local spatial structure modeling and long-range dependency modeling across bands in hyperspectral images; guiding visual feature aggregation through object pseudomasks provides pixel-level semantic cues for visual question answering tasks, improving complex relationship reasoning and object analysis capabilities; and enhancing question text understanding and cross-modal semantic reasoning capabilities by introducing a large language model and an efficient LoRA parameter fine-tuning strategy. Compared with existing methods, the large-model-guided hyperspectral image understanding method of the present invention achieves better results in terms of time and reconstructed image quality. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of a three-dimensional U-shaped Mamba visual encoding-decoding network used in an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of the network structure of the feature extraction layer in an embodiment of the present invention.
[0020] Figure 4 This is a schematic diagram of the network structure of the upsampling module in an embodiment of the present invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0022] like Figure 1As shown, the large-model-guided hyperspectral image understanding method in this embodiment includes the following steps: hyperspectral visual encoding of the input hyperspectral image using a three-dimensional U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features; language encoding of the input question text using a large model to obtain contextual semantic features; aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features; and performing fine-grained visual-linguistic fusion and answer selection using the object pseudomask, visual features, and cross-modal text features. The hyperspectral visual question answering method in this embodiment includes three parts: hyperspectral visual encoding (using a three-dimensional U-shaped Mamba visual encoder-decoder network), language encoding (using a large model), and fine-grained visual-linguistic fusion.
[0023] Let the input hyperspectral image be Hyperspectral images It can be represented as: ; in, and These represent the spatial height and spatial width of the hyperspectral image, respectively. This represents the number of spectral bands.
[0024] Let the input question text be The input question text It can be represented as: ; in, For the question text The length of the lexicon. ~ For the first to T words, any Let i be the i-th word element.
[0025] In this embodiment, the input hyperspectral image is hyperspectral visually encoded using a three-dimensional U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features, including: performing block preprocessing on the hyperspectral image and performing patch embedding to obtain initial visual features. ; in, As initial visual features, For patch embedding operations, For convolutional mappings used to unify spectral dimensions, To perform block operations based on preset size pairs, Given a hyperspectral image as input, the initial visual features are encoded and decoded using a 3D U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features.
[0026] like Figure 2 As shown, in this embodiment, the three-dimensional U-shaped Mamba visual encoder-decoder network is an encoder-decoder structure composed of stacked multi-level encoders and stacked multi-level decoders. The function expression for obtaining the output features of any i-th level encoder is: ; in, For the output features of the i-th level encoder, For the PatchMerge (PM) module of the i-th level encoder, This is the visual state space (VSS) layer of the i-th level encoder. This is the feature extraction layer (CatchFeat, CF) of the i-th level encoder. For the input features of the i-th level encoder, , The number of encoder layers is specified. The Visual State Space (VSS) layer is based on existing state space visual modeling methods, but in this embodiment, it is used for hyperspectral 3D spectral-spatial feature modeling and improved by combining 3D selective scanning. The feature extraction layer (CatchFeat (CF)) is a feature extraction structure designed to meet the requirements of joint extraction of shallow and deep features from hyperspectral data.
[0027] In this embodiment, the function expression for obtaining the output features of any j-th level decoder is: ; in, For the output features of the j-th level decoder, For the j-th level decoder, the convolutional fusion module is used to map and optimize the fusion result of the upsampled features and skip connection features obtained by the upsampled module; For upsampling module, For the input features of the j-th level decoder, Indicates splicing or residual fusion. For the skip connection features of the j-th level decoder, , The decoder layer number represents the number of layers. Finally, the functional expression for the 3D U-shaped Mamba visual encoder-decoder network to obtain the object pseudomask and visual features is: ; ; in, For object pseudomask, For convolutional mapping operations used to generate object pseudomasks, For the first The output features obtained by passing the output features of the level decoder through a 1×1 convolution are as follows: As a visual feature, This is a linear projection operation used to project decoded features into a visual feature space. This represents the output feature of the first-level decoder.
[0028] like Figure 3 As shown, the feature extraction layer (CatchFeat, CF) includes shallow and deep branches. The shallow branch uses 1×1×1 convolutions to quickly extract shallow spectral features, while the deep branch uses 1×1×1 and 3×3×3 convolutions to extract richer spatial-spectral context features. The feature extraction layer of the i-th level encoder processes the input features of the i-th level encoder as follows: [The text abruptly ends here, likely due to an incomplete translation or source material.] Features are extracted using parallel shallow and deep branches respectively: ; ; in, The shallow spectral features are extracted from the shallow branches. Spatial-spectral context features extracted for deep branches, It is the ReLU activation function. This indicates the instance normalization operation. It is a 3×3×3 convolution. The convolution is 1×1×1; the shallow spectral features extracted from the shallow branches and the spatial-spectral context features extracted from the deep branches are aggregated into aggregated features: ; in, This is an aggregated feature; through the above structure, both shallow detail information and deep context information can be preserved with low computational overhead. Instance Normalization (IN) is a well-known normalization method. Unlike batch normalization (BN), which typically calculates the mean and variance of a batch, instance normalization normalizes each individual sample and channel separately, without relying on batch size statistics. Its advantage lies in reducing instability caused by inter-batch statistical fluctuations when the training batch is small and the distribution differences between samples are large. Considering the common problems of small batches and significant sample differences in hyperspectral data training, instance normalization is used in the CatchFeat module in this embodiment to improve the stability of feature extraction. Finally, the aggregated features are passed through a 3×3×3 convolution to generate the output features of the feature extraction layer of the i-th level encoder: ; in, represents the output features of the feature extraction layer of the i-th level encoder.
[0029] The Visual State Space (VSS) layer and 3D selective state space scanning are improvements on the existing SS2D module of the SwinUMamba framework. The processing of input features by the Visual State Space (VSS) layer of the i-th level encoder includes: Divide the input features into multiple local windows: ; in, For the nth local window, For input features Window partitioning operation, dividing the window into 7×7 local windows, n=1,2,…,N w N w This represents the number of local windows; Each local window is processed by multiple sequential vision modules VSSBlock: ; in, The local window features are obtained after processing by multiple vision modules VSSBlock in a serial manner; ~ These represent the VSSBlocks from the 1st to the Lth visual modules, where L is the total number of VSSBlocks in the sequence; the local window features of all local windows are merged to generate the merged window feature: ; in, To merge window features, This is a local window feature merging operation.
[0030] In the serially generated multiple vision modules VSSBlock, each VSSBlock employs a 3D selective state-space scanning mechanism (SS3D) to perform orientation-sensitive long-range dependency modeling of the input features. The processing of input features by any VSSBlock vision module includes: linearly projecting the input features and then segmenting them into main branch features and gated branch features. ; in, Main branch features, This is a gated branch feature. For the segmentation operation, The weight matrix is the linear projection. The main branch features are used as input features; local spatial-spectral context features are extracted from the main branch features through 3D deep convolution. ; in, For local spatial-spectral context features, This is a 3D depthwise convolution; the local spatial-spectral context features are transformed in multiple directions to obtain multiple directional sequences: ; in, For the m-th direction sequence, For feature flattening operation, For the m-th direction transformation of local spatial-spectral context features, the direction transformation includes some or all of the original, flipped, and rotated transformations; each direction sequence is fed into a Mamba module to perform a selective scan using the input-related state-space parameters to obtain the corresponding output features. ; Output features corresponding to each directional sequence After fusion, the gated branch features are modulated and projected back to the original feature dimensions: ; in, The output features of the vision module VSSBlock, To output the weight matrix, Representation layer normalization, To extract the output features corresponding to each directional sequence Summation operation after shaping This represents element-wise multiplication. The SiLU activation function is used. This is a gated branch feature. Through the above structure, it is possible to simultaneously model the spatial locality and spectral continuity of hyperspectral images.
[0031] In this embodiment, the Mamba module performs a selection scan by inputting relevant state-space parameters to obtain the corresponding output features. At that time, for each direction sequence The selection scan is performed using input-related state-space parameters, which can be written as: ; The corresponding discrete state update and output are: ; in, For the input matrix, For the output matrix, To discretize the step size, , and The weight matrix is a learnable matrix. It is a soft positive function used to ensure that the discretization step size is greater than 0 in order to guarantee numerical stability. The direction sequence at time t; Let be the hidden state at time t. Let t be the discretized state transition matrix (diagonal matrix) at time t. Let be the hidden state at time t-1. The output feature at time t.
[0032] To reduce the computational complexity of decoding large-scale hyperspectral data, this embodiment employs a combination of bilinear interpolation and convolution for upsampling. For example... Figure 4 As shown, in any j-th level decoder, the upsampling module processes the input features by: adjusting the number of feature channels through convolution operations and performing preliminary integration of deep semantic information; and then using interpolation upsampling to restore the spatial size of the input feature map so that it is consistent with the skip connection features output by the corresponding coding layer in terms of spatial scale. ; in, This refers to the scale-aligned features obtained after interpolation and upsampling. For interpolation upsampling operation, For convolution operations, The input features of the upsampling module are used; the scale-aligned features obtained after interpolation and upsampling are fused with the skip connection features passed from the encoder: ; in, As a feature of fusion, For feature splicing operations, For the first The skip connection features from the layer encoder are processed by a residual block to obtain the features; the fused features can be processed through two residual blocks to enhance the feature representation capability and obtain the output features of the upsampling module. ; in, For the first The output features of the upsampling module in the layer decoder, and It is a residual block.
[0033] In this embodiment, when processing the input features, the upsampling module uses bilinear interpolation to restore the spatial scale and convolution to further aggregate spectral and spatial features, thereby maintaining high reconstruction efficiency while avoiding the additional computational burden of pure transposed convolution. First, deep features from the previous decoding layer are input into the convolutional layer for feature mapping. This convolution operation is mainly used to adjust the number of feature channels and to initially integrate deep semantic information. Subsequently, the module uses interpolation upsampling to restore the spatial size of the input feature map, ensuring it is consistent with the skip connection features output by the corresponding encoding layer in terms of spatial scale. After scale alignment, the upsampled deep features are fused with the skip connection features from the encoder. Deep features contain strong global semantic information, while skip connection features retain more edge, texture, and local spatial detail information. Through their fusion, the decoder can restore spatial resolution while reducing information loss during downsampling. The fused features are then further input into the residual block for feature enhancement. The residual block re-integrates the fused features through local convolutional transformations and residual connections, enabling the network to extract more discriminative spatial-semantic features while retaining the original effective information. The two residual blocks shown in the figure can be used to enhance feature representation capabilities and improve the feature recovery quality during the decoding stage. The processing of input features by the upsampling module in this embodiment can be summarized as follows: ; Functionally, this structure utilizes interpolation to achieve lightweight upsampling, avoiding the additional computational burden of complex transposed convolutions. Simultaneously, it supplements shallow detail information through skip connections and enhances the expressive power of fused features using residual blocks. Therefore, this module can effectively recover hyperspectral image features with lower computational complexity, providing a more complete multi-scale feature representation for subsequent pixel-level semantic prediction.
[0034] In this embodiment, the input question text is used to perform language encoding to obtain contextual semantic features, including: segmenting, truncating, and padding the question text Q to obtain the word index sequence I and the attention mask A, which can be represented as: ; in, This is for language encoding operations using large models.
[0035] Inputting the lexical index sequence I and the attention mask A into a pre-trained large model yields contextual semantic features: ; in, For contextual semantic features, For pre-trained large models; The pre-trained large model is a large model fine-tuned and optimized by introducing a LoRA parameter fine-tuning strategy in the attention projection layer; aligning contextual semantic features with visual feature dimensions to obtain cross-modal text features refers to using linear projection and layer normalization on contextual semantic features to obtain cross-modal text features aligned with visual feature dimensions. ; in, For cross-modal text features, For layer normalization, For contextual semantic features, and These represent the learnable projection parameters. To reduce training costs, this embodiment introduces an efficient LoRA parameter fine-tuning strategy in the attention projection layer of the large language model to reduce the number of trainable parameters while preserving pre-trained semantic knowledge.
[0036] In this embodiment, the fine-grained visual-linguistic fusion of object pseudomasks, visual features, and cross-modal text features for answer selection includes: S101, the object pseudomask is scale-matched to the same spatial resolution as the visual features, and then the scale-matched object pseudomask is coupled with the visual features to form coupled features through the object-guided attention module to highlight the object region and pixel-level semantic cues, which can be represented as: ; in This indicates a scaling or resampling operation used to adjust the object pseudomask M to match the visual features. Same spatial resolution, This refers to the Object-Guided Attention module. The Object-Guided Attention module is a well-known existing network module, and its implementation details will not be elaborated here. S102, the coupled features are processed through a multi-head self-attention operation and then concatenated with the original coupled features. The concatenated features are then processed through a feedforward neural network to obtain the visual features after visual object context modeling. , can be represented as: ; in, This refers to multi-head self-attention operations. This refers to a feed-forward network. S103, Visual features after modeling the context of a visual object Cross-modal text features are obtained by using a bidirectional cross-attention mechanism to obtain cross-modal enhanced visual features. and text features , can be represented as: ; ; in, Let Q be the attention calculation function, and let Q, K, and V be the query vector, key vector, and value vector, respectively. Visual features enhanced from text features These are the text features updated in reverse from the visual features; S104, visual features enhanced across modalities and text features After concatenation and pooling, the predicted probability distribution or predicted count result of the answer category is obtained through a classification head or a counting regression head: ; ; in, The predicted probability distribution of the answer category obtained from the classification head. For normalized activation functions, and These are the weight matrix and bias term of the classification head, respectively; The predicted count results obtained from the count regression head. and These represent the weight matrix and bias term of the counting regression head, respectively. This refers to the output features obtained from the pooling aggregation operation.
[0037] As an optional implementation, in this embodiment, the large model-guided hyperspectral image understanding method employs normalized difference loss to enhance the perception of counting error magnitude when training the overall network consisting of the three-dimensional U-shaped Mamba visual encoder-decoder network and the visual-language fine-grained fusion and answer selection components. ; in, To normalize the difference loss, Represents the normalization factor. For the number of categories, For class distance penalty function, For the true count category, Let be the predicted probability of the model for the k-th counting category; ; ; in, As a penalty factor, As a sensitive factor, Used to control the overall severity of punishment. Used to control the sensitivity to counting distance, Z is the normalization factor. When When the value is 0, the normalized difference loss degenerates into the ordinary classification cross-entropy loss.
[0038] To verify the effectiveness of the large model-guided hyperspectral image understanding method in this embodiment, this embodiment was validated on hyperspectral datasets such as the WHU-HI-HongHu dataset, the XiongAn dataset, and the Huston2013 dataset. Existing U-net, Segnet, Resnet, SSDGL, FPGA, Fusion Att, and SwinUMamba were used for comparison. Overall accuracy OA (%), average class accuracy AA (%), and Kappa coefficient κ were used as performance indicators. The comparative experimental results are shown in Tables 1 and 2.
[0039] Table 1: Comparison of results between the method in this embodiment and the comparative method in semantic segmentation tasks
[0040] As shown in Table 1, compared with existing U-net, Segnet, Resnet, SSDGL, FPGA, Fusion Att and SwinUMamba, the method in this embodiment achieves excellent results in overall accuracy OA (%), average class accuracy AA (%) and Kappa coefficient κ on the WHU-HI-HongHu dataset, XiongAn dataset and Huston2013 dataset, indicating that the semantic segmentation accuracy of the method in this embodiment reaches a high level on multiple datasets.
[0041] Table 2: Comparison of results between the method of this embodiment and the comparative method in the visual question answering task.
[0042] Table 2 shows the tasks designed for overall accuracy and overall root mean square error, including: Basic Judgment (Bas Ju), Inference Judgment (Rel Ju), Basic Counting (Bas Co), Complexity Analysis (Com An), Inference Counting (Rel Co), Object Analysis (ObjAn), Overall Accuracy (OA), and Overall Root Mean Square Error (OR). As can be seen from Table 2, compared with existing methods such as SSDGL, FPGA, Fusion Att, and ResNet, the method in this embodiment achieves excellent results in accuracy and overall mean square error for Basic Judgment (Bas Ju), Inference Judgment (Rel Ju), Basic Counting (Bas Co), Complexity Analysis (Com An), and Inference Counting (Rel Co) on the WHU-HI-HongHu and Huston2013 datasets. This indicates that under the same visual question answering framework, the method in this embodiment outperforms the main comparison methods in terms of both overall accuracy and overall mean square error. The experimental results show that the large model-guided hyperspectral image understanding method in this embodiment can effectively improve the semantic understanding and relational reasoning performance in visual question answering tasks while preserving the fine spatial-spectral representation capabilities of hyperspectral images, demonstrating promising application prospects.
[0043] Those skilled in the art will understand that the technical solutions provided by this invention may take the form of a method, system, or computer program product. Therefore, this invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this invention may take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce an implementation of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0044] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A large-model-guided hyperspectral image understanding method, characterized in that, The process includes the following steps: using a 3D U-shaped Mamba visual encoder-decoder network to perform hyperspectral visual encoding on the input hyperspectral image to obtain object pseudomasks and visual features; using a large model to perform language encoding on the input question text to obtain contextual semantic features; and aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features. Fine-grained visual-linguistic fusion and answer selection are performed on object pseudomasks, visual features, and cross-modal text features: S101, the object pseudomask is scale-matched to the same spatial resolution as the visual features, and then the scale-matched object pseudomask and visual features are coupled into coupled features through an object-guided attention module to highlight object regions and pixel-level semantic cues; S102, the coupled features are processed through a multi-head self-attention operation and then concatenated with the original coupled features. The concatenated features are then processed through a feedforward neural network to obtain the visual features after visual object context modeling. S103, Visual features after modeling the context of a visual object. Cross-modal text features are obtained by using a bidirectional cross-attention mechanism to obtain cross-modal enhanced visual features. and text features S104, enhancing visual features across modalities. and text features After concatenation and pooling, the predicted probability distribution or predicted count result of the answer category is obtained through a classification head or a counting regression head: ; ; in, The predicted probability distribution of the answer category obtained from the classification head. For normalized activation functions, and These are the weight matrix and bias term of the classification head, respectively; The predicted count results obtained from the count regression head. and These represent the weight matrix and bias term of the counting regression head, respectively. This refers to the output features obtained from the pooling aggregation operation.
2. The large-model-guided hyperspectral image understanding method according to claim 1, characterized in that, The step of using a three-dimensional U-shaped Mamba visual encoder-decoder network to perform hyperspectral visual encoding on the input hyperspectral image to obtain object pseudomasks and visual features includes: performing block preprocessing on the hyperspectral image and performing patch embedding to obtain initial visual features: ; in, As initial visual features, For patch embedding operations, For convolutional mappings used to unify spectral dimensions, To perform block operations based on preset size pairs, Given a hyperspectral image as input, the initial visual features are encoded and decoded using a 3D U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features.
3. The large-model-guided hyperspectral image understanding method according to claim 2, characterized in that, In step S102, the three-dimensional U-shaped Mamba visual encoder-decoder network is an encoder-decoder structure composed of stacked multi-level encoders and stacked multi-level decoders. The functional expression for the output features acquired by any i-th level encoder is: ; in, For the output features of the i-th level encoder, For the patch fusion module of the i-th level encoder, This is the visual state space layer of the i-th level encoder. This is the feature extraction layer of the i-th level encoder. For the input features of the i-th level encoder, , Let be the number of encoder layers; the function expression for obtaining the output features of any j-th level decoder is: ; in, For the output features of the j-th level decoder, For the j-th level decoder, the convolutional fusion module is used to map and optimize the fusion result of the upsampled features obtained by the upsampling module and the skip connection features; For upsampling module, For the input features of the j-th level decoder, Indicates splicing or residual fusion. For the skip connection features of the j-th level decoder, , The number of decoder layers; the functional expression for the three-dimensional U-shaped Mamba visual encoder-decoder network to obtain object pseudomasks and visual features is: ; ; in, For object pseudomask, For convolutional mapping operations used to generate object pseudomasks, For the first The output features obtained by passing the output features of the level decoder through a 1×1 convolution are as follows: As a visual feature, This is a linear projection operation used to project decoded features into a visual feature space. This represents the output feature of the first-level decoder.
4. The large-model-guided hyperspectral image understanding method according to claim 3, characterized in that, The feature extraction layer of the i-th level encoder processes the input features of the i-th level encoder as follows: Features are extracted using parallel shallow and deep branches respectively: ; ; in, The shallow spectral features are extracted from the shallow branches. Spatial-spectral context features extracted for deep branches, It is the ReLU activation function. This indicates the instance normalization operation. It is a 3×3×3 convolution. The convolution is 1×1×1; the shallow spectral features extracted from the shallow branches and the spatial-spectral context features extracted from the deep branches are aggregated into aggregated features: ; in, The aggregated features are used to generate the output features of the feature extraction layer of the i-th level encoder. ; in, The output features of the feature extraction layer of the i-th level encoder are denoted as .
5. The large-model-guided hyperspectral image understanding method according to claim 3, characterized in that, The visual state space layer of the i-th level encoder processes the input features by dividing the input features into multiple local windows: ; in, For the nth local window, For input features Window partitioning operation, dividing the window into 7×7 local windows, n=1,2,…,N w N w The number of local windows; each local window is input into multiple serial vision modules VSSBlock for processing: ; in, The local window features are obtained after processing by multiple vision modules VSSBlock in a serial manner; ~ These represent the VSSBlocks from the 1st to the Lth visual modules, where L is the total number of VSSBlocks in the sequence; the local window features of all local windows are merged to generate the merged window feature: ; in, To merge window features, This is a local window feature merging operation; among the multiple serial vision modules VSSBlock, the processing of input features by any vision module VSSBlock includes: linearly projecting the input features and then dividing them into main branch features and gated branch features: ; in, Main branch features, This is a gated branch feature. For the segmentation operation, The weight matrix is the linear projection. The main branch features are used as input features; local spatial-spectral context features are extracted from the main branch features through 3D deep convolution. ; in, For local spatial-spectral context features, This is a 3D depthwise convolution; the local spatial-spectral context features are transformed in multiple directions to obtain multiple directional sequences: ; in, For the m-th direction sequence, For feature flattening operation, For the m-th direction transformation of local spatial-spectral context features, the direction transformation includes some or all of the original, flipped, and rotated transformations; each direction sequence is fed into a Mamba module to perform a selective scan using the input-related state-space parameters to obtain the corresponding output features. ; Output features corresponding to each directional sequence After fusion, the gated branch features are modulated and projected back to the original feature dimensions: ; in, The output features of the vision module VSSBlock, To output the weight matrix, Representation layer normalization, To extract the output features corresponding to each directional sequence Summation operation after shaping This represents element-wise multiplication. The SiLU activation function is used. This is a gated branch feature.
6. The large-model-guided hyperspectral image understanding method according to claim 3, characterized in that, In any j-th level decoder, the upsampling module processes the input features by: adjusting the number of feature channels through convolution operations and performing preliminary integration of deep semantic information; and then using interpolation upsampling to restore the spatial size of the input feature map so that it is consistent with the skip connection features output by the corresponding coding layer in terms of spatial scale. ; in, This refers to the scale-aligned features obtained after interpolation and upsampling. For interpolation upsampling operation, For convolution operations, The input features of the upsampling module are used; the scale-aligned features obtained after interpolation and upsampling are fused with the skip connection features passed from the encoder: ; in, As a feature of fusion, For feature splicing operations, For the first The skip connection features from the layer encoder are processed by a residual block to obtain the features; the fused features can be processed through two residual blocks to enhance the feature representation capability and obtain the output features of the upsampling module. ; in, For the first The output features of the upsampling module in the layer decoder, and It is a residual block.
7. The large-model-guided hyperspectral image understanding method according to claim 1, characterized in that, The step of using a large model to encode the input question text to obtain contextual semantic features includes: segmenting, truncating, and padding the question text Q to obtain a word index sequence I and an attention mask A; inputting the word index sequence I and the attention mask A into a pre-trained large model to obtain contextual semantic features. The pre-trained large model is a large model fine-tuned and optimized by introducing a LoRA parameter fine-tuning strategy in the attention projection layer. The step of aligning the contextual semantic features with the visual feature dimensions to obtain cross-modal text features refers to using linear projection and layer normalization on the contextual semantic features to obtain cross-modal text features aligned with the visual feature dimensions. ; in, For cross-modal text features, For layer normalization, For contextual semantic features, and These represent the learnable projection parameters.
8. A large-model-guided hyperspectral image understanding system, comprising an interconnected microprocessor and a memory, characterized in that, The microprocessor is programmed or configured to execute the large model-guided hyperspectral image understanding method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the large model-guided hyperspectral image understanding method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the large model-guided hyperspectral image understanding method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Visual question-answering model training method and device, visual question-answering method and device, equipment and medium
CN113392253A
Remote sensing task general processing method combined with fine-grained feature enhancement
CN121883842A