Automatic multi-band image description method based on memory enhancement and soft mask
Patent Information
- Application Number
- CN202410822066.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-25
AI Technical Summary
[0006]本发明为了解决未配准多波段图像描述中往往存在目标幻影的问题,提出了一种基于内存增强和软掩膜的多波段图像自动描述方法
[0015]本发明实现了输入同一时刻下同一场景由不用成像设备拍摄的多波段图像,能够输出一条自然语言描述,为高精度智能探测系统提供技术支持。
Smart Images

Figure CN118736576B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image description methods and image fusion methods, and more particularly to a multi-band image automatic description method, specifically a multi-band image automatic description method based on memory enhancement and soft masking. Background Technology
[0002] Image captioning (IC), as a cross-modal task, aims to generate a natural language description for a given image. Image captioning not only requires the accurate detection and identification of targets, their attributes, and their interactions within the image, but also involves understanding and learning background knowledge. Image captioning technology has significant application value and research implications in scenarios such as medical auxiliary diagnosis, human-computer interaction, and security monitoring.
[0003] Broadband multispectral detection is one of the core technologies of high-end, intelligent equipment. Intelligence means the platform must be fully autonomous from perception and cognition to decision-making and action, and automatic image description is key to achieving this intelligence in such high-end equipment. Looking at existing image description technologies, apart from a small number of remote sensing images and medical images, they mainly focus on single visible light images. This is far from the practical needs of current high-precision intelligent detection systems that generally employ multiband imaging technology for all-weather detection, such as using infrared imaging to detect thermal targets, polarization imaging to detect certain metal minerals, and ultraviolet imaging to detect spark discharges. In these cases, describing only visible light images inevitably discards the complementarity of different bands, making it impossible to cope with the effects of camouflage, obstruction, long distances, darkness, and haze, thus leading to misjudgments and missed detections.
[0004] Thanks to breakthroughs in deep learning technology, image description techniques have made significant progress in terms of accuracy and information completeness. However, current mainstream image description models do not yet support multiple input images. This is mainly because multi-band images acquired by existing unmanned platforms are often unregistered, making them prone to producing target phantoms. Therefore, exploring automatic multi-band image description techniques is necessary.
[0005] To address this, a multi-band image automatic description method based on deep learning technology is proposed for simultaneous visible light and infrared imaging commonly used in high-end unmanned equipment platforms. This method improves the accuracy of image description in applications such as security monitoring and military reconnaissance, and can also broaden new areas of image description research, promote innovation in cross-modal understanding and generation technologies, and lay the foundation for intelligent decision-making research in related fields. Summary of the Invention
[0006] To address the problem of target phantoms often existing in the description of unregistered multi-band images, this invention proposes an automatic description method for multi-band images based on memory enhancement and soft masking.
[0007] This invention is achieved using the following technical solution: a multi-band image automatic description method based on memory enhancement and soft masking, comprising the following steps:
[0008] A multi-band image automatic description model was designed and constructed, consisting of a memory-enhanced encoder and a soft-mask-guided decoder. First, a Faster R-CNN model pre-trained on the Visual Genome dataset was used to extract features from visible light and infrared images. Then, based on the traditional Transformer, P learnable vectors were initialized as a memory-enhanced module to store the intrinsic correlations between multi-band image features and linguistic contextual information. Finally, a soft-mask mechanism was introduced in the decoder to filter visual features and memory vectors, ensuring that the model focuses on key visual or linguistic information.
[0009] A memory augmentation module is designed in the encoder, consisting of memory vectors, modal labels, and a linear layer. First, P (P=3) memory vectors with trainable attributes are randomly initialized. Then, modal labels are applied to visible light image features, infrared image features, and memory vectors to distinguish different types of inputs: visible light image features are labeled as 0, infrared image features as 1, and memory vectors as 2. Finally, the three types of labeled features are concatenated and input into the linear layer for dimensionality reduction.
[0010] A soft-mask-guided dual attention module is designed in the decoder. First, the encoder output features are mapped to key and value vectors. The text features processed by the masked multi-head attention layer are mapped to query vectors. Scaled dot product attention is used to calculate the similarity score between the query and key vectors, which is then normalized using the Softmax function and mapped to a probability distribution (attention weights). Next, the attention weights are filtered, retaining only the top k (k=9) most salient regions, while the rest are multiplied by a learnable mask. Finally, the attention weights are normalized using the Softmax function to generate a new attention weight, which is then multiplied by the value vector to further enhance the focus on key visual features.
[0011] The aforementioned multi-band image automatic description method based on memory enhancement and soft masking provides five natural language descriptions for each selected image pair. Each description includes rich texture information and details from visible light images, as well as thermal target information specific to infrared images. To enhance the richness and diversity of the samples, a series of data augmentation operations are performed on each selected image, including rotation, cropping, and radiometric transformation. Furthermore, back-translation technology (English-German-English) is used to enrich each description statement. After these processes, the original 2000 image pairs are expanded to 4000 pairs, generating a dictionary of size 3038, using special characters. <start>As the starting marker of the descriptive statement, <end>As the end marker of the descriptive statement, <pad>Characters are used to fill in sentences to ensure that the length of all descriptions meets the requirement of a maximum of 50.
[0012] The aforementioned multi-band image automatic description method based on memory enhancement and soft masks consists of two training phases. The goal of the cross-entropy training phase is to minimize the negative log-likelihood of words. Cross-entropy training stops when the metric CIDEr stops increasing after 5 epochs. The reinforcement learning optimization phase uses the metric CIDEr as the reward function and employs a gradient strategy algorithm to optimize and minimize the negative expected reward. Reinforcement learning training stops when the metric CIDEr stops increasing after 5 epochs.
[0013] The aforementioned multi-band image automatic description method based on memory enhancement and soft masks is experimentally performed using the PyTorch framework. The Dropout algorithm is used to prevent model overfitting, and the Adam gradient optimization algorithm is employed for parameter updates. A Faster R-CNN model combined with a ResNet-101 model fine-tuned on the Visual Genome dataset is used to extract feature vectors for each region, with a channel dimension of 2048. The multi-head attention mechanism in the Transformer is set to an 8-head attention mechanism, and the dimension of d_model is 512. A fixed value of 5e is used during the reinforcement learning optimization phase. -6 The learning rate.
[0014] The process of image description generation is essentially a word prediction process. Predicting each word is essentially classifying it, with the classification label being a word from the dictionary. Therefore, this invention applies cross-entropy loss at all time steps and accumulates it to represent the overall loss for the entire sentence. However, inconsistencies between the training and testing processes often lead to exposure bias. To address this, a self-critical training strategy is employed for optimization, using a consensus-based Image Description Evaluation (CIDEr) metric as the reward function for reinforcement learning. Experiments are conducted using the PyTorch framework, employing the Dropout algorithm to prevent overfitting and the Adam gradient optimization algorithm for parameter updates.
[0015] This invention enables the input of multi-band images of the same scene captured by different imaging devices at the same time, and outputs a natural language description, providing technical support for high-precision intelligent detection systems. Attached Figure Description
[0016] Figure 1 This is a diagram of the overall framework of the model.
[0017] Figure 2 This is a diagram of the encoder structure.
[0018] Figure 3 This is a diagram of the decoder structure.
[0019] Figure 4 This is a structural diagram of a dual attention module guided by a soft mask.
[0020] Figure 5 Example a is a schematic diagram of the result of multi-band image description.
[0021] Figure 6 Example b is a schematic diagram of the result of multi-band image description.
[0022] Figure 7 Example c is a schematic diagram of the result of multi-band image description. Detailed Implementation
[0023] A method for automatic description of multi-band images based on memory enhancement and soft masks includes the following steps:
[0024] 1. Design and build a memory-enhanced encoder.
[0025] The visual encoder consists of multiple Transformer encoding layers. A memory enhancement module is designed in the input section to model the intrinsic correlation between visible light image features and infrared image features. Simultaneously, this memory enhancement module effectively maintains and updates the state of the language context during subsequent text generation, thereby distinguishing between visual and linguistic features.
[0026] (1) Memory enhancement module
[0027] First, given a visible light image and an infrared image, feature extraction is performed using Faster R-CNN to obtain features V for the visible light image and features S for the infrared image. Wherein, m and n represent the number of visible light region features and infrared region features, respectively, and C represents the channel dimension. Let M represent the set of real numbers. The memory augmentation matrix contains a set of p randomly initialized memory vectors with trainable properties, denoted as M = {m1, m2, ..., mn}. p }, Modal labeling is applied to visible light image features, infrared image features, and memory features to distinguish different types of input data. The formula is as follows: V * =[V,token_v],S * = [S, token_s], M * = [M, token_m], where token_v represents visible light image feature tokens, token_s represents infrared image feature tokens, token_m represents memory feature tokens, and [·,·] represents the concatenation operation. Then, V * S * and M * The input features are obtained by concatenation. It can be represented as I * =[V * ,S * M * Finally, I * The joint input features are obtained by dimensionality reduction through a linear layer. C′ represents the number of channels after dimensionality reduction.
[0028] (2) Coding layer
[0029] First, the joint input features I are linearly projected into the query vector Q. e Key vector K e Sum vector V e Its formula is as follows: Q e =W1I,K e =W2I,V e =W3I, where W1, W2, and W3 represent learnable parameter matrices; secondly, scaled dot product attention is used to compute the query vector Q. e and bond vector K e The similarity score is then normalized using the Softmax function, mapping it to a probability distribution, and then compared with the value vector V. e A weighted sum is used to obtain a feature representation, which is formulated as follows: Where, d k Represents the scaling factor. This represents one feature representation; then, the calculation is repeated 8 times to obtain 8 feature representations, which are then concatenated and mapped through a linear layer to obtain the result of multi-head attention, as shown in the formula. Where Concat(·) is the concatenation operator, and W4 is the learnable matrix parameter. In the Transformer model, the application of layer normalization and residual structures is one of the key factors contributing to its excellent performance. Layer normalization, by standardizing the output of each layer, ensures the scale consistency of different features, thereby effectively improving the model's generalization ability. Simultaneously, residual connections allow gradients to be directly propagated to deeper layers, ensuring that the model can learn and capture deeper feature representations. Its formula is as follows: I′=LayerNorm(MHSA(Q e ,K e V e )+I), where I′ represents the feature enhanced by the multi-head attention mechanism; finally, the nonlinear expressive power of the model is increased by the feedforward neural network layer, which is formulated as FFN(I′)=ReLU(W5I′+b1)W6+b2, I″=LayerNorm(FFN(I′)+I′), where W5 and W6 represent learnable matrix parameters, b1 and b2 represent learnable bias parameters, and I″ represents the feature enhanced by the feedforward neural network layer.
[0030] 2. Design and build a soft mask-guided decoder
[0031] The text decoder consists of multiple Transformer decoding layers. In the traditional cross-attention mechanism, a Soft-mask-guided Dual Attention Module (SMDAM) is designed. This module uses the soft mask mechanism to process the cross-attention weights and, by redistributing the attention weights, achieves a prominent emphasis on important visual features.
[0032] (1) Decoding layer
[0033] First, given the words from the first t time steps, word embedding is performed to convert each word into a corresponding word vector representation. Next, these word vector representations are combined with positional encoding and used as input to the decoder, which can be represented as X. <t =Embedding(x <t )+pos, where x <t X represents the word at the previous t time steps. <t =(X0,X1,…,X) t-1 ) T Let X represent the corresponding word vector representation, and pos represent the corresponding positional encoding. Then, through a masked multi-head attention layer, SMDAM, and a feedforward neural network layer, the output result can be obtained, which is formulated as X. <t ′=maskMHSA(X <t ), y = SMDAM(X <t ′,I″),y * =FFN(y), where X <t ' represents the text features processed by the masked multi-head attention layer, and y represents the multimodal features processed by the soft-mask-guided dual attention module. * This represents the output features after processing by the feedforward neural network layer.
[0034] (2) Soft mask-guided dual attention module
[0035] First, the encoder output feature I″ is mapped to a key vector and a value vector, and the text feature X... <t The mapping '' to the query vector is formalized as Q. d =W7X <t ′,K d =W8I″,V d =W9I″, where W7, W8, and W9 represent the learnable parameter matrices. Next, scaled dot product attention is used to compute the query vector Q. d and bond vector K d The similarity scores are then normalized using the Softmax function, mapping them to a probability distribution. Next, the attention distribution (att) is filtered, retaining only the top k most salient regions, while the rest are multiplied by a learnable mask. Finally, it is normalized using the Softmax function to generate a new attention distribution, which is then compared with the value vector V. d Multiplication further enhances the focus on key visual features, and its formula is as follows: α = topK(att) ⊙ mask, y = Softmax(α ⊙ att) V d , where topK(·) represents selecting the top k important regions of the original attention distribution, and mask represents the learnable mask.
[0036] 3. Neural Network Training
[0037] First, the model is trained using cross-entropy loss (XE), and then optimized using reinforcement learning. The cross-entropy loss function is as follows: Among them, y 1:t-1 y represents the generated word sequence. t Let θ represent the word predicted at the current time step, T represent the sentence length, and θ represent the model parameters. The loss function for reinforcement learning is as follows: Among them, S i Let represent the i-th sentence, k represent the width of the beam search, and r(·) represent the reward function. Let θ represent the reward baseline, which is the mean of the reward for k-th beam search, and θ represent the model parameters.
[0038] 4. Automatic description of multi-band images based on memory enhancement and soft masking
[0039] (1) The Faster RCNN model was used in combination with a ResNet-101 model finely tuned on the Visual Genome dataset to extract visible light image features and infrared image features;
[0040] (2) Initialize the memory vector, place the visible light image features, infrared image features and memory vector at the same representation level, add modal labels to the three, perform feature dimensionality reduction through a linear layer, and input the result into the Transformer encoder;
[0041] (3) Design a soft mask-guided dual attention module to filter the original cross attention distribution, retain only the top k important regions, multiply the rest with the learnable mask, recalculate the attention distribution, and generate image descriptions.
[0042] The aforementioned multi-band image automatic description method based on memory enhancement and soft masks faces the challenge of a severe shortage of open-source datasets for describing unregistered visible-infrared images. Therefore, a meticulous manual screening process was conducted on existing visible-infrared image tracking datasets, including CAMEL Dataset, CSR_GTOT, and RGB-T234, resulting in 2000 pairs of nighttime scenes, images with occlusions, and low resolution. To meet the needs of unregistered images, a series of data augmentation techniques were applied, including rotation, cropping, and raying. Five detailed natural language descriptions were provided for each image pair, encompassing both the rich texture information of the visible light image and the thermal target information of the infrared image.
[0043] The aforementioned multi-band image automatic description method based on memory enhancement and soft masks follows the Transformer model's training strategy during the cross-entropy training phase; and uses 5e during the reinforcement learning optimization phase. -6 A fixed learning rate was used, and the Adam optimizer was employed for optimization, with the batch size set to 50 and the beam search size set to 5.
[0044] This invention relates to an automatic multi-band image description method, specifically an automatic multi-band image description method based on memory enhancement and soft masks. The invention proceeds according to the following steps: First, an unregistered visible light image-infrared image description dataset is constructed; second, an automatic multi-band image description model based on memory enhancement and soft masks is constructed; finally, the image description model is trained. The model flow is as follows: using a Faster R-CNN model pre-trained on the Visual Genome dataset, features are extracted from both visible light and infrared images; then, P learnable vectors are initialized as a memory enhancement module to store the deep correlation between visible light and infrared image features and the linguistic context; finally, a soft mask mechanism is introduced in the decoder to filter visual features and memory vectors. This invention achieves automatic description of unregistered multi-band images, alleviating the target phantom problem caused by unregistration, and providing technical support for high-precision intelligent detection systems.< / pad> < / end> < / start>
Claims
1. A method for automatic description of multi-band images based on memory enhancement and soft masks, characterized in that: Design and build a multi-band image automatic description generation model. The automatic description generation model is based on the Transformer architecture, including a visual encoder and a text decoder. The visual encoder consists of a memory enhancement module and a Transformer encoding layer. The text decoder includes a Transformer decoding layer. A soft mask-guided dual attention module is designed in the Transformer decoding layer. The memory enhancement module consists of memory vectors, modal labels, and a linear layer. P memory vectors with trainable attributes are randomly initialized. Then, modal labels are applied to visible light image features, infrared image features, and memory vectors to distinguish different types of inputs. Finally, the three types of labeled features are concatenated and input into the linear layer for dimensionality reduction to obtain joint input features, which are then input into the Transformer encoding layer. The Transformer decoding layer includes a masked multi-head attention layer, a soft-mask-guided dual attention module, and a feedforward neural network layer. It maps the encoder output features to key and value vectors. The text features processed by the masked multi-head attention layer are mapped to query vectors. A scaled dot product attention algorithm is used to calculate the similarity score between the query and key vectors, which is then normalized using the Softmax function and mapped to attention weights. Next, the attention weights are filtered, retaining only the k most salient regions, while the rest are multiplied by a learnable mask. Finally, the attention weights are normalized using the Softmax function to generate new attention weights, which are then multiplied by the value vectors to output multimodal features. These multimodal features are then processed by the feedforward neural network layer to generate an image description.
2. The multi-band image automatic description method based on memory enhancement and soft masking according to claim 1, characterized in that: For each selected image pair, five natural language descriptions are provided. Each description includes both the rich texture information and details in the visible light image and the thermal target information unique to the infrared image. To enhance the richness and diversity of the sample, a series of data augmentation operations are performed on each selected image, including rotation, cropping, and radiometric transformation. In addition, back-translation technology is used to enrich each descriptive statement.
3. The automatic multi-band image description method based on memory enhancement and soft masking according to claim 2, characterized in that: The multi-band image automatic description method consists of two training phases. The goal of the cross-entropy training phase is to minimize the negative log-likelihood of words. Cross-entropy training stops when the metric CIDEr stops increasing after 5 epochs. The reinforcement learning optimization phase uses the metric CIDEr as the reward function and employs a gradient strategy algorithm to optimize and minimize the negative expected reward. Reinforcement learning training stops when the metric CIDEr stops increasing after 5 epochs.
4. The automatic multi-band image description method based on memory enhancement and soft masking according to claim 3, characterized in that: The experiments were conducted using the PyTorch framework. Dropout was used to prevent overfitting, and Adam gradient optimization was employed for parameter updates. A Faster R-CNN model, finely tuned on the Visual Genome dataset, was used to extract feature vectors for each region, with a channel dimension of 2048. The multi-head attention mechanism in the Transformer was set to an 8-head attention mechanism, and the dimension of d_model was 512. A fixed value of 5e was used during the reinforcement learning optimization phase. -6 The learning rate.