Underwater image subtitle generation method and system based on multi-modal information fusion

Through multi-scale image feature extraction and hierarchical text feature fusion, combined with multi-head attention mechanism, the problem of insufficient text features and insufficient fusion in underwater image subtitles is solved, and high-quality underwater image subtitles are achieved.

CN120281862AActive Publication Date: 2025-07-08CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510405514.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In the existing underwater image subtitle generation technology, insufficient text feature extraction and insufficient image integration with text features lead to insufficient quality and accuracy of generated subtitles.

Method used

Multi-scale image feature extraction, hierarchical text feature extraction and multi-modal information fusion methods are used to extract multi-scale image features using Faster R-CNN and CLIP models, and text feature fusion is combined with K-mean clustering and multi-head attention mechanism to generate high-quality underwater image subtitles.

Benefits of technology

It improves the generation quality and accuracy of underwater image subtitles, enhances the understanding and description ability of the underwater environment, and the generated subtitles are more in line with grammatical norms and semantic correlation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281862A_ABST
    Figure CN120281862A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater image subtitle generation method and system based on multi-modal information fusion, and the method comprises the steps: firstly, extracting the multi-scale image features of an underwater image through a Faster R-CNN, including a full image feature and a region feature, and capturing the scene and salient target information of the underwater image; then, text word embedding codes related to the underwater image content are generated through a CLIP model, hierarchical text features are extracted through multi-level clustering by means of a K mean value, and the hierarchical structure of text information is further analyzed; then, a fusion method based on a multi-head attention mechanism is adopted, image features and text features are effectively fused, and the understanding ability of the model to underwater images is enhanced; and finally, inputting the fused multi-modal features into an image subtitle generator based on Transform, and generating underwater image subtitles related to image contents and contexts. According to the method, the accuracy and robustness of subtitle generation of the underwater image can be effectively improved, and the method has relatively high practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater image caption generation, and particularly relates to a method and system for underwater image caption generation based on multimodal information fusion. Background Art

[0002] The underwater space is a vast and mysterious environment, which is home to various organisms and rich natural resources. At present, many advanced technologies have been adopted in the fields of underwater photography and measurement, including optical imaging, laser imaging, and acoustic sensing, etc. Among them, underwater optical imaging is the most widely used due to its high resolution and ability to capture detailed image information. Underwater cameras mounted on equipment such as underwater robots or unmanned boats can obtain a large number of underwater optical images. However, interpreting and analyzing these images of marine scenes requires extensive professional knowledge and resources, especially in professional fields such as marine resource exploration, marine ecological protection, and coral reef monitoring. Based on these needs, underwater image caption generation technology can be used to describe and analyze underwater images. This technology can automatically generate accurate and detailed descriptions of marine scenes and targets through artificial intelligence algorithms, thus assisting professionals in marine environmental monitoring and resource management. This method of automatically generating descriptive captions can be embedded in underwater detection equipment or applied to batch image analysis, which not only improves work efficiency but also reduces the dependence on a large number of manual analyses. Therefore, it provides a more efficient and scalable solution for underwater detection and promotes the transition of underwater image analysis from visual perception to semantic understanding.

[0003] The system input of underwater image caption generation is an underwater optical image, and the output is a set of sentences that describe the image scene and content and conform to grammar norms and language fluency. It is a very typical cross-modal information processing system. Currently, the structure of an image encoder and a language generator is mostly adopted. The image encoder is used to extract features that effectively represent the information in the image, and the language generator synthesizes the information of the image encoder and outputs sentences describing the image content according to the semantics of the context. Specific implementation methods are as follows. How to represent the features of the image more accurately and comprehensively and how to generate sentences with complete semantics are the key points of research.

[0004] The limitations of existing underwater image caption generation technologies are on the one hand the lack of attention to the effective representation and feature extraction of text, and on the other hand the insufficient fusion of text features and image features.

[0005] (1) Insufficient consideration of text features

[0006] Underwater image caption generation is a typical cross-modal task, which involves generating text descriptions from images. However, traditional models often over-focus on optimizing the extraction of image features. For example, global and regional features based on convolutional neural networks can improve the representation performance of image features, while ignoring the role of text features in enhancing model performance. This limitation reduces the quality of the generated descriptions and affects the expressive ability of the words in the generated captions. Therefore, effectively extracting text features to improve underwater image caption generation methods poses a major challenge.

[0007] (2) Insufficient fusion of image and text features

[0008] After extracting text features, achieving deep fusion with image features to enhance the description ability is still a key challenge. Currently, most underwater image caption generation models rely on sequence generators such as the Long Short-Term Memory (LSTM) model for text generation; however, they show obvious limitations in processing multi-modal information and are difficult to fully capture the deep correlations between text and image features. Therefore, developing an efficient cross-modal alignment mechanism and designing a flexible feature fusion module are crucial for improving underwater image caption performance. Summary of the Invention

[0009] The present invention aims to provide an underwater image caption generation method and system based on multi-modal information fusion to solve the technical problems existing in the background technology.

[0010] To solve the technical problems, the technical solution of the present invention is as follows:

[0011] An underwater image caption generation method based on multi-modal information fusion, the method includes:

[0012] Multi-scale image feature extraction: Obtain an underwater image, and use the Faster Region-based Convolutional Neural Network (Faster R-CNN) to extract multi-scale image features, including full-image features and regional features;

[0013] Hierarchical text feature extraction: Use the Contrastive Language-Image Pre-Training (CLIP) model for text word embedding encoding to obtain semantic information associated with the image content, and perform hierarchical text feature extraction using K-means clustering, analyze the data distribution, and extract text features at different levels;

[0014] Multi-modal Information Fusion: Construct a multi-modal information fusion method based on the multi-head attention mechanism to effectively fuse image features and text features at different scales to obtain fused features;

[0015] Caption Generation: Based on the fused features, use a pre-set underwater image caption generator of Transformer to generate image captions that express the image content and are contextually related for underwater images.

[0016] Furthermore, in the image feature extraction stage, use Faster R-CNN for image feature extraction, with the pre-trained Residual Network (ResNet) ResNet101 as the backbone network; the multi-scale image features expand the global and local information expressions of underwater images from two aspects of full-image features and regional features;

[0017] Assume that an input is a three-channel true-color underwater image I. ResNet101 is a typical convolutional neural network containing residual modules. A series of convolutional operations are performed using different types of convolutional kernels, and it outputs a full-image feature map F, representing the scene information of the underwater image;

[0018] F = ResNet101(I). (1)

[0019] Faster R-CNN uses the Region Proposal Network (RPN) to generate N proposed target regions, denoted as {r1, r2,..., r N}, which predicts whether each position contains a target and calculates the offset of the bounding box of the target. Region of Interest Pooling (RoI Pooling) is used to process regions of interest with different sizes of inputs, and it outputs N region feature vectors labeled as R, corresponding to the N proposed targets, as follows:

[0020] R = RoIpooling(r1,r2,...,r N ). (2)

[0021] Finally, through a fully connected network combined with bounding box regression and the softmax function, the bounding boxes and categories of candidate targets are predicted, and the multi-scale image information provides support for word sequence prediction in the underwater image caption generator.

[0022] Furthermore, the text feature extraction stage specifically includes:

[0023] Word Embedding Encoding Based on CLIP:

[0024] Tokenize and count the text annotated in the underwater image caption dataset to construct a corresponding vocabulary. The words in the vocabulary are encoded with word embeddings through a pre-trained CLIP model to generate word vectors related to the image content. The shape of the word vectors is [N×512], where N represents the number of words in the vocabulary, and 512 is the dimension of the word embedding vector output by the CLIP model;

[0025] Generate text clustering features:

[0026] Use the K-means algorithm to cluster the word embedding vectors generated by CLIP. First, perform low-level clustering to cluster the word vectors of all words to obtain low-level clustering centers. The shape of the clustering result is [N1×512], where N1 represents the number of low-level clustering centers;

[0027] Generate high-level clustering:

[0028] Based on the vector set of low-level clustering centers, perform high-level clustering; further cluster the low-level clustering center vectors to obtain high-level clustering centers. The shape of the clustering result is [N2×512], where N2 represents the number of high-level clustering centers;

[0029] The multi-level clustering is respectively called "low-level clustering" and "high-level clustering" to obtain "hierarchical text features".

[0030] Text feature representation of hierarchical clustering centers:

[0031] The clustering center represents the samples similar to it. Therefore, each clustering center serves as the classification feature of these words to enhance the expression of text semantics. The low-level clustering center represents more detailed semantic categories, while the high-level clustering center represents more abstract and broad categories;

[0032] Information fusion:

[0033] Gradually fuse the text features of low-level clustering and high-level clustering with the features of underwater images to provide rich semantic support for the subsequent image caption generation task; through the above steps, the word embeddings generated by CLIP and the two-level clustering method based on the K-means algorithm effectively extract the hierarchical text features of underwater image captions and enhance the multi-modal information association between text and images.

[0034] Furthermore, the multi-modal information fusion stage specifically includes:

[0035] Assume that at the time step t of the underwater image caption generation model, the full-image feature F t and the region feature R t =[r1, r2,..., r n, extracted through object detection, and then they are concatenated to form a combined image representation, denoted as V t , where V t aggregates multi-scale image information and is defined as:

[0036] V t = [F t , R t . (3)

[0037] In the multi-head attention module, by taking the image information V t and the output of the j-th level text clustering module, representing the extracted text feature vector C j , the attention mechanism is applied to achieve feature fusion, and the attention weights are calculated as:

[0038]

[0039] where W q and W k are the query and key-value matrices for implementing linear dimensional transformation in the self-attention mechanism, the parameter d k is the dimension of (W k C j ), the parameter α t represents the attention weight, quantifying the correlation between the combined full-image features and region features of V t and the text clustering feature C j ;

[0040] Next, the weighted text clustering feature is calculated using the following formula:

[0041]

[0042] In the formula represents the weighted text clustering feature, and W v represents the linear transformation matrix;

[0043] Subsequently, and V t perform a residual connection and then layer normalization, and the formula is:

[0044]

[0045] where the function LayerNorm(·) represents the layer normalization operation, which scales and offsets based on the learnable parameters of all samples, X t is the output of the layer normalization operation. Layer normalization stabilizes the output distribution by normalizing the output values of the activation function within each batch, thereby accelerating the training process and enhancing the generalization ability of the model;

[0046] X tIt is fed into a feed-forward neural network, then added to itself, and layer normalization is performed for normalization;

[0047] Fusion vt = LayerNorm(X t + FFN(X t ))), (7)

[0048] where Fusion vt is the feature of the fusion of image and text information. The function FFN(·) represents the operation of a two-layer feed-forward network, which is a supplement to the self-attention mechanism and enhances the expressive ability of the model. Finally, the output fusion feature Fusion vt is sent to a Transformer-based underwater image caption generator to predict the next word in the sentence sequence;

[0049] Through two-level clustering, text feature vectors of two sets of clustering centers are obtained. During the information fusion process, they are combined sequentially level by level according to the global attention mechanism fusion mode to obtain a vector that fuses more feature information. A hierarchical information fusion strategy is adopted. The text clustering features of the higher level are fused with the image features of the previous layer, mainly focusing on identifying the main categories of words. The text clustering features of the lower level are fused in the subsequent layers to express more specific information, and finally the fusion feature of the image and text is obtained.

[0050] Furthermore, the model training in the caption generation stage specifically includes:

[0051] Assume that the description y of the underwater image caption generation model for image I is:

[0052] y = (y1, y2,..., y T )(8)

[0053] where y i is the i-th word generated, and T is the length of the description;

[0054] The training process of the underwater image caption generator includes two stages: cross-entropy optimization and optimization of the evaluation metric CIDEr of the similarity between the image description text and the reference value. In the cross-entropy optimization stage, the cross-entropy loss function is used to quantify the ability of the model to predict the word sequence, and its definition is:

[0055]

[0056] where y t is the output of the generator at time t, y 1:t-1 is the output of the generator before time t. Given the parameters θ of the underwater image caption generation model, the symbol p θ represents that the output of the generator is yt and the previous word is y 1:t-1 probability distribution;

[0057] The generator is trained by minimizing the cross-entropy loss to generate sentences similar to the reference sentences. When the cross-entropy optimization reaches a certain level and overfitting or training stagnation occurs, reinforcement learning techniques are further applied to directly optimize the numerical value of the evaluation parameter CIDEr of the generated text;

[0058] The Self-Critical Sequence Training (SCST) method is adopted to optimize the model based on the expected reward. The loss function of reinforcement learning is defined as:

[0059]

[0060] where the reward function r(·) is the CIDEr parameter value, which is used to measure the consistency between the description generated by the model and the reference text. By maximizing the mathematical expectation of the CIDEr parameter, the ability of the model to generate high-quality underwater image captions is further improved.

[0061] An underwater image caption generation system based on multi-modal information fusion, the system is applied to any one of the above methods, and the system includes:

[0062] Image feature extraction module: Obtain underwater images and use Faster R-CNN to extract multi-scale image features, including full-image features and regional features;

[0063] Text feature extraction module: Use CLIP to perform text word embedding encoding to obtain semantic information associated with the image content, and adopt K-means clustering for hierarchical text feature extraction, analyze the data distribution, and extract text features at different levels;

[0064] Multi-modal information fusion module: Construct a multi-modal information fusion method based on the multi-head attention mechanism to effectively fuse image features at different scales and hierarchical text features to obtain fusion features;

[0065] Caption generation module: Based on the fusion features, use a preset underwater image caption generator of Transformer to generate image captions that express the image content and are contextually related for underwater images.

[0066] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements an underwater image caption generation method based on multi-modal information fusion described in any one of the above.

[0067] A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the method for generating underwater image captions based on multimodal information fusion described in any one of the above.

[0068] Compared with the prior art, the advantages of the present invention are as follows:

[0069] (1) A new method is developed to extract hierarchical text features to enhance the ability of underwater image caption generation. The CLIP model based on contrastive learning for language-image pre-training is used to obtain word embedding encodings with semantic contexts consistent with underwater image content. This method helps generate captions closely corresponding to the image content. In addition, a hierarchical clustering method is implemented to extract hierarchical text features, which capture features such as typical underwater objects, scenes, and activities in the text. This achieves more accurate text prediction by using semantic similarity and promotes a deeper understanding of the context relationship of underwater image descriptions.

[0070] (2) A multimodal information (i.e., image and text features) fusion method based on the multi-head attention mechanism is developed to effectively fuse full-image and region features with hierarchical text clustering features. By adopting an attention-based weight distribution in the underwater image caption generation task, it focuses on significant objects and their related semantic text features. Therefore, this method can generate more accurate and grammatically correct underwater image captions. Description of the Drawings

[0071] Figure 1 The technical roadmap of the method for generating underwater image captions based on multimodal information fusion of the present invention;

[0072] Figure 2 The multi-scale image feature extraction diagram based on Faster RCNN;

[0073] Figure 3 The CLIP word embedding encoding of the text in the underwater image caption corresponding dictionary and the hierarchical text clustering diagram;

[0074] Figure 4 The multimodal information fusion diagram of image and text features based on the multi-head attention mechanism;

[0075] Figure 5 The multi-level image and text feature fusion diagram based on the attention mechanism;

[0076] Figure 6 The visualization example diagram of underwater image detection based on the Faster R-CNN model;

[0077] Figure 7, the content word and two - level clustering center diagram of the underwater image caption dataset thesaurus;

[0078] Figure 8 , the influence diagram of the number of high - level text clustering centers on the model performance. Specific implementation manners

[0079] The specific implementation manners of the present invention will be described below in conjunction with embodiments:

[0080] It should be noted that the structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0081] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" cited in this specification are only for the convenience of clear narration, and are not used to limit the scope of implementation of the present invention. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope where the present invention can be implemented.

[0082] Embodiment 1:

[0083] As Figure 1 shown, this embodiment provides an overall framework of an underwater image caption generation method based on multi - modal information fusion;

[0084] First, in the image feature extraction stage, Faster R - CNN is used to extract multi - scale image features of underwater images, including full - image features and regional features. These features effectively capture the information of the scene and significant objects in the image. At the same time, in the text feature extraction stage, the CLIP model based on contrastive learning is used to obtain the word - embedding encoding associated with the image content. Subsequently, the K - means clustering algorithm is used to analyze the data distribution and extract hierarchical text features related to the text. Then a method based on the multi - head attention mechanism is developed to effectively fuse the multi - modal information of the extracted various image and text features. Finally, these features are input into a Transformer - based language generator to generate image captions that express the image content and are context - related for underwater images.

[0085] I. Multi - scale image feature extraction

[0086] Faster R-CNN has become one of the popular solutions for real-time object detection tasks due to its balanced characteristics of high accuracy and reasonable speed. For the underwater image caption generation task, Faster R-CNN is used for image feature extraction, with the pre-trained ResNet101 as the backbone network. The parameters in the first few layers exhibit good generalization ability, enabling it to handle the basic features of underwater images, while the subsequent layers are fine-tuned on the underwater image caption dataset to capture more detailed and task-specific features. These extracted features provide rich image information to support the underwater image caption generation task. The multi-scale image features expand the global and local information expression of underwater images from both the full-image features and regional features. The specific process is as Figure 2 shown:

[0087] Assume an input three-channel true-color underwater image I. ResNet101 is a typical CNN containing residual modules, which performs a series of convolution operations using different types of convolutional kernels. It outputs the full-image feature map F, representing the scene information of the underwater image, as shown in the figure:

[0088] F = ResNet101(I). (1)

[0089] Faster R-CNN uses the Region Proposal Network (RPN) to generate N proposed target regions, denoted as {r1, r2,..., r N}. It predicts whether each location contains an object and calculates the offset of the bounding box of the object. Region of Interest (RoI) pooling is used to process the regions of interest with different sizes of inputs. It outputs N region feature vectors labeled as R, corresponding to the N proposed targets, as follows

[0090] R = RoIpooling(r1,r2,...,r N ). (2)

[0091] The final fully connected network, along with bounding box regression and the softmax function, predicts the bounding box candidate targets and the categories of the objects. All of these multi-scale image information can be used in the underwater image caption generation task, enabling it to be fused and used to predict the word sequence in the underwater image caption generator.

[0092] In summary, this module not only extracts the full-image features of the underwater image to represent the scene information, but also obtains the corresponding regional features through the region proposal network. These multi-scale image features provide rich image representations for subsequent multi-modal information fusion, thereby enhancing the understanding ability of the underwater image caption generation model for complex underwater environments.

[0093] II. Hierarchical Text Feature Extraction

[0094] CLIP-based Word Embedding Text Feature Representation:

[0095] The encoding of text information is crucial for underwater image captioning. Word embedding technology maps words into a continuous high-dimensional space, generating dense vector representations that facilitate semantic understanding. To more effectively extract text features and establish the relationship between text and images, the CLIP model emerged. By training on a large-scale dataset of images and their corresponding text descriptions, CLIP can effectively capture the semantic associations between text and images, significantly enhancing image understanding and text generation capabilities.

[0096] Underwater image captioning faces complex visual environments and special imaging conditions, which pose higher requirements for the model's multi-modal information understanding ability. Considering the advantages of CLIP in semantic extraction and cross-modal features, this model performs very well in addressing the challenges of underwater image caption generation. Therefore, using the word embeddings generated by CLIP as text features can effectively improve the performance of the model in the underwater image caption generation task.

[0097] Hierarchical Text Feature Extraction Based on Two-level Clustering:

[0098] The regional features derived from underwater images correspond to the objects and categories present in the image content. When fusing information from text, it is important to not only consider context encoding but also add category-specific information, which reflects higher-level abstract information in the text data. In the absence of predefined text classification labels, an unsupervised clustering method can be used to generate a sample set of word vector content words, where the cluster centers represent the category features of semantically similar words. A process based on the K-means algorithm was implemented to perform this.

[0099] During the clustering process, it is not easy to accurately determine the number of cluster centers. An inappropriate number may lead to unclear boundaries between clusters. Single-level clustering may not be able to capture the complexity of the data. To achieve more reasonable sample clustering, multi-level clustering can be used to obtain better text clustering features, further reducing the differences within clusters, thereby improving the final clustering quality. It can also reveal complex patterns in the dataset, such as nested clustering, sub-cluster structures, etc. Lower-level clustering may result in broader categories, while higher-level clustering can help find more representative subgroups within each cluster. This can make the final clustering results easier to interpret.

[0100] Adopt such as Figure 3The two-stage K-means clustering operation shown above. In the first stage, the word embedding vectors of all words in the thesaurus are clustered to obtain N1 clustering centers, called low-level clustering centers, denoted as C1. In the second stage, the vector set of the first-stage clustering centers is clustered to obtain N2 clustering centers, called high-level clustering centers, denoted as C2.

[0101] The clustering center represents samples similar to it, so it can be used as a classification feature for these words. When the present invention uses two-stage clustering, it enhances the hierarchical relationship in semantic expression, making the text information more comprehensive. They will be used for the fusion of image and text information.

[0102] III. Multimodal Information Fusion Based on Multi-Head Attention Mechanism

[0103] The multi-head attention mechanism performs excellently in modeling long-term dependencies and capturing subtle context changes. It can simultaneously focus on multiple aspects of the input, thereby enhancing the ability to model complex dependencies within the data, and thus achieving better performance in generation tasks. Therefore, in underwater image caption generation, the Transformer model is used to replace the traditional LSTM to improve the efficiency and accuracy of caption generation. However, relying solely on image features may not be sufficient to fully understand the complex underwater situation. However, relying solely on visual features may not be sufficient to fully understand the complex underwater scene. Therefore, a multimodal information fusion method based on the multi-head attention mechanism is proposed, which effectively fuses the full-image features, regional features, and hierarchical clustering text features. By making full use of the complementarity of different modalities, this method enhances the model's understanding and description ability of underwater images. The overall fusion process is as Figure 4 shown, which demonstrates how to fuse text and visual features using the multi-head attention mechanism.

[0104] Assume that at time step t of the underwater image caption generation model, the full-image feature is F t , the regional feature R t = [r1, r2,..., r n , and the regional features are extracted through object detection and then concatenated to form a unified image representation. This combined representation is V t , where V tt aggregates multi-scale image information and is defined as:

[0105] V t = [F t , R t . (3)

[0106] In the multi-head attention module, through the image information V t and the output of the j-th layer word clustering module (representing the extracted text feature vector C j) Feature fusion is achieved by applying the attention mechanism. The attention weights are calculated as

[0107]

[0108] where W q and W k are the query and key-value matrices for implementing linear dimensional transformation in the self-attention mechanism. The parameter d k is the dimension of (W k C j ). The parameter α t represents the attention weight, which quantifies the correlation between the combined global image and regional feature V t and the text clustering feature C j .

[0109] Next, the weighted text clustering feature is calculated using the following formula

[0110]

[0111] In the formula represents the weighted text clustering feature, and W v represents the linear transformation matrix

[0112] Subsequently, and V t perform a residual connection and then layer normalization. The formula is

[0113]

[0114] where the function LayerNorm(·) represents the layer normalization operation, which scales and offsets based on the learnable parameters of all samples. X t is the output of the layer normalization operation. Layer normalization stabilizes the output distribution by normalizing the output values of the activation function within each batch, thereby accelerating the training process and enhancing the generalization ability of the model.

[0115] X t is fed into the feed-forward neural network, then added to itself, and layer normalization is performed for normalization, as shown in the figure

[0116] Fusion vt =LayerNorm(X t +FFN(X t ))), (7)

[0117] where Fusion vt is the feature of the fusion of image and text information. The function FFN(·) represents the two-layer feed-forward network operation, which is a supplement to the self-attention mechanism and can enhance the expressive ability of the model. Finally, the output fusion feature Fusion vtinto the Transformer-based generator to predict the next word in the sentence sequence.

[0118] Two-level clustering is performed to obtain text feature vectors of two sets of clustering centers. During the information fusion process, they can be fused step by step according to the global attention mechanism fusion mode to obtain a vector that fuses more feature information, as Figure 5 shown

[0119] Inspired by the cognitive process of humans from general reasoning to specific reasoning, the method proposed in the present invention adopts a multi-level information fusion strategy. The higher-level text clustering features are fused with the image features of the previous layer, mainly focusing on identifying the main categories of words, and the lower-level text clustering features are fused in the subsequent layers to express more specific information.

[0120] IV. Training Based on Reinforcement Learning

[0121] Suppose the description y of the underwater image caption generation model for generating image I is:

[0122] y = (y1, y2,..., y T ) (8)

[0123] where y i is the i-th word generated, and T is the length of the description.

[0124] The training process of the underwater image caption generator includes two stages: cross-entropy optimization and optimization of the evaluation metric CIDEr of the similarity between the image description text and the reference value; in the cross-entropy optimization stage, the cross-entropy loss function is used to quantify the ability of the model to predict the word sequence, and its definition is:

[0125]

[0126] where y t is the output of the generator at time t, and y 1:t-1 is the output of the generator before time t. Given the parameters θ of the underwater image caption generation model, the symbol p θ represents the probability distribution that the output of the generator is y t , and the previous word is y 1:t-1 .

[0127] The generator is trained by minimizing the cross-entropy loss to generate a sentence similar to the reference sentence. When the cross-entropy optimization reaches a certain level and overfitting or training stagnation occurs, the reinforcement learning technique is further applied to directly optimize the value of the evaluation parameter CIDEr of the generated text; CIDEr is an evaluation parameter widely used to evaluate image descriptions, and it is mainly used to evaluate the similarity between the generated sentence and the annotation.

[0128] The Self-Critical Sequence Training (SCST) method is adopted to optimize the model based on the expected reward. The loss function of reinforcement learning is defined as:

[0129]

[0130] Among them, the reward function r(·) is the CIDEr parameter value, which is used to measure the consistency between the descriptions generated by the model and the annotations. By maximizing the mathematical expectation of the CIDEr parameter, the ability of the model to generate high-quality underwater image captions is further improved.

[0131] Example 2:

[0132] This second example is applied to the first example. The implementation process of the technical solution of this example specifically includes:

[0133] Dataset

[0134] The experiment was carried out on an underwater image caption dataset newly annotated with images from the Underwater Image Enhancement Benchmark (UIEB). Each image was manually annotated with 5 different English sentences. The 915 underwater images in the dataset included 765 training samples, 90 validation samples, and 60 test samples divided. The images in the UIEB dataset contain a variety of real underwater scenes with rich target categories, including typical targets such as divers, different types of fish and corals, sea turtles, coral reefs, shipwrecks, antiques, and stone statues. This diversity makes it particularly valuable, ensuring the diversity and variability of the training samples.

[0135] The underwater image caption dataset contains a total of 49,278 words, of which 1,469 words are unique. These unique words constitute the candidate vocabulary library for caption generation. Figure 7 The statistical distribution of the 20 most frequently occurring words is given. It is worth noting that these high-frequency words are mainly related to the marine environment and underwater activities, including terms such as "sea", "fish", "coral", "diver", "coral reef", and "rock". In addition, the part-of-speech analysis of the words in the dataset is carried out, and the approximate percentage distribution of each part-of-speech category is shown in Table 1. The most frequently occurring categories include nouns (noun), proper nouns (PROPN), verbs (VERB), adjectives (ADJ), and adverbs (ADV), which constitute the core vocabulary components of the sentence structure. In addition, function words such as determiners (DET) and adpositions (ADP) also exist, which help to explain the text of the grammatical structure.

[0136] Table 1 - Percentage of part-of-speech categories

[0137]

[0138] Experimental setup

[0139] The experiment was conducted on a computer equipped with an Intel(R) Xeon(R) Silver 4214R CPU @ 2.40 GHz, 128 GB of RAM, and an NVIDIA GeForce RTX 2080Ti GPU. The operating system used was Ubuntu 20.04, PyTorch 1.0.4 as the deep learning framework, and CUDA 10.1 for GPU acceleration.

[0140] For the training process of the developed underwater image caption generation model, the model parameters were set as follows: the dimensions of both the full-image feature vector and the region feature vector were 2048, while the dimension of the contrastive learning-based word embedding vector was 512. In the two-level text clustering, the lower level contained 100 cluster centers, and the higher level contained 10 cluster centers. The text features and image features were fused using a three-layer structure, and the three-layer text cluster centers were set to 10, 10, and 100 respectively. The dimension of the Transformer-based generator was 512, with 8 attention heads, and the internal dimension of the feed-forward network module was 2048. The training parameters were set as the batch sample size of 25 samples per batch, the initial learning rate of cross-entropy training was 1e-4, and the initial learning rate of reinforcement learning was 5e-6. The training patience was set to 5, which means that if the cross-entropy loss was less than the threshold for 5 consecutive generations of gradients, it would switch to the reinforcement learning mode, or if the CIDEr value decreased for 5 consecutive generations, the training process would stop.

[0141] Evaluation metrics

[0142] To ensure objective evaluation, the present invention adopted four widely used standard metrics: BLEU-N, METEOR, ROUGE-L, CIDEr, and S*, to comprehensively evaluate the quality of automatically generated underwater image captions. For any given underwater image, a labeled reference caption was used to compare with the candidate caption generated by the model. The definitions of these five metrics are as follows:

[0143] (1) BLEU-N is used to evaluate the matching degree between the candidate caption and the reference caption. It is a precision-based metric that measures the overlap of N-grams between the candidate caption and one or more reference captions.

[0144] (2) METEOR calculates the harmonic mean of the precision and recall between the candidate points and the reference points. It compares the generated caption with the reference caption based on multiple language factors (including synonym matching, stemming, word order, and paraphrasing) to evaluate the generated caption.

[0145] (3) ROUGE-L measures the similarity between candidate captions and reference captions by calculating the F-measure based on the longest common subsequence, effectively capturing overlapping information.

[0146] (4) CIDEr treats each caption as a document, calculates the cosine similarity of the term frequency-inverse document frequency vectors, and aggregates the similarities of n graphs of different lengths to obtain the final score.

[0147] (5) S* is defined as the average of four evaluation metrics: BLEU-4, METEOR, ROUGE-L, and CIDEr. It provides a comprehensive evaluation of the performance of the generated candidate captions.

[0148] By using multiple evaluation metrics, this study comprehensively evaluated the development of an underwater image caption generation model from multiple perspectives, thus more accurately reflecting its overall effectiveness. For consistency and clarity, the results of subsequent experiments are presented in decimal form, with higher scores indicating better performance.

[0149] Image Feature Extraction

[0150] In the image feature extraction stage, the input underwater images are uniformly resized to 224×224 pixels. Then, the pre-trained ResNet101 backbone is used to extract the full-image features, obtaining 2048-dimensional feature vectors. Next, the present invention uses Faster R-CNN to extract the target region features from the underwater images. Each region is represented by a 2048-dimensional feature vector, and the region proposal network, RoI pooling, and fully connected layers are used to obtain the bounding box coordinates and target category information.

[0151] Figure 6 The detection results of two underwater sample images are shown, where target regions such as sunken shipwrecks and rocks are effectively detected. However, due to the limited number of underwater scene samples in the pre-trained dataset, there may be differences between the labeled categories and the actual categories. For example, in Figure 6 (a), the sunken ship is misidentified as an elephant, while in Figure 6 (b), the coral is misidentified as a tree. However, this inconsistency does not affect Figure 6 the effectiveness of the region feature vectors in underwater image caption generation. In addition, the attribute information of significant objects, such as "brown", "dark", "large", and "rocky", is usually accurate, providing valuable support for underwater image caption generation.

[0152] Text Feature Extraction

[0153] In the candidate word library, 317 independent words are identified, including adjectives, comparative adjectives, superlative adjectives, nouns, plural nouns, proper nouns, and verbs. Since there are no pre-defined lexical text classification labels marked, the present invention applies a clustering method to extract the text features of these words. The word embedding encoding of the clustering centers represents underwater objects and activities, which will all be fused with visual features to enhance the ability to generate underwater image captions.

[0154] Figure 7 Hierarchical text clustering centers are given, which consist of nouns, verbs, and adjectives related to the ocean, marine life, diving, underwater exploration, and related fields. The word clustering centers have the following characteristics, mainly referring to marine life (such as sharks, eels, fish) and the marine environment (such as corals, seagrass, shipwrecks). Other nouns and verbs are related to diving activities, underwater expeditions, and human equipment (such as gloves, clothes, boxes, equipment). Adjectives and adverbs mainly describe the attributes of objects. Terms such as cylindrical, thin, huge, colorful, and golden describe the shape and attributes of objects. These text features show the significant features of typical underwater objects, scenes, and activities. When these clustering features are fused with visual features to produce an output, the predicted words will also contain underwater information corresponding to the text-visual fusion features.

[0155] Comparison with typical image caption generation models

[0156] The underwater image captioning model developed in this invention is compared with some typical image captioning models, including the Neural Image Captioning (NIC) (Vinyals et al., 2015), Soft attention captioning model (Xu et al., 2015), Bi-directional LSTM (Bi-LSTM) captioning model (Wang et al., 2016), Adaptive attention captioning model (Lu et al., 2017), Spatial attention (Lu et al., 2017), Up-Down attention captioning model (Anderson et al., 2018), Attention on Attention Network (AoANet) captioning model (Li et al., 2018), and Meshed-Memory Transformer (M2 Transformer) captioning model based on grid memory (Cornia et al., 2020). As shown in Table 2, the experimental results of the underwater image captioning model proposed in this invention and some typical image captioning models on the underwater image captioning dataset are presented. The best results of the evaluation parameters are presented in bold. The arrow "↑" on the right side of each evaluation parameter in the first row of the table indicates that the "(RL)" model of this invention has improved over other models in this value.

[0157] Table 2 - Experimental results of image captioning generation models on the underwater image captioning dataset

[0158]

[0159] The NIC image captioning model used as a baseline shows relatively low performance, with scores on all evaluation metrics lagging behind the subsequently improved models. After introducing the attention mechanism, the soft-attention captioning model shows improvements in metrics such as BLEU and CIDEr. However, its overall performance is still limited. The Bi-LSTM captioning model enhances the ability to model context information, resulting in a slight improvement in BLEU and METEOR values. The adaptive attention model and spatial attention model further improve the evaluation parameter values by refining the attention mechanism. However, their generalization ability in complex underwater scenarios is still limited. The bottom-up and top-down attention captioning models significantly improve by fusing regional and global features, producing strong results in the BLEU, METEOR, and CIDEr metrics. The AoANet captioning model enhances the attention to image details through a multi-layer attention mechanism. M2 Transformer is an image captioning based on Transformer that uses multi-scale feature fusion to achieve near-optimal BLEU and CIDEr scores.

[0160] Based on these studies, the underwater image captioning model developed in the present invention, after being trained with cross-entropy (XE) and reinforcement learning (RL), achieves state-of-the-art performance on most evaluation metrics. After being optimized using reinforcement learning, it further improves the quality of underwater image captions and leads in terms of BLEU, METEOR, ROUGE-L, CIDEr, and S* scores. This shows that the captions generated by the model of the present invention not only align more closely with the reference captions but also exhibit greater semantic accuracy and diversity, enabling it to more effectively capture the fine-grained details of complex underwater scenarios.

[0161] Ablation experiments

[0162] The present invention conducts a series of ablation experiments to further verify the effectiveness of the developed underwater image captioning model based on multi-model information fusion. Specifically:

[0163] (a) Influence of text feature fusion

[0164] The present invention uses a Transformer as a generator and studies the role of text features in the underwater image caption generation model developed in the present invention. The experimental results are shown in Table 3. The analysis evaluation indicators show that the evaluation parameters have relatively low performance when only using regional features as the model input, especially in BLEU-4 (0.1994) and CIDEr (0.2753). This indicates that the caption generation model using only regional features has obvious limitations in generating high-quality underwater image captions. The model of the present invention performs multi-modal information fusion of images and texts and is trained using cross-entropy loss. Table 3 shows the improvements in various evaluation indicators, especially in BLEU-2, BLEU-3, BLEU-4, and CIDEr, where its performance is significantly better. It is worth noting that CIDEr increases from 0.2753 to 0.5214, indicating a substantial improvement in caption generation quality. In addition, the model trained by the present invention using reinforcement learning outperforms the model trained by cross-entropy and the baseline in all indicators and achieves the best performance under various evaluation criteria. Overall, these results verify the effectiveness of multi-modal information fusion, which combines multi-scale image features and text clustering features to enhance the performance of underwater image caption generation.

[0165] Table 3 - Experimental Results of the Effect of Text Feature Fusion in the Underwater Image Caption Dataset

[0166]

[0167] "Transformer (Reg.)" refers to the model that only uses regional features. "The present invention's (XE)" and "The present invention's (RL)" refer to the model of the present invention that combines text features and is trained using cross-entropy loss and reinforcement learning, respectively.

[0168] (b) Influence of Using Different Types of Image Features

[0169] The present invention conducted two groups of experiments to evaluate the performance of the model of the present invention in fusing different image features and text features. The first group only utilized the region features extracted by Faster R-CNN as the input, while the second group combined the combination of the full-image features and region features. The model of the present invention adopts a multi-modal information fusion mode and fuses image and text information through a two-level text clustering mechanism. Specifically, one set of experimental configurations consists of 100 low-level clustering centers and 10 high-level clustering centers, while the other set of experimental configurations consists of 100 low-level clustering centers and 30 high-level clustering centers. Their fusion of image features and text clustering features adopts a three-layer fusion mode: two layers fuse high-level clustering centers, and one layer fuses low-level clustering centers. The experimental results are shown in Table 4. When the input image features consist of the fusion of full-image features and region features, the image caption generation model shows significant improvements in multiple evaluation metrics on the dataset. For example, in two comparative experiments, the comprehensive evaluation metric S * reaches the values of 0.4282 and 0.4196, showing a significant performance improvement compared with 0.4149 and 0.4029 obtained when only using region features. These experimental results strongly verify the effectiveness of multi-scale images and the role of feature fusion of reaction scenes and targets in improving the quality of underwater image caption generation.

[0170] Table 4 - Experimental results of different visual features. RF represents only using region features, and FI-RF combines full-image features and region features

[0171]

[0172] (c) Influence of different training modes on the model

[0173] During the training process of the underwater image caption generation model of the present invention, both cross-entropy loss optimization and reinforcement learning to improve the evaluation parameter CIDEr can be used.

[0174] Table 5 presents the results under different training modes. Based on the evaluation metrics, the model based on reinforcement learning always performs better than the model trained using cross-entropy loss in all metrics. These results highlight that combining reinforcement learning strategies can significantly improve the performance of the underwater image caption generation model, especially in achieving a better balance among text fluency, quality, and generation diversity.

[0175] Table 5 - Experimental results of different training procedures

[0176]

[0177] (d) Influence of the number of high-level text clustering centers in the model

[0178] During the fusion process, five different combination results of the two-level clustering centers are as follows Figure 8 shown. This figure illustrates the influence of high-level and low-level text clustering centers with different configurations on the evaluation metrics. The two-level text clustering features are still combined with image features through a three-layer fusion mode. The configuration of the three-layer fusion captures the text feature variations at different levels through the combination of 100 low-level and different high-level clustering centers, enhancing the model's ability to extract and utilize rich semantic information.

[0179] Figure 8 (a) shows the comparison of METEOR and CIDEr evaluation parameters for different numbers of high-order clustering centers. When the number of high-level text clustering centers is 30, the METEOR value is the highest, while when the number of high-level text clustering centers is 10, the CIDEr value reaches its peak. This indicates that in the latter case, the model achieves a higher semantic similarity and sentence quality. Figure 8 (b) presents the comparison of ROUGE_L and S* scores. When the number of high-level text clustering centers is 30, the ROUGE_L score is the highest, but the combination of 10 high-level text clustering centers obtains the highest S* score, which indicates the overall performance of the system and provides the most balanced results. This configuration performs well in both N-gram matching and natural language relevance.

[0180] (e) Influence of the number of fusion layers in the model

[0181] When different numbers of clustering centers or fusion layers are adopted, the results are shown in Table 6. In the previous experiment, when the two-level text clustering features are fused with image features across a three-layer structure of multi-head attention, the 10-10-100 clustering center configuration obtains the highest evaluation metrics. To compare the performance when there are fewer fusion layers, the present invention designs a two-layer structure using a combination of 10-100 clustering features, and another structure using 100 clustering centers to perform a single fusion layer. Different levels of fusion operations enhance features at different depths or strengthen features at specific levels through multiple fusions. From each metric, the three-layer model is superior to the single-layer and double-layer models on multiple evaluation matrices, especially BLEU-4 (0.3084) and CIDEr (0.5989). In terms of the average value S* of the four evaluation parameters, the three-layer fusion in the model of the present invention also shows a significant improvement, indicating that the overall performance of the model is much superior. Generally speaking, increasing the number of layers helps to improve the performance of the model in underwater image caption generation tasks, especially in terms of expression quality and generation ability.

[0182] Table 6 Experimental results of different fusion layers in the model

[0183]

[0184] Example 3:

[0185] This embodiment provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of a method for generating underwater image captions based on multi-modal information fusion, including the following steps:

[0186] Image feature extraction: Obtain an underwater image, and use Faster R-CNN to extract multi-scale image features, including full-image features and regional features;

[0187] Text feature extraction: Use CLIP to perform text word embedding encoding to obtain semantic information associated with the image content, and use K-means clustering for hierarchical text feature extraction to analyze the data distribution and extract text features at different levels;

[0188] Multi-modal information fusion: Construct a multi-modal information fusion method based on the multi-head attention mechanism, effectively fuse image features and text features at different scales, and obtain fused features;

[0189] Caption generation: Based on the fused features, use a preset underwater image caption generator of Transformer to generate an image caption that expresses the image content and is contextually related for the underwater image.

[0190] Example 4:

[0191] This embodiment provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.

[0192] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps in the above embodiment regarding a method for generating underwater image captions based on multimodal information fusion; one or more instructions in the computer-readable storage medium are loaded and executed by a processor to perform the following steps:

[0193] Image feature extraction: Obtain an underwater image and use Faster R-CNN to extract multi-scale image features, including full-image features and regional features;

[0194] Text feature extraction: Use CLIP to perform text word embedding encoding to obtain semantic information associated with the image content, and use K-means clustering for hierarchical text feature extraction to analyze the data distribution and extract text features at different levels;

[0195] Multimodal information fusion: Construct a multimodal information fusion method based on the multi-head attention mechanism to effectively fuse image features and text features at different scales to obtain fused features;

[0196] Caption generation: Based on the fused features, use a preset underwater image caption generator of Transformer to generate an image caption that expresses the image content and is contextually related for the underwater image.

[0197] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0198] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices create means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or in one or more of the blocks.

[0199] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or in one or more of the blocks.

[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or in one or more of the blocks.

[0201] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the above embodiments. Without departing from the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

[0202] Many other changes and modifications can be made without departing from the concept and scope of the present invention. It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.

Claims

1. An underwater image caption generation method based on multimodal information fusion, characterized in that, The method includes: Multi-scale image feature extraction: Obtain an underwater image, and use the faster Region-based Convolutional Neural Network (Faster R-CNN) to extract multi-scale image features, including full-image features and regional features; Hierarchical text feature extraction: Utilize the Contrastive Language-Image Pretraining (CLIP) model based on contrastive learning for text word embedding encoding to obtain semantic information associated with the image content. Adopt K-means clustering for hierarchical text feature extraction, analyze the data distribution, and extract text features at different levels; Multi-modal information fusion: Construct a multi-modal information fusion method based on the multi-head attention mechanism to effectively fuse image features and text features at different scales to obtain fused features; Caption generation: Based on the fused features, use a preset underwater image caption generator of the Transformer to generate an image caption that expresses the image content and is contextually related for the underwater image.

2. The underwater image caption generation method based on multimodal information fusion according to claim 1, wherein, In the image feature extraction stage, use Faster R-CNN for image feature extraction, with the pre-trained Residual Network (ResNet101) as the backbone network; the multi-scale image features expand the global and local information expression of the underwater image from two aspects: full-image features and regional features; Suppose an input three-channel true-color underwater image I is provided. ResNet101 is a typical convolutional neural network containing residual modules. A series of convolutional operations are performed using different types of convolutional kernels, and it outputs a full-image feature map F, representing the scene information of the underwater image; F = ResNet101(I). (1) Faster R-CNN uses the Region Proposal Network (RPN) to generate N proposed target regions, denoted as {r1, r2, …, r N}, which predicts whether each location contains a target and calculates the offset of the bounding box of the target. Region of Interest Pooling (RoIPooling) is used to process regions of interest of different sizes of input, and it outputs N region feature vectors labeled as R, corresponding to the N proposed targets, as follows: R = RoIpooling(r1, r2,..., r N ). (2) Finally, through a fully connected network combined with bounding box regression and the softmax function, predict the bounding boxes and categories of the candidate objects. The multi-scale image information provides support for the word sequence prediction in the underwater image caption generator.

3. A method for generating underwater image captions based on multi-modal information fusion according to claim 1, It is characterized in that In the text feature extraction stage, it specifically includes: Word embedding encoding based on CLIP: Tokenize and count the text annotated in the underwater image caption dataset to construct a corresponding vocabulary. The vocabulary in the vocabulary is encoded by the pre-trained CLIP model for word embedding to generate word vectors related to the image content. The shape of the word vectors is [N×512], where N represents the number of vocabulary in the vocabulary, and 512 is the dimension of the word embedding vector output by the CLIP model; Generate text clustering features: Use the K-means algorithm to cluster the word embedding vectors generated by CLIP. First, perform low-level clustering by clustering the word vectors of all vocabulary to obtain low-level clustering centers; the shape of the clustering result is [N1×512], where N1 represents the number of low-level clustering centers; Generate high-level clustering: Based on the vector set of the low-level clustering centers, perform high-level clustering; further cluster the low-level clustering center vectors to obtain high-level clustering centers. The shape of the clustering result is [N2×512], where N2 represents the number of high-level clustering centers; Text feature representation of hierarchical clustering centers: The cluster centers represent the samples similar to them. Therefore, each cluster center serves as the classification feature of these words, enhancing the expression of text semantics. The lower-level cluster centers represent more detailed semantic categories, while the higher-level cluster centers represent more abstract and broad categories; Information fusion: The text features of the lower-level clustering and the higher-level clustering are gradually fused with the features of the underwater image to provide rich semantic support for the subsequent image caption generation task; Through the above steps, the word embeddings generated by CLIP and the two-level clustering method based on the K-means algorithm effectively extract the hierarchical text features of the underwater image caption, enhancing the multi-modal information association between the text and the image.

4. A method for generating underwater image captions based on multimodal information fusion according to claim 1, characterized in that, The multi-modal information fusion stage specifically includes: Assume that at time step t of the underwater image caption generation model, the full-image feature F t and the region feature R t =[r1, r2, …, r n , which are extracted by object detection, and then they are concatenated to form a combined image representation, denoted as V t , where V t aggregates multi-scale image information and is defined as: V t = [F t , R t . (3) In the multi-head attention module, by using the image information V t and the output of the j-th level text clustering module, representing the extracted text feature vector C j , the attention mechanism is applied to achieve feature fusion, and the attention weight is calculated as: where W q and W k are the query and key-value matrices that implement linear dimensionality transformation in the self-attention mechanism, and the parameter d k is the dimension of (W k C j ), the parameter α t represents the attention weight, which quantifies the correlation between the combined full-image features and the regional features of V t and the text clustering feature C j ; Next, use the following formula to calculate the weighted text clustering features: In the formula represents the weighted text clustering feature, and W v represents the linear transformation matrix; Subsequently, and V t perform a residual connection and then perform layer normalization. The formula is: Among them, the function LayerNorm(·) represents the layer normalization operation, which performs scaling and offset based on the learnable parameters of all samples, and X t is the output of the layer normalization operation. The layer normalization stabilizes the output distribution by normalizing the output values of the activation function within each batch, thereby accelerating the training process and enhancing the generalization ability of the model; X t It is fed into a feedforward neural network, then added to itself, and layer normalization is performed for normalization; Fusion vt = LayerNorm(X t + FFN(X t ))), (7) Among them, Fusion vt is the feature of the fusion of image and text information. The function FFN(·) represents the operation of a two-layer feed-forward network, which is a supplement to the self-attention mechanism to enhance the expressive power of the model. Finally, the output fusion feature Fusion vt is sent to the Transformer-based underwater image caption generator to predict the next word in the sentence sequence; Through two-level clustering, the text feature vectors of the two sets of cluster centers are obtained. During the information fusion process, they are combined sequentially level by level according to the global attention mechanism fusion mode to obtain a vector that integrates more feature information. A hierarchical information fusion strategy is adopted. The higher-level text clustering features are fused with the image features of the previous layer, mainly focusing on identifying the main categories of words. The lower-level text clustering features are fused in the subsequent layers to express more concrete information, and finally the fusion features of the image and the text are obtained.

5. The underwater image caption generation method based on multimodal information fusion according to claim 1, wherein The model training in the caption generation stage specifically includes: Assume that the description y of the underwater image caption generation model for generating image I is: y = (y1, y2,..., y T ) (8) where y i is the i-th word generated, and T is the length of the description; The training process of the underwater image caption generator includes two stages: cross-entropy optimization and the optimization of the evaluation metric CIDEr for the similarity between the image description text and the reference value; In the cross-entropy optimization stage, the cross-entropy loss function is used to quantify the ability of the model to predict the word sequence, and its definition is: where y t is the output of the generator at time t, and y 1:t-1 is the output of the generator before time t. Given the underwater image caption generation model parameters θ, the symbol p θ represents the probability distribution that the output of the generator is y t , and the previous word is y 1:t-1 ; The generator is trained by minimizing the cross-entropy loss to generate sentences similar to the reference sentences. When the cross-entropy optimization reaches a certain level and overfitting or training stagnation occurs, the reinforcement learning technique is further applied to directly optimize the numerical value of the evaluation parameter CIDEr of the generated text; The self-critical sequence training SCST method is adopted to optimize the model based on the expected reward, and the loss function of the reinforcement learning is defined as: Among them, the reward function r(·) is the CIDEr parameter value, which is used to measure the consistency between the description generated by the model and the reference text. By maximizing the mathematical expectation of the CIDEr parameter, the ability of the model to generate high-quality underwater image captions is further improved.

6. An underwater image caption generation system based on multimodal information fusion, characterized in that, The system is applied to the method described in any one of claims 1-5. The system includes: Image feature extraction module: Obtain the underwater image and use Faster R-CNN to extract multi-scale image features, including the full-image features and regional features; Text feature extraction module: Use CLIP to perform text word embedding encoding to obtain semantic information associated with the image content, and adopt K-means clustering for hierarchical text feature extraction, analyze the data distribution, and extract text features at different levels; Multi-modal information fusion module: Construct a multi-modal information fusion method based on the multi-head attention mechanism to effectively fuse the image features at different scales and the hierarchical text features to obtain the fusion features; Subtitle generation module: Based on the fused features, an underwater image subtitle generator using a preset Transformer is used to generate image subtitles that express the image content and are contextually relevant for underwater images.

7. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements an underwater image subtitle generation method according to any one of claims 1 to 5 based on multimodal information fusion.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the program is executed by the processor, it implements an underwater image subtitle generation method according to any one of claims 1 to 5 based on multimodal information fusion.

Citation Information

Patent Citations

  • Image caption generation method of multi-attention fusion network based on multi-granularity reward mechanism

    CN112116685A

  • Modal alignment and multi-scale extraction remote sensing image description generation method, system and device based on remote sensing image-text comparison pre-training features and medium

    CN119131196A

  • Automated content filtering using image retrieval models

    US12105755B1

  • Super-resolution on text-to-image synthesis with gans

    US20240281924A1

  • Open-vocabulary object detection based on frozen vision and language models

    WO2024006340A1