An Ocean Ship Graphic Location Method Based on Cross-Modal Interaction
By constructing a cross-modal interaction vision-language encoding and decoding module in the graphic positioning task of marine ships, the problem of difficulty in alignment of visual and language information in graphic positioning of marine ships is solved, and more efficient and accurate object query and positioning is achieved.
Patent Information
- Application Number
- CN202411895993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The prior art has problems in the graphic positioning task of marine ships with difficulty in alignment of visual and language information, unclear object query semantics, and ignoring the rich context of the image itself, resulting in low positioning accuracy.
A method of graphic and text positioning of marine ships based on cross-modal interaction is proposed. By constructing a visual-language encoding module and a visual-language decoding module, the spectrum text participatory interaction module, global visual participatory interaction module, coordinate prior module and discriminant fusion module are used to enhance the semantic understanding and cross-modal interaction of multimodal features.
It effectively improves the efficiency and accuracy of the graphic positioning task of marine ships, and improves the accuracy and positioning accuracy of object query by enhancing the semantic understanding of visual and language features and cross-modal interaction.
Smart Images

Figure CN119360001B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of intelligent ocean and computer vision, and particularly relates to a method for locating ocean ship images and texts based on cross-modal interaction. Background Art
[0002] With the support of computer vision technology, ocean ship image and text location can identify and locate the corresponding objects in ocean ship images according to the given language description. This task not only requires the model to have excellent image parsing ability to understand visual information in the image, such as the shape, color, position of the ship, etc.; but also requires the model to have powerful natural language processing ability to parse semantic information in the language description, such as "a large cargo ship located in the center of the picture" or "a fishing boat with a red flag hanging on the right side of the stern", etc., so as to establish an accurate mapping relationship between vision and language. The input of the image and text location task is an expression and an image, and the output is one or more location boxes, which can accurately mark the positions of the objects in the image corresponding to the language description. This technology has greatly improved the efficiency and accuracy of ocean environment monitoring, and is of great significance for fields such as ocean environment monitoring, ocean resource exploration, and maritime rescue.
[0003] With the rise of Transformer, using multi-modal Transformer to construct feature representations and using cross-attention mechanisms to achieve cross-modal reasoning have achieved excellent performance in the image and text location task. However, this method still has problems such as difficult alignment of visual and language information, unclear object query semantics, and ignoring the rich context of the image itself in ocean ship images, thus reducing the location accuracy of the model in this scenario. Therefore, how to align single-modal visual and language features, enhance multi-modal understanding and cross-modal interaction, and obtain more accurate object queries has now become a key task faced by ocean ship image and text location. Summary of the Invention
[0004] Aiming at the deficiencies in the background art, the purpose of the present invention is to propose a method for locating ocean ship images and texts based on cross-modal interaction, which associates single-modal features with other modal features, improves the semantic understanding of visual and language features by the model, uses coordinate priors to enhance object queries, introduces a discriminative fusion module to strengthen the semantic consistency of multi-modal features, and finally effectively improves the efficiency and accuracy of the ship image and text location task, including the following steps:
[0005] Step S1, collect ocean ship images and text data, and construct an ocean ship image and text location dataset;
[0006] Step S2, construct an ocean ship image and text location model based on cross-modal interaction, and the model structure includes two parts: a visual-language encoding module and a visual-language decoding module;
[0007] Among them, 1) the visual-linguistic encoding module includes a spectral text participatory interaction module and a global visual participatory interaction module, and the specific steps are as follows:
[0008] Step S21, extract the original visual features of the image and the original linguistic features of the text ;
[0009] Step S22, the spectral text participatory interaction module takes the original visual features and the original linguistic features as inputs to generate text-involved visual features ;
[0010] Step S23, the global visual participatory interaction module takes the generated text-involved visual features and the original linguistic features as inputs to generate visually-involved linguistic features ;
[0011] 2) the visual-linguistic decoding module includes a coordinate prior module and a discriminative fusion module, and the specific steps are as follows:
[0012] Step S24, the coordinate prior module takes the original linguistic features as an input to generate a coordinate prior ;
[0013] Step S25, the discriminative fusion module takes the generated text-involved visual features , visually-involved linguistic features and the original linguistic features as inputs to generate an alignment feature ;
[0014] Step S26, use the coordinate prior to enhance the object query , obtain a query , and interact the query with the extracted multi-modal visual features and the linguistic features ;
[0015] Step S3, input the training set in the dataset into the model, calculate the total loss function value, perform backpropagation, and optimize the connection weights through an optimizer and corresponding parameters. After multiple rounds of training, obtain the final image-text localization model.
[0016] Preferably, the visual-linguistic encoding module in Step S21 extracts the original visual features and the original linguistic features The process is as follows: ResNet is used to extract visual features ;
[0017] BERT is used to extract sentence features and word features , and and are concatenated into language features .
[0018] Preferably, the spectrum text participation interaction module in step S22 generates visual features involved in text The process is as follows: The original visual features are dimensionally reduced and reshaped into features , is used as the input of the first layer, and the visual features of the input of the th layer are denoted as , the original language features after dimensional reduction are denoted as , and the original language features after dimensional reduction for each layer are . In each layer of the encoder, the visual features are fused with the original language features after dimensional reduction through the spectrum text participation interaction module, and the feature output of the last layer is represented as visual features involved in text ;
[0019] The specific steps are as follows:
[0020] Step Sa1, use the softmax function to achieve global fine-grained interaction between visual features and language features , and and are multiplied matrix-wise to obtain multimodal visual features . The calculation formula is as follows:
[0021]
[0022]
[0023] where , matrix , , , are learnable parameters with a shape of , and is the softmax function;
[0024] Step Sa2, for the multimodal visual features Perform a two-dimensional discrete Fourier transform twice and set low-pass filtering to obtain visual features , and record the operation processes of the Fourier transform and low-pass filtering as . The calculation formula is as follows:
[0025]
[0026]
[0027] where represents the fast Fourier transform, represents the inverse fast Fourier transform, represents using an adaptive Gaussian smoothing filter to perform low-pass filtering. This filter matches the spatial size of and adapts to the input data with a bandwidth of . represents the convolution operation.
[0028] Preferably, the process of generating the language features of visual participation in the global visual participation interaction module in step S23 is as follows: Use the multi-head cross-modal attention mechanism (MHCA) to capture the image context under the text condition , and then use another multi-head cross-modal attention mechanism to fuse the image context and the language features to obtain the language features of visual participation ;
[0029] The calculation formula is as follows:
[0030]
[0031]
[0032] where represents the original language features after dimensionality reduction , represents the visual features participated by the text after dimensionality reduction , represents the multi-head cross-modal attention mechanism, and the input parameters , and represent the query, key, and value respectively.
[0033] Preferably, the process of generating the coordinate prior in the coordinate prior module in step S24 is as follows: Perform average pooling on the original language features to obtain the features , and use to model the potential spatial position clues to generate the coordinate prior ;
[0034] The calculation formula is as follows:
[0035]
[0036]
[0037]
[0038]
[0039] Among them, represents two fully connected layers with ReLU activation functions, represents the sigmoid function, is a two-dimensional anchor point, represents the number of output channels, is the exponent along the channel dimension, is the temperature.
[0040] Preferably, the discriminative fusion module in step S25 generates the aligned feature The process is as follows: using the visual feature involved in the text and the language feature involved in the vision as the query and key to calculate the correlation between the multimodal visual feature and the language feature, and then applying it to the original language feature to collect the relevant text semantics , multiplying and to obtain the aligned feature ;
[0041] The calculation formula is as follows:
[0042]
[0043]
[0044] Among them, , , are learnable projection parameters.
[0045] Preferably, the vision-language decoding module in step S26 iteratively interacts the object query with the multimodal language and visual features. The specific steps are as follows:
[0046] Step Sb1, at the t-th layer, first add the coordinate prior to the object query to obtain , and then use as the query, As keys and values, they are input into the first multi-head cross-modal attention module for interaction to obtain text semantics , and the calculation formula is as follows:
[0047]
[0048]
[0049] Among them, is the layer normalization operation, is the object query at the t-th layer;
[0050] Step Sb2: Take as the query, take as the keys and values, input them into the second multi-head cross-modal attention module for interaction, and then perform the normalization operation to collect relevant visual semantics ; After enhancing the expression ability of the features using the feed-forward network, perform the normalization operation and output the features , and the calculation formula is as follows:
[0051]
[0052]
[0053] Among them, is the feed-forward network operation.
[0054] Preferably, in the model training process of step S3: Use three fully connected layers and a sigmoid function to directly regress the bounding box coordinates. Let be the object query at the t-th layer projected bounding box, represents the ground truth bounding box, and the predicted bounding box of the last layer of the decoder is used as the final prediction;
[0055] The loss function formula between the predicted bounding box and the ground truth bounding box of the last layer is as follows:
[0056]
[0057] Among them, the smooth L1 loss and the GIoU loss are used to supervise the model training, and the hyperparameters and are used to balance these two losses during the training process, is the number of decoder layers.
[0058] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned method for marine ship graphic and text positioning based on cross-modal interaction.
[0059] Compared with the prior art, the present invention provides a method for marine ship graphic and text positioning based on cross-modal interaction, which has the following advantages and beneficial effects:
[0060] 1. The spectrum text participatory interaction module designed by the present invention enhances visual feature representation by capturing and enriching global visual language relationships. The global visual participatory interaction module designed by the present invention combines relevant image contexts to interpret the input text and generates more descriptive language features. The encoding module links unimodal features with relevant global contexts of other modalities. This process better fuses the visual features extracted by the visual encoder with the language descriptions generated by the text encoder, and realizes effective interaction between modalities through the cross-attention mechanism, improving the model's semantic understanding of visual and language features.
[0061] 2. The present invention uses the input text to generate coordinate priors, enhances the spatial information of the object query representation, suppresses inaccurate cross-modal reasoning, and further strengthens the semantic consistency of multi-modal features through the discriminative fusion module, so as to more comprehensively understand the scene, realize efficient decoding, and finally effectively improve the efficiency and accuracy of the graphic and text positioning task. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is the overall flowchart of a method for marine ship graphic and text positioning based on cross-modal interaction proposed by the present invention;
[0063] Figure 2 is the framework diagram of the marine ship graphic and text positioning model proposed by the present invention;
[0064] Figure 3 is the processing flowchart of the spectrum text participatory interaction module proposed by the present invention;
[0065] Figure 4 is the processing flowchart of the global visual participatory interaction module proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] Next, the technical solutions in the embodiments of the present application will be further clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It should be noted that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0067] To make the invention purpose, technical solution and advantages of this application clearer, the following further elaborates on the embodiments of this application with reference to the accompanying drawings of the specification: To better understand the above-mentioned purpose, features and advantages of the present invention, the following will further illustrate the advantages of the present invention through a comparison of embodiments in combination with the drawings and specific implementation manners.
[0068] In an embodiment, the present invention proposes a method for graphic location of ocean vessels based on cross-modal interaction. The flowchart of this method is as Figure 1 shown below. The steps of this method will be elaborated in detail as follows:
[0069] Step S1: Collect ocean vessel image and text data, and construct an ocean vessel graphic location dataset using the OPT-RSVG, OpenSARShip, and HRSC2016 datasets.
[0070] Step S2: Construct an ocean vessel graphic location model based on cross-modal interaction. The model structure includes two parts: a vision-language encoding module and a vision-language decoding module, as Figure 2 shown;
[0071] Among them, 1) The vision-language encoding module includes a spectral text participatory interaction module and a global vision participatory interaction module. The specific steps are as follows:
[0072] Step S21: Extract the original visual features of the image and the original language features of the text ;
[0073] Furthermore, the process of the vision-language encoding module in step S21 for extracting the original visual features and the original language features is as follows: Set the input image resolution to 640×640, set the length of the input text to 40, use ResNet to extract visual features , use BERT to extract sentence features and word features , and connect and to form language features ;
[0074] Step S22: The spectral text participatory interaction module takes the original visual features and the original language features as inputs to generate text-participated visual features , as Figure 3 shown;
[0075] Furthermore, the spectral text participatory interaction module in step S22 generates text-participated visual features The process is as follows: For the original visual features perform dimensionality reduction and reshape them into features , take as the input of the first layer, and denote the visual features of the input of the th layer as , denote the dimensionality-reduced original language features as , and the dimensionality-reduced original language features of each layer are . In each layer of the encoder, the visual features are fused with the dimensionality-reduced original language features through the spectral text participatory interaction module, and represent the feature output of the last layer as the text-participated visual features ;
[0076] The specific steps are as follows:
[0077] Step Sa1, use the softmax function to achieve the global fine-grained interaction of the visual features and the language features , perform matrix multiplication on and to obtain the multi-modal visual features . The calculation formula is as follows:
[0078]
[0079]
[0080] where , is the visual feature of the th layer, matrices , , , are learnable parameters with a shape of , is the softmax function, is the th layer's multi-modal visual feature;
[0081] Step Sa2, perform a two-dimensional discrete Fourier transform (FFT) on the multi-modal visual features . The first discrete Fourier transform (FFT) gives the frequency-domain representation . Use the adaptive Gaussian smoothing filter to perform a convolution operation on the frequency-domain representation to obtain the convolution result. Perform the second inverse Fourier transform (FFT) on the convolution result to obtain the time-domain representation, and set a low-pass filter for the time domain to obtain the visual features. Denote the operation process of the Fourier transform and the low-pass filter as , the calculation formula is as follows:
[0082]
[0083]
[0084] Among them, represents the fast Fourier transform, represents the inverse fast Fourier transform, represents using an adaptive Gaussian smoothing filter to perform low-pass filtering. This filter matches the spatial size and adapts to the input data with ; represents the convolution operation.
[0085] Step S23, the global visual participatory interaction module will generate visual features involving text and the original language features as inputs to generate language features involving vision , as Figure 4 shown;
[0086] Furthermore, the process of the global visual participatory interaction module in step S23 to generate language features involving vision is as follows: Use the multi-head cross-modal attention mechanism (MHCA) to capture the image context under text conditions , then use another multi-head cross-modal attention mechanism to fuse the image context and language features to obtain the language features involving vision ;
[0087] The calculation formula is as follows:
[0088]
[0089]
[0090] Among them, represents the original language features after dimensionality reduction , represents the visual features involving text after dimensionality reduction , represents the multi-head cross-modal attention mechanism, and the input parameters , and represent the query, key, and value respectively.
[0091] 2) The vision-language decoding module includes a coordinate prior module and a discriminant fusion module. The specific steps are as follows:
[0092] Step S24, the coordinate prior module takes the original language feature as input to generate a coordinate prior ;
[0093] Further, the process of the coordinate prior module in step S24 generating the coordinate prior is as follows: perform average pooling on the original language feature to obtain a feature , and through a fully connected layer and an activation function, convert the feature into an intermediate representation , utilize to model potential spatial location cues, and generate the coordinate prior through a position encoding function. Then, calculate the values at odd positions in the coordinate prior through a sine function, and calculate the values at even positions in the coordinate prior through a cosine function;
[0094] The calculation formula is as follows:
[0095]
[0096]
[0097]
[0098]
[0099] Among them, represents two fully connected layers with ReLU activation functions, represents the sigmoid function, is a two-dimensional anchor point, represents the number of output channels, is the exponent along the channel dimension, is the temperature.
[0100] Step S25, the discriminative fusion module takes the visual features participated by the generated text , the language features participated by vision and the original language feature as input to generate alignment features ;
[0101] Further, the process of the discriminative fusion module in step S25 generating the alignment features is as follows: take the visual features participated by the text and the language features participated by vision as queries and keys to calculate the correlation between multimodal visual features and language features, and then apply it to the original language feature to collect relevant text semantics , for and perform a multiplication operation to obtain alignment features ;
[0102] The calculation formula is as follows:
[0103]
[0104]
[0105] wherein, is the multi-head attention mechanism, , , is the learnable projection parameter.
[0106] Step S26, use the coordinate prior to enhance the object query , obtain the query , and interact the query with the extracted multi-modal visual features and the language features ;
[0107] Furthermore, in the visual-linguistic decoding module in step S26, the object query iteratively interacts with the multi-modal language and visual features to obtain a more accurate representation of the referenced object. The specific steps are as follows:
[0108] Step Sb1, at the t-th layer, first add the coordinate prior to the object query to obtain , then use as the query, as the key and value, input them into the first multi-head cross-modal attention module for interaction, collect relevant text semantics, add the output of the attention module to and then perform normalization to obtain the final text semantics . The calculation formula is as follows:
[0109]
[0110]
[0111] wherein, is the layer normalization operation, is the object query at the t-th layer;
[0112] Step Sb2, use as the query, and use Input as keys and values into the second multi-head cross-modal attention module for interaction, and then perform a normalization operation to collect relevant visual semantics After enhancing the expression ability of features using a feed-forward network, perform a normalization operation and output the features , and the calculation formula is as follows:
[0113]
[0114]
[0115] Among them, is the visual feature representing the current time step t, is the cross-modal attention weight, is the feed-forward network operation.
[0116] Step S3: Input the training set in the dataset into the model, calculate the total loss function value, perform backpropagation, optimize the connection weights through the optimizer and corresponding parameters, and obtain the final image-text localization model after multiple rounds of training.
[0117] Specifically, the model training process in step S3: Use three fully connected layers and a sigmoid function to directly regress the bounding box coordinates. Let be the object query at the t-th layer The projected bounding box, represents the true annotation box, and the predicted bounding box of the last layer of the decoder is used as the final prediction;
[0118]
[0119] The loss function formula between the predicted box and the true annotation box of the last layer is as follows:
[0120] Among them, the smooth L1 loss and the GIoU loss are used to supervise the model training, and the hyperparameters and are used to balance these two losses during the training process, is the number of decoder layers.
[0121] In addition, an embodiment of the present invention also proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the above-mentioned image-text localization method for ocean ships based on cross-modal interaction.
[0122] The present invention collects ocean ship image and text data, constructs an ocean ship image-text localization dataset; constructs an ocean ship image-text localization model based on cross-modal interaction following an encoder-decoder structure; extracts visual features And language features , generate visual features involved in the generated text And language features involved in vision ; fuse the above features to generate coordinate priors And alignment features , utilize To enhance object queries , and interact with multi-modal visual and language features And ; input the training set in the dataset into the model, calculate the loss function value of the entire model, perform backpropagation, optimize the connection weights through the selected optimizer and corresponding parameters, and obtain the final model after multiple rounds of training. The present invention relates single-modal features to other modal features, improving the model's semantic understanding of visual and language features; utilizes coordinate priors to enhance object queries, introduces a discriminative fusion module to strengthen the semantic consistency of multi-modal features, and ultimately effectively improves the efficiency and accuracy of the marine ship graphic and text localization task.
[0123] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0124] Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for positioning marine vessels based on cross-modal interaction, characterized in that: include: Step S1, collecting ocean ship images and text data to construct an ocean ship image and text positioning dataset; Step S2, constructing a marine ship image and text positioning model based on cross-modal interaction, the model structure includes two parts: a visual-language encoding module and a visual-language decoding module; Among them, 1) the visual-language encoding module includes a spectrum text participatory interaction module and a global visual participatory interaction module, and the specific steps are as follows: Step S21, extracting the original visual features of the image and the original language features of the text ; Step S22: the spectrum text participatory interaction module converts the original visual features and original language features Visual features used as input to generate textual participation ; Step S23, the global visual participatory interaction module generates visual features of text participation and original language features Generate visually engaged linguistic features as input ; The specific steps are as follows: Capturing image context conditioned on text using multi-head cross-modal attention mechanism Then, another multi-head cross-modal attention mechanism is used to fuse the image context and language features to obtain the visually involved language features ; The calculation formula is as follows: ; ; in, Represents the original language features after dimensionality reduction , Visual features representing textual involvement after dimensionality reduction , Represents a multi-head cross-modal attention mechanism, with input parameters , and Represents query, key, and value respectively; 2) The visual-language decoding module includes a coordinate prior module and a discriminant fusion module, and the specific steps are as follows: Step S24: the coordinate prior module converts the original language features Generate coordinate priors as input ; Step S25, the discrimination fusion module generates visual features of text participation , Language Features of Visual Participation and original language features Generate alignment features for input ; Step S26, in the tth layer of the decoder, use the coordinate prior Enhanced object query , is the t-th layer object query, and the query is , query and extracted multimodal visual features and language features Interact; Step S3, input the training set in the data set into the model, calculate the total loss function value, perform back propagation, optimize the connection weights through the optimizer and corresponding parameters, and obtain the final image and text positioning model after multiple rounds of training.
2. According to claim 1, a method for positioning marine vessels based on cross-modal interaction is characterized in that: The visual-language encoding module in step S21 extracts the original visual features and original language features The process is as follows: ResNet is used to extract visual features , using BERT to extract sentence features and word features ,Will and Connecting into language features .
3. According to the method for positioning marine vessels based on cross-modal interaction, the method is characterized in that: The spectrum text participatory interaction module in step S22 generates visual features of text participation The process is as follows: the original visual features Perform dimensionality reduction and reshape into features ,Will As the first layer input, The visual features of the layer input are recorded as , the original language features after dimension reduction Recorded as , No. The input of the layer is and In each layer of the encoder, the visual features are combined with the original language features after dimensionality reduction through the spectrum text participatory interaction module Fusion, the feature output of the last layer is represented as the visual feature of the text ; The specific steps are as follows: Step Sa1, use the softmax function to realize visual features and language features Global fine-grained interaction , and Perform matrix multiplication to obtain multimodal visual features , the calculation formula is as follows: ; ; in, ,matrix , , , The shape is The learnable parameters of is the softmax function; Step Sa2, multimodal visual features Perform a secondary two-dimensional discrete Fourier transform and set a low-pass filter to obtain visual features , the operation process of Fourier transform and low-pass filtering is recorded as , the calculation formula is as follows: ; ; in, represents the fast Fourier transform, represents the inverse fast Fourier transform, Indicates the use of adaptive Gaussian smoothing filter Perform low-pass filtering, Represents a convolution operation.
4. The method for positioning marine vessels based on cross-modal interaction according to claim 1, characterized in that: The coordinate prior module in step S24 generates a coordinate prior The process is as follows: the original language features Perform average pooling to obtain features ,use Modeling latent spatial location cues and generating coordinate priors ; The calculation formula is as follows: ; ; ; ; in, represents two fully connected layers with ReLU activation function, represents the sigmoid function, is a two-dimensional anchor point, Indicates the number of output channels, is the index along the channel dimension, It's the temperature.
5. The method for positioning marine vessels based on cross-modal interaction according to claim 1, characterized in that: The discriminant fusion module in step S25 generates an alignment feature The process is as follows: Visual features involving text and visual involvement of language features As query and key to calculate the correlation between multimodal visual features and language features, which is then applied to the original language features To collect relevant text semantics ,right and Perform multiplication to obtain the alignment feature ; The calculation formula is as follows: ; ; in, , , are the learnable projection parameters.
6. The method for positioning marine vessels based on cross-modal interaction according to claim 1, characterized in that: In step S26, the visual-language decoding module decodes the object query Iteratively interact with multimodal language and visual features. The specific steps are as follows: Step Sb1, in the tth layer of the decoder, firstly transform the coordinate prior Add to Object Query In, get , and then As a query, As keys and values, they are input into the first multi-head cross-modal attention module for interaction to obtain text semantics , the calculation formula is as follows: ; ; in, is the layer normalization operation, is the t-th layer object query; Step Sb2, As a query, As keys and values, they are fed into the second multi-head cross-modal attention module for interaction to collect relevant visual semantics , using the feedforward network to enhance the expressiveness of features and output features , the calculation formula is as follows: ; ; in, is a feed-forward network operation.
7. The method for positioning marine vessels based on cross-modal interaction according to claim 1, characterized in that: The model training process in step S3 is as follows: three fully connected layers and a sigmoid function are used to directly regress the bounding box coordinates. Query for the tth layer object The projected bounding box, represents the true annotation box, and the predicted bounding box of the last layer of the decoder is used as the final prediction; Prediction box and final The loss function formula between the true annotation boxes of the layers is as follows: ; Among them, smooth L1 loss and GIoU loss For supervised model training, hyperparameters and Used to balance these two losses during training, is the number of decoder layers.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cross-modal retrieval method and system based on global and local semantic comparative learning
CN117150069A
Cross-modality processing method and apparatus, and computer storage medium
US20210303921A1