Water surface target detection method and system based on visual language large model

By adopting a Transformer structure-based model and a cross-modal cross-attention fusion network in the water surface object detection algorithm, combining position coding and self-attention mechanism, the problems of large and poor computational overhead and poor robustness of existing algorithms in complex scenarios and multi-scale object detection are solved, and more efficient and accurate surface object detection is achieved.

CN120107690APending Publication Date: 2025-06-06XIDIAN UNIV +1

Patent Information

Application Number
CN202510261639.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing surface object detection algorithms have problems such as large computational overhead and poor robustness when dealing with complex scenarios and multi-scale targets. It is difficult to capture global context information and long-distance dependencies, and have poor separation ability on dense targets.

Method used

A model based on Transformer structure is adopted, combining position coding and self-attention mechanisms, multi-scale features of the image are extracted, and feature fusion and information interaction are performed through dynamic feature aggregation network and cross-modal cross-attention fusion network. Language features are extracted using BERT to enhance the model's understanding and reasoning ability.

Benefits of technology

The detection accuracy of complex scenarios and multi-scale targets is improved, the ability to capture global and local features is enhanced, the separation ability and generalization performance of dense targets is improved, and more accurate surface target detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107690A_ABST
    Figure CN120107690A_ABST
Patent Text Reader

Abstract

The invention discloses a water surface target detection method based on a visual language large model, which mainly solves the problem of inaccurate water surface target detection in the prior art, and the implementation scheme comprises the following steps: constructing a water surface target detection data set; extracting multi-scale features of the image by using the backbone network; constructing a dynamic feature aggregation network to fuse the multi-scale features of the image; extracting language features of the target category text; constructing a cross-modal cross attention fusion network, and fusing the visual dynamic aggregation features and the language features; utilizing a dynamic target detection head to obtain a prediction target category and a target frame; forming a water surface target detection model based on visual language multi-modal fusion by using the backbone network, the dynamic feature aggregation network, the text encoder, the cross-modal cross attention fusion network and the dynamic target detection head, and training the water surface target detection model; and obtaining a water surface target detection result by using the trained detection model. The method can effectively detect the water surface target, is high in accuracy and robustness, and can be used for intelligent ship driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and further designs a surface target detection method and system, which can be used for intelligent ship driving. Background Art

[0002] my country is a country with abundant water resources. It not only has numerous lakes and rivers, but also has a vast ocean area. With the continuous advancement of cutting-edge technologies such as artificial intelligence and computer vision, the continuous replacement and optimization of deep learning models and frameworks, and the object detection algorithm based on deep learning has gradually been applied to the shipping field. The accurate detection of surface targets provides important guarantees for reducing the incidence of ship accidents, improving ship navigation efficiency, and reducing personnel and economic losses. It also plays a vital role in the development of maritime, military, environmental protection and other fields towards intelligence, automation, and informatization. However, the visible light image of the water surface covers a wide range of viewing angles, and contains targets of different sizes, which exacerbates the diversity and complexity of the target scale. Traditional target detection algorithms rely on manually designed features and traditional machine learning models for target recognition. These methods usually have problems such as high computational overhead and poor robustness. They are sensitive to changes in scale, rotation, and deformation, and it is difficult to handle complex scenes and multiple targets. They also lack the automatic feature learning and global context modeling capabilities of deep learning methods.

[0003] The two-stage target detection algorithm and the single-stage target detection algorithm based on deep learning have difficulty in capturing long-distance dependencies and have limitations in capturing global context information. In addition, the two-stage target detection algorithm needs to generate candidate regions, which is computationally expensive and difficult to meet real-time requirements. In addition, convolutional neural networks usually extract features through local convolution windows, which limits their ability to identify occluded or highly similar targets in complex scenes. In addition, the convolutional neural network downsampling layer by layer may cause information loss of small objects. Therefore, the above algorithms have limitations for the case of densely distributed and mutually occluded targets in water surface images. The model cannot capture global and local features well, has poor processing capabilities for multi-scale information and separation capabilities for dense targets, cannot locate and identify small and dense targets more accurately, and has weak generalization capabilities in practical applications.

[0004] The patent document with application number CN202410975253.4 discloses a lightweight real-time integrated detection and recognition method for surface ship targets. It constructs a surface ship image detection model and uses large kernel convolution and pyramid segmentation attention modules to extract image features. However, this method ignores the interaction of multi-level features, resulting in some information being masked or lost, affecting the effective combination of semantic information of low-level features and high-level features. In addition, the convolutional neural network used in this method is difficult to obtain the dependency between features with long spatial distances, which reduces the accuracy of surface target detection.

[0005] In order to solve these problems, researchers proposed a model based on the Transformer structure, which adopts an encoder-decoder network structure, in which the encoder is mainly composed of multiple identical multi-head attention layers, normalization layers, and multi-layer perceptron layers, and the residual structure in the residual neural network is used between the encoders. This model combines position encoding and self-attention mechanisms to expand the receptive field, better capture global context information, and solve the problem that local features in traditional methods cannot capture long-distance context dependencies. It can better handle complex target and background relationships, especially in the detection of multiple targets and overlapping targets. In addition, more and more researchers use visual language large models to handle computer vision tasks and use visual language large models for target detection. By jointly learning visual and language information, rich context information is introduced to enhance the model's understanding and reasoning ability of data, improve the model's adaptability to complex scenes and its ability to identify targets; it can also learn the connection and difference between natural language and image representation, and has higher accuracy and generalization ability in detecting targets in actual application scenarios.

[0006] Patent document CN202410169523.2 discloses a surface target detection method and system based on improved Deformable DETR, which uses a Transformer structure containing deformable convolution. Although this method can better capture global context information and is conducive to processing the relationship between complex targets and backgrounds, it also combines channel attention and spatial attention to enhance feature extraction and improve the detection accuracy of surface targets. However, it is unable to effectively fuse features at different levels, and its ability to process multi-scale information and separate dense targets is poor, and it is unable to more accurately locate and identify small targets and dense targets, resulting in weak generalization ability in practical applications. Summary of the invention

[0007] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a surface target detection method based on a large visual language model, so as to understand more complex semantic relationships through the combination of images and natural language, better capture the global and local information of the image, promote information interaction between different levels, and achieve accurate detection of surface targets.

[0008] The technical solution for achieving the purpose of the present invention includes:

[0009] (1) Using the Transformer-based backbone network to extract multi-scale features of the image 1 ,C 2 ,C 3 ,C 4};

[0010] (2) Construct a dynamic feature aggregation network including a dynamic feature extraction module, a channel attention module, and a context aggregation module, and use the dynamic feature aggregation network to effectively fuse the multi-scale features of the image to obtain a visual dynamic aggregation feature {P 2 ,P 3 ,P 4};

[0011] (3) Use BERT as a text encoder to extract the language features E of the target category text L ;

[0012] (4) Construct a cross-modal cross-attention fusion network including a multi-scale expansion attention module, a flattened splicing layer, and a cross-attention module. Use this cross-modal cross-attention fusion network to fuse visual dynamic aggregation features and language features to obtain a visual fusion embedding vector F containing rich semantic information. V and language fusion embedding vector F L , enhance the ability to understand real scenes and improve the generalization performance of target detection;

[0013] (5) Using the above backbone network, dynamic feature aggregation network, text encoder, cross-modal cross-attention fusion network and dynamic target detection head, a surface target detection model based on visual language multimodal fusion is constructed;

[0014] (6) The gradient descent method is used to train the water surface target detection model. The test set images are input into the trained water surface target detection model, and the target detection bounding box, classification label and corresponding confidence are output to obtain the water surface target detection result.

[0015] Furthermore, the structures of the modules in the dynamic feature aggregation network in step (2) are as follows:

[0016] The dynamic feature extraction module includes an upsampling module, a standard convolution, and a deformable convolution;

[0017] The channel attention module includes a global maximum pooling layer, a global average pooling layer, a standard convolution, and a Sigmoid function;

[0018] The context aggregation module includes three branches including a standard convolution, a Sigmoid function and a Softmax function, wherein the first branch is composed of a standard convolution and a Softmax function connected in series; the second branch is composed of a standard convolution; and the third branch is composed of a standard convolution and a Sigmoid function connected in series;

[0019] The dynamic feature extraction module, channel attention module, and context aggregation module are connected in series in sequence to form a dynamic feature aggregation network.

[0020] Furthermore, the modules that constitute the cross-modal attention fusion network in step (4) are structured as follows:

[0021] The multi-scale dilated attention module includes a projection layer, a sliding window dilated attention head, a linear layer, a multi-layer perceptron, and batch normalization, wherein the projection layer includes three parallel linear layers;

[0022] The flattened and spliced ​​layer comprises three parallel characteristic flattened layers and a spliced ​​layer;

[0023] The criss-cross attention module includes a linear layer and a Softmax function;

[0024] The multi-scale dilated attention module, the flattened splicing layer, and the cross attention module are connected in series in sequence to form a cross-modal cross attention fusion network.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] First, the present invention constructs a dynamic feature aggregation network based on the different distribution and representation capabilities of features at different levels. It can understand the spatial relationship and long-distance dependency between objects while capturing the local edge and texture information of the target, and uses this module to fuse features at different levels, which not only retains the global and local features, but also promotes the network to interact with multi-scale contextual information between features at different levels.

[0027] Second, since the present invention constructs a cross-modal cross-attention fusion network, uses multi-scale dilated attention and cross-attention to enhance the ability to capture multi-scale information, and dynamically associates and effectively aligns image and text information, it can better interactively fuse visual language information, improve adaptability to complex scenes and target recognition capabilities, and thereby improve the accuracy of surface target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flow chart for realizing the method of the present invention;

[0029] Figure 2 It is a block diagram of a water surface target detection model based on visual language multimodal fusion in the present invention;

[0030] Figure 3 It is a block diagram of the dynamic feature aggregation network in the present invention;

[0031] Figure 4 It is a block diagram of the cross-modal cross-attention fusion network in the present invention;

[0032] Figure 5 It is a structural block diagram of the system of the present invention;

[0033] Figure 6 This is a comparison chart of subjective results of the present invention and four existing methods on a water surface target detection dataset. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme and effects in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the present invention, not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by ordinary technicians in this field without creative work should all belong to the protection scope of the present invention.

[0035] It should be noted that the step numbers in the specification and claims of the present invention are only for a clear description of the implementation scheme of the present invention to facilitate understanding, and the order of the step numbers is not limited.

[0036] Reference Figure 1 The implementation steps of the surface target detection method based on the visual language large model in this embodiment are as follows:

[0037] Step 1: Build a water surface target detection dataset.

[0038] Collect water surface videos under different scenes, time and weather conditions, extract the videos frame by frame to obtain images, select the images and annotate them using Labelme software to obtain one-to-one corresponding annotation information files and images;

[0039] The obtained annotation information files and images are combined with the existing WSODD and FloW datasets to form a surface target detection dataset, which is then divided into a training set and a test set in a ratio of 7:3.

[0040] Step 2: read the surface target detection data set to obtain training tensor data and test tensor data.

[0041] Load the images, target texts and label information of the surface target detection training set, and perform random cropping, flipping, scaling and normalization operations on the images to obtain the tensor data corresponding to the training images.

[0042] Load the surface target detection test set images, target text and label information, adjust the image size to the same size as the training image, and then perform normalization to obtain the tensor data corresponding to the test image.

[0043] Step 3: Build a surface target detection model based on visual language multimodal fusion.

[0044] Reference Figure 2 The specific implementation of this step includes the following:

[0045] 3.1) The existing Swin Transformer-ACmix network is selected as the backbone network, and the image tensor data is used as the input of the backbone network. The visual multi-scale features are obtained through four feature extraction levels {C 1 ,C 2 ,C 3 ,C 4};

[0046] 3.2) Construct a dynamic feature aggregation network to obtain visual dynamic aggregation features {P 2 ,P 3 ,P 4}:

[0047] Reference Figure 3 The specific implementation of this step includes the following:

[0048] 3.2.1) Establish a dynamic feature extraction module including upsampling module, standard convolution and deformable convolution;

[0049] 3.2.2) Establish a channel attention module including global maximum pooling layer, global average pooling layer, standard convolution, and Sigmoid function:

[0050] 3.2.3) Establish a context aggregation module including three branches; the first branch is composed of a standard convolution and a Softmax function in series; the second branch is composed of a standard convolution; the third branch is composed of a standard convolution and a Sigmoid function in series. The Sigmoid function and the Softmax function are respectively expressed as follows:

[0051]

[0052] Among them, the Sigmoid function is as follows: <1> As shown, it is used to map the input data into a nonlinear range, so that the network can handle more complex nonlinear relationships; the Softmax function is to convert [z 1 ,z 2 ,…,z n ] as the input vector, and each element of the output vector is obtained as follows: <2> As shown, it is used to normalize each element of the input vector and convert it into an output vector that conforms to the probability distribution. Each element value of the output vector is between 0 and 1, and the sum of all element values ​​is 1.

[0053] 3.2.4) Connecting the dynamic feature extraction module, the channel attention module, and the context aggregation module in series to obtain a dynamic feature aggregation network;

[0054] 3.2.5) The i-th level feature C obtained in step 3.1) i and the adjacent i+1th level feature C i+1 Input dynamic feature extraction module, C i+1 Through the up-sampling module to sample to C i The same resolution, get the upsampled feature C i ' +1 ,; Upsample feature C i ' +1 and low-level features C i After deformable convolution and standard convolution, they are spliced ​​in the channel dimension, and then the dynamic fusion feature X is obtained through standard convolution. i ;

[0055] 3.2.6) Use the channel attention module to obtain the channel attention feature Y i , which will dynamically fuse feature X i After passing through the global maximum pooling layer and the global average pooling layer, they are concatenated in the channel dimension, and then the fusion weight w is obtained through standard convolution and Sigmoid function. i ; According to the low-level feature C i The weight w i And the upsampled features C i ' +1 The weight of 1-w i , C i With C i ' +1 After weighted fusion, the channel attention feature Y is obtained through standard convolution i ;

[0056] 3.2.7) Channel attention feature Y iThrough the three branches of the context aggregation module, the output results of the first branch and the second branch are matrix multiplied, and then a standard convolution is performed on the output result of the third branch to obtain the visual dynamic aggregation feature P i ;

[0057] 3.2.8) Let i = 2, 3, repeat steps 3.2.5) - 3.2.7) twice, and get two visual dynamic aggregation features P 2 ,P 3 ; and C 4 Keep unchanged, that is, get the visual dynamic aggregation feature P 4 , and finally obtain the visual dynamic aggregation feature {P 2 ,P 3 ,P 4};

[0058] 3.3) Use BERT as a text encoder to extract the language features E of the target category text L ;

[0059] 3.3) Construct a cross-modal cross-attention fusion network to obtain the visual fusion embedding vector F V and language fusion embedding vector F L :

[0060] Reference Figure 4 The specific implementation of this step includes the following:

[0061] 3.3.1) Establish a multi-scale dilated attention module including a projection layer, a sliding window dilated attention head, a linear layer, a multi-layer perceptron, and batch normalization, where the projection layer includes three parallel linear layers;

[0062] 3.3.2) establishing a flattened splicing layer including three parallel feature flattening layers and one splicing layer;

[0063] 3.3.3) Establish a cross-attention module including a linear layer and a Softmax function;

[0064] 3.3.4) The multi-scale dilated attention module, the flattened splicing layer, and the cross attention module are sequentially connected in series to form a cross-modal cross attention fusion network;

[0065] 3.3.5) Use the multi-scale dilated attention module to obtain the multi-scale dilated attention feature {D 2 ,D 3 ,D 4}:

[0066] 3.3.5.1) The i-th level visual dynamic aggregation feature P obtained in step 3.2) is iFlattened into a two-dimensional vector form, three sets of query, key and value vectors are obtained through three parallel projection layers, represented as {q1 i ,k1 i ,v1 i},

[0067] {q2 i ,k2 i ,v2 i}, {q3 i ,k3 i ,v3 i};

[0068] 3.3.5.2) The three sets of vectors are expanded through three sliding windows with different expansion rates, and local expansion attention calculation is performed inside the sliding window to obtain the expanded attention vector {M i 1 ,M i 2 ,M i 3};

[0069] 3.3.5.3) Expand the attention vector {M i 1 ,M i 2 ,M i 3} are concatenated in sequence and then passed through a linear layer with the visual dynamic aggregation feature P i Add together to get the multi-scale fusion feature O i , the O i After passing through the multi-layer perceptron and batch normalization, it is combined with O i Add it to itself to get the multi-scale expansion attention feature D i ;

[0070] 3.3.5.4) Let i = 1, 2, 3, repeat steps 3.3.5.1)-3.3.5.3) to obtain the multi-scale dilated attention feature {D 2 ,D 3 ,D 4};

[0071] 3.3.6) Multi-scale dilated attention features {D 2 ,D 3 ,D 4}Through three parallel feature flattening layers, three vectors are obtained in the length and width dimensions, and then the three vectors are concatenated to obtain the multi-scale feature embedding vector E V ;

[0072] 3.3.7) Use the cross attention module to obtain the visual fusion embedding vector FV and language fusion embedding vector F L :

[0073] 3.3.7.1) Embed the multi-scale features into vector E V Through two parallel linear layers, we get two independent matrices V V , Q V , the language feature E L Through two parallel linear layers, we get two other independent matrices K L 、V L ;

[0074] 3.3.7.2) The matrix Q V and the matrix K L Multiply the transpose of , and then use the Softmax function to get the normalized attention distribution matrix C;

[0075] 3.3.7.3) The attention distribution matrix C is respectively combined with the matrix V V and the matrix V L Multiply and pass through two parallel linear layers to obtain two weighted matrices V′ V , V′ L ;

[0076] 3.3.7.4) The weight matrix V′ V and the multi-scale feature embedding vector E V Add together to get the visual fusion embedding vector F V , and the weight matrix V′ L With language feature E L Add together to get the language fusion embedding vector F L ;

[0077] 3.4) Using the dynamic object detection head to perform object classification and bounding box regression simultaneously: embedding the visual fusion into the vector F V and language fusion embedding vector F L The dynamic target detection head is input to obtain the predicted target category and target bounding box, and the prediction results are filtered using the non-maximum suppression algorithm to obtain the final output detection result;

[0078] 3.5) The above-mentioned backbone network and dynamic feature aggregation network are connected in series and then in parallel with the text encoder, and then pass through the cross-modal cross-attention fusion network and the dynamic target detection head in sequence to form a surface target detection model based on visual language multimodal fusion.

[0079] Step 4: Construct the loss function L of the water surface target detection model based on visual language multimodal fusion.

[0080] Since the present invention combines text and visual modalities, its loss function L can be composed of the alignment loss L of visual features and language features: Alignment and the bounding box regression loss L IoU The implementation includes the following:

[0081] 4.1) Calculate the intersection over union (IoU) of the predicted box and the true box to measure the degree of overlap between the predicted box and the true box. The calculation method is as follows:

[0082]

[0083] Among them, Bounding Box represents the predicted box, and Ground Truth represents the real box;

[0084] 4.2) Use the obtained intersection over union (IoU) to calculate the bounding box regression loss L IoU , through L IoU The similarity between the bounding box predicted by the model and the true bounding box can be quantified to measure the performance of the model in detecting the target location. The calculation method is as follows:

[0085] L IoU = 1-IoU;

[0086] 4.3) Calculate the alignment loss L of visual features and language features Alignment , which is used to promote the model to bring the features of similar samples closer and push the features of dissimilar samples farther away, thereby facilitating the effective matching of images and texts. The calculation formula is as follows:

[0087]

[0088] Among them, v i is the image feature, l i is with v i The corresponding language description features, i.e., positive samples, l j is with v i Incompatible language description features, i.e., negative samples; τ is a temperature parameter used to adjust the distribution of similarity scores. If τ is large, the model has a weaker ability to distinguish between positive and negative samples; if τ is small, the model can more strictly distinguish between positive and negative samples, but there may be a risk of overfitting when the samples are more complex or noisy.

[0089] 4.4) Using the alignment loss L of the obtained visual features and language features Alignment and the bounding box regression loss L IoU Calculate the loss function L:

[0090] L=L Alignment +L IoU ;

[0091] Step 5: Train the water surface target detection model based on visual language multimodal fusion.

[0092] 5.1) Set training parameters:

[0093] Set the image size to 1333×800, the batch size to 2, the optimizer to use AdamW, the number of training rounds to 30, and the initial learning rate to 10 -5 , weight decay is 10 -4 ;

[0094] 5.2) Divide all driving image sequences in the training set into different batches according to the batch size;

[0095] 5.3) Input a batch of training set image samples into the surface target detection model based on visual language multimodal fusion to obtain the output results, calculate the loss value based on the output results and the true label, and calculate the gradient of the loss function for each model parameter through back propagation, and then use the AdamW optimizer to update the network parameters according to the gradient calculated by back propagation;

[0096] 5.4) Repeat step 5.3) until the loss function stops decreasing or reaches the maximum number of training rounds, the training stops, and a trained surface target detection model based on visual language multimodal fusion is obtained;

[0097] Step 6: Input the test set images into the trained surface target detection model, output the target detection bounding box, classification label and corresponding confidence, and obtain the surface target detection result.

[0098] Reference Figure 5 ,This example provides a water surface target detection system based on a ,visual language large model, which includes a feature extraction module for ,extracting multi-scale features of an image;

[0099] Dynamic feature aggregation module, used to effectively fuse multi-scale features of images;

[0100] BERT text encoder module, used to extract language features of target category text;

[0101] Cross-modal cross-attention fusion module, used to fuse visual dynamic aggregation features and language features;

[0102] The dynamic target detection head module is used to predict the target detection bounding box, classification label and corresponding confidence to obtain the surface target detection results.

[0103] The effects of the present invention are further described in detail below in conjunction with simulation experiments.

[0104] 1. Simulation Experiment Conditions

[0105] The computer processor used is Intel(R) Core(TM) i5-6600 CPU@3.30GHz, 16GB of memory, and the graphics card is two NVIDIA RTX 3090GPUs with 24GB of video memory.

[0106] The operating system is 64-bit Ubuntu 18.04, the algorithm simulation adopts Python language, and the deep learning framework PyTorch version 2.1.0 is used.

[0107] The evaluation indicators are AP50 and mAP50. AP50 represents the area under the precision-recall curve when the IoU threshold is 0.50, which is used to measure the detection accuracy of a single category; mAP50 represents the average AP of all categories when the IoU threshold is 0.50, which is used to comprehensively evaluate the performance of the model on multiple categories. The calculation formulas are as follows:

[0108]

[0109] Among them, P represents precision, R represents recall, P(R) is the P value corresponding to R on the precision-recall curve; N represents the total number of target detection categories.

[0110] 2. Simulation experiment content and results

[0111] Simulation experiment 1: Under the above simulation experiment conditions, the present invention and the existing six surface target detection methods are used to perform training and testing on a data set containing seven surface targets, respectively, to obtain surface target detection results, and the above evaluation indicators are calculated to evaluate the surface target detection results. The results are shown in Table 1:

[0112] Table 1 AP50 and mAP50 of each category of the present invention and the existing six methods on the water surface target detection dataset

[0113]

[0114] In Table 1, rock, ship, buoy, bridge, boat, rubbish, and harbor are 7 different surface targets;

[0115] The 6 existing methods in Table 1 are:

[0116] Co-DETR: A novel collaborative hybrid assignment training scheme for object detection proposed by Zong et al.

[0117] InternImage-H: A large-scale object detection model based on deformable convolutional neural networks proposed by Wang et al.

[0118] YOLOv8m: The Ultralytics team proposed the eighth version of the YOLO series algorithm, which further improved the detection accuracy, inference speed and versatility of the model.

[0119] GLIP-T: Li et al. proposed a pre-trained model that combines language and visual information for object detection tasks. It uses the relationship between images and text through multimodal alignment to improve the model's understanding ability.

[0120] YOLO-World: A real-time open-world object detection algorithm based on visual language modeling and large-scale datasets proposed by Cheng et al.

[0121] MQ-Det: A multimodal query based object detection algorithm proposed by Xu et al., which incorporates visual queries into existing language query detectors to augment category text with category-level visual information.

[0122] It can be seen from Table 1 that in the real water surface scene, the method of the present invention is superior to the existing model in terms of mAP50 and AP50 indicators of each category, indicating that the method of the present invention is more accurate in detecting water surface targets.

[0123] Simulation experiment 2: Under the above simulation experiment conditions, the present invention and the four existing target detection methods are used to train on the training set collected in step 1, and four images are selected for testing to obtain the subjective results of surface target detection as follows: Figure 6 shown.

[0124] from Figure 6 It can be seen that the method of the present invention is more accurate in detecting small targets and has a better separation effect on dense targets. Compared with other methods, there are fewer missed detections and false detections, and the predicted target frame can cover the complete target and is closer to the size of the real target.

[0125] The above simulation results show that the present invention has higher accuracy and robustness in detecting surface targets in complex scenarios.

Claims

1. A surface target detection method based on a large visual language model, characterized in that: include: (1) Extract multi-scale features {C1, C2, C3, C4} of the image using a Transformer-based backbone network; (2) Construct a dynamic feature aggregation network including a dynamic feature extraction module, a channel attention module, and a context aggregation module, and use the dynamic feature aggregation network to effectively fuse the multi-scale features of the image to obtain visual dynamic aggregation features {P2, P3, P4} containing rich context information; (3) Use BERT as a text encoder to extract the language features E of the target category text L ; (4) Construct a cross-modal cross-attention fusion network including a multi-scale expansion attention module, a flattened splicing layer, and a cross-attention module. Use this cross-modal cross-attention fusion network to fuse visual dynamic aggregation features and language features to obtain a visual fusion embedding vector F containing rich semantic information. V and language fusion embedding vector F L , enhance the ability to understand real scenes and improve the generalization performance of target detection; (5) Using the above backbone network, dynamic feature aggregation network, text encoder, cross-modal cross-attention fusion network and dynamic target detection head, a surface target detection model based on visual language multimodal fusion is constructed; (6) The gradient descent method is used to train the water surface target detection model. The test set images are input into the trained water surface target detection model, and the target detection bounding box, classification label and corresponding confidence are output to obtain the water surface target detection result.

2. The method according to claim 1, characterized in that The multi-scale features of the image are extracted using a Transformer-based backbone network in step (1), including the following: 2a) Perform random cropping, flipping, scaling and normalization on the image to obtain the tensor data corresponding to the image; 2b) The image tensor data is used as the input of the Transformer-based backbone network, and visual multi-scale features {C1, C2, C3, C4} are obtained through four feature extraction levels.

3. The method according to claim 1, characterized in that The structure of each module in the dynamic feature aggregation network in step (2) is as follows: The dynamic feature extraction module includes an upsampling module, a standard convolution, and a deformable convolution; The channel attention module includes a global maximum pooling layer, a global average pooling layer, a standard convolution, and a Sigmoid function; The context aggregation module includes three branches including a standard convolution, a Sigmoid function and a Softmax function, wherein the first branch is composed of a standard convolution and a Softmax function connected in series; the second branch is composed of a standard convolution; and the third branch is composed of a standard convolution and a Sigmoid function connected in series; The dynamic feature extraction module, channel attention module, and context aggregation module are connected in series in sequence to form a dynamic feature aggregation network.

4. The method according to claim 1, characterized in that In step (2), a dynamic feature aggregation network is used to effectively fuse the multi-scale features of the image, and its implementation includes the following: 2a) Use the dynamic feature extraction module to obtain the dynamic fusion feature X i : The high-level features C in the multi-scale features {C1, C2, C3, C4} i+1 Through the upsampling module, the sample is sampled to the adjacent low-level feature C i The same resolution, get the upsampled feature C i ' +1 ; The upsampled feature C i ' +1 and low-level features C i After deformable convolution and standard convolution, they are spliced ​​in the channel dimension, and then the dynamic fusion feature X is obtained through standard convolution. i ; 2b) Use the channel attention module to obtain the channel attention feature Y i : Dynamically fusion feature X i After passing through the global maximum pooling layer and the global average pooling layer, they are concatenated in the channel dimension, and then the fusion weight w is obtained through standard convolution and Sigmoid function. i ; According to the low-level feature C i The weight w i And the upsampled features C i ' +1 The weight of 1-w i , C i With C i ' +1 After weighted fusion, the channel attention feature Y is obtained through standard convolution i ; 2c) Use the context aggregation module to obtain dynamic aggregation features P i : The channel attention feature Y i Through the three branches of the context aggregation module, the output results of the first branch and the second branch are matrix multiplied, and then a standard convolution is performed on the output result of the third branch to obtain the visual dynamic aggregation feature P i ; 2d) Let i = 2, 3, repeat steps 2a) to 2c) twice, and obtain two visual dynamic aggregate features P2 and P3; and C4 remains unchanged, that is, the visual dynamic aggregate feature P4 is obtained, and finally the visual dynamic aggregate feature {P2, P3, P4} containing rich contextual information is obtained.

5. The method according to claim 1, characterized in that The modules that constitute the cross-modal attention fusion network in step (4) are as follows: The multi-scale dilated attention module includes a projection layer, a sliding window dilated attention head, a linear layer, a multi-layer perceptron, and batch normalization, wherein the projection layer includes three parallel linear layers; The flattened and spliced ​​layer comprises three parallel characteristic flattened layers and a spliced ​​layer; The criss-cross attention module includes a linear layer and a Softmax function; The multi-scale dilated attention module, the flattened splicing layer, and the cross attention module are connected in series in sequence to form a cross-modal cross attention fusion network.

6. The method according to claim 1, characterized in that In step (4), the cross-modal cross-attention fusion network is used to fuse the visual dynamic aggregation features and the language features to obtain a visual fusion embedding vector F containing rich semantic information. V and language fusion embedding vector F L , which is implemented as follows: 4a) Use the multi-scale dilated attention module to obtain the multi-scale dilated attention features {D2, D3, D4}: 4a1) Aggregate the visual dynamic features {P2, P3, P4} into P i Flattened into a two-dimensional vector form, three sets of query, key and value vectors are obtained through three parallel projection layers, represented as {q1 i ,k1 i ,v1 i }, {q2 i ,k2 i ,v2 i }, {q3 i ,k3 i ,v3 i }; 4a2) The three sets of vectors are expanded through three sliding windows with different expansion rates, and local expansion attention calculation is performed inside the sliding window to obtain the expanded attention vector {M i 1 ,M i 2 ,M i 3 }; 4a3) Expand the attention vector {M i 1 ,M i 2 ,M i 3 } are concatenated in sequence and then passed through a linear layer with the visual dynamic aggregation feature P i Add together to get the multi-scale fusion feature O i , the O i After passing through the multi-layer perceptron and batch normalization, it is combined with O i Add it to itself to get the multi-scale expansion attention feature D i ; 4a4) Let i = 1, 2, 3, repeat steps 4a1)-4a3) to obtain multi-scale dilated attention features {D2, D3, D4}; 4b) The multi-scale dilated attention features {D2, D3, D4} are expanded in length and width dimensions through three parallel feature flattening layers to obtain three vectors, and then the three vectors are concatenated to obtain the multi-scale feature embedding vector E V ; 4c) Use the cross attention module to obtain the visual fusion embedding vector F V and language fusion embedding vector F L : 4c1) Embed the multi-scale features into vector E V Through two parallel linear layers, we get two independent matrices V V , Q V , the language feature E L Through two parallel linear layers, we get two other independent matrices K L 、V L ; 4c2) The matrix Q V and the matrix K L Multiply the transpose of , and then use the Softmax function to get the normalized attention distribution matrix C; 4c3) The attention distribution matrix C is respectively combined with the matrix V V and the matrix V L Multiply and pass through two parallel linear layers to get two weighted matrices V V '、V L '; 4c4) The weight matrix V V 'With the multi-scale feature embedding vector E V Add together to get the visual fusion embedding vector F V , and the weight matrix V L ' and language feature E L Add together to get the language fusion embedding vector F L .

7. The method according to claim 3, characterized in that The Sigmoid function and Softmax function in the context aggregation module are expressed as follows: Among them, the Sigmoid function is as follows: <1> As shown, it is used to map the input data into a nonlinear range, so that the network can handle more complex nonlinear relationships; The Softmax function converts [z1,z2,…,z n ] as the input vector, and each element of the output vector is obtained as follows: <2> As shown, it is used to normalize each element of the input vector and convert it into an output vector that conforms to the probability distribution. Each element value of the output vector is between 0 and 1, and the sum of all element values ​​is 1.

8. The method according to claim 1, characterized in that In step (6), the gradient descent method is used to train the water surface target detection model, including the following: 6a) Collect water surface videos under different scenes, time and weather conditions, extract the videos frame by frame to obtain images, select images and use Labelme software to annotate them, obtain one-to-one corresponding annotation information files and images; and combine the existing WSODD and FloW datasets to form a water surface target detection dataset, and then divide it into a training set and a test set in a ratio of 7:3; 6b) Based on the characteristics of combining text and visual modalities, the visual feature and language feature alignment loss L is used Alignment and the bounding box regression loss L IoU Construct the loss function L: L=L Alignment +L IoU 6c) Set the batch size to 2, the optimizer to use AdamW, the number of training rounds to 30, and the initial learning rate to 10 -5 , weight decay is 10 -4 ; 6d) Input a batch of training set image samples into the water surface target detection model based on visual language multimodal fusion, use AdamW optimizer to calculate its loss value according to the loss function L, and update the network parameters; 6e) Repeat step 6d) until the loss function stops decreasing or the maximum number of training rounds is reached, and then the training stops.

9. A surface target detection system based on a large visual language model, characterized in that: include: Feature extraction module, used to extract multi-scale features of images; Dynamic feature aggregation module, used to effectively fuse multi-scale features of images; BERT text encoder module, used to extract language features of target category text; Cross-modal cross-attention fusion module, used to fuse visual dynamic aggregation features and language features; The dynamic target detection head module is used to predict the target detection bounding box, classification label and corresponding confidence to obtain the surface target detection results.

Citation Information

Patent Citations

  • Water surface target detection method and system based on improved Deformable DETR

    CN118015255A

  • Lightweight real-time detection and identification integrated method for sea surface ship target

    CN118823326A

Cited By

  • Marine target detection method and device based on multi-modal fusion

    CN120431322A

  • Structured target detection method, device and equipment based on multi-modal language model

    CN120580514A

  • Structured object detection method, device and equipment based on multi-modal language model

    CN120580514B

  • Open target detection method and device, equipment and storage medium

    CN120707834A

  • Flying and hanging object identification method and device of lightweight open-set detection model, and medium

    CN120747744A