Multi-dimensional Fusion Recognition Method and System for Waterborne Targets

By combining the multi-dimensional fusion recognition method of EfficientDet and Motionformer networks, the shortcomings of existing water target recognition technology in multimodal perception and dynamic trajectory modeling are solved, and high-precision recognition and dynamic behavior understanding of water targets in complex water environments are achieved.

CN119992275BActive Publication Date: 2025-06-24ANHUI GUANGCHENG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510473530.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-06-24
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing water target recognition technology has shortcomings in multimodal perception, space-time feature fusion, dynamic trajectory modeling, identification report generation and model adaptive update, and it is difficult to meet the high-precision identification needs in complex water environments.

Method used

A multi-dimensional fusion recognition method for water targets is proposed, combining the EfficientDet target detection network and the Motionformer network to build a recognition framework that coordinates multimodal perception and timing modeling to achieve accurate detection of water targets, dynamic behavior modeling and identification report generation.

Benefits of technology

It improves the recognition accuracy, behavioral understanding ability and system stability of multiple targets in complex water environments, and has the advantages of strong robustness, stable identification, strong behavioral modeling ability and adaptable optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992275B_ABST
    Figure CN119992275B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-dimensional fusion recognition method and system for water targets, comprising the following steps: S1, collecting perceptual image data in the water surface environment and performing preprocessing; S2, performing multi-modal fusion; S3, inputting the fused image data into the EfficientDet target detection network, extracting feature maps, and outputting the class labels and spatial position information of the water targets through the bounding box regression sub-network and the target classification sub-network; S4, constructing a time series input, inputting it into the Motionformer network, and extracting the motion trajectory features of the water targets; S5, generating a recognition report; S6, online updating the network parameters in an incremental learning manner. The present invention fuses the EfficientDet and Motionformer networks to achieve accurate detection and dynamic recognition of water targets, with high accuracy, strong stability, and adaptive optimization capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly to a multi-dimensional fusion recognition method and system for water targets. Background Art

[0002] With the continuous development of artificial intelligence, remote sensing image analysis, and edge computing technologies, water target detection and recognition based on computer vision have been widely applied in the fields of intelligent shipping, maritime patrol, water area monitoring, ecological protection, etc. By deploying shore-based or shipborne visual perception devices to collect water surface image data and perform target recognition and behavior analysis, it has become an important technical path for realizing intelligent water surface supervision and dynamic early warning. However, due to the complex water target scenarios, diverse imaging modalities, and frequent changes in target states, traditional recognition methods have many deficiencies in terms of robustness, real-time performance, and generalization ability.

[0003] Most of the current mainstream water target recognition technologies rely on single-modal image data for processing, such as target detection methods like YOLO series, Faster R-CNN, RetinaNet based on visible light images. These methods have high detection accuracy under ideal lighting conditions, but their performance drops significantly under night, water mist, strong reflection, or occlusion conditions, making it difficult to meet the requirements of all-weather monitoring. To improve the reliability of recognition, some studies have introduced multi-source data such as infrared images and radar images, but the common practices are mostly simple fusion or independent processing, failing to fully exploit the complementarity between different modalities. In addition, most traditional methods lack the ability to continuously model target behaviors and often perform detection based on only a single frame or short-time frames, unable to recognize the dynamic changes of targets in the time series, which limits the effective recognition of abnormal behaviors and high-risk events.

[0004] In terms of model structure, existing methods generally use two-dimensional convolutional neural networks to extract static image features, making it difficult to model the spatial structure and time evolution relationship simultaneously. Some improved methods attempt to introduce temporal neural networks (such as LSTM, GRU) or 3D convolutional modules to enhance the time modeling ability, but they have problems such as large computational complexity, difficult training, and insufficient target motion modeling. In addition, the lack of a unified end-to-end framework to integrate image fusion, target detection, trajectory modeling, and behavior recognition as a whole results in independent operation of each module, inconsistent feature expressions, and affects the overall recognition performance.

[0005] In terms of fusion, although existing research has explored deep fusion methods for multi-modal images, most of them focus on specific modal pairs (such as infrared and visible light) or simple splicing at the feature layer, failing to achieve deep collaborative expression of cross-modal features. At the same time, existing multi-modal fusion methods generally ignore the impacts brought by image quality differences and spatial registration errors between modalities, easily leading to redundant or distorted fusion information, which in turn affects the subsequent recognition effect.

[0006] In the aspect of object detection, as an efficient object detection framework, EfficientDet has relatively high detection accuracy and computational efficiency, and has been gradually applied to the water surface object detection task. However, in a complex water environment, it is difficult to support higher-level behavior analysis and risk warning only relying on object bounding box coordinates and class labels. Especially in scenarios such as multi-object interaction, occlusion crossing, and drastic changes in moving speed, traditional detection results are difficult to provide complete object state information and lack a deep understanding of object trajectories and behavior patterns.

[0007] To sum up, the existing water surface object recognition technologies have varying degrees of deficiencies in multi-modal perception, spatio-temporal feature fusion, dynamic trajectory modeling, recognition report generation, and model adaptive update. They lack a unified and efficient multi-dimensional fusion recognition framework and are difficult to meet the comprehensive requirements of accurate object recognition and dynamic behavior understanding in a complex water environment. Summary of the Invention

[0008] An object of the present invention is to propose a multi-dimensional fusion recognition method and system for water surface objects. The present invention integrates the EfficientDet object detection network and the Motionformer network to construct a recognition framework that coordinates multi-modal perception and temporal modeling, realizes accurate detection of water surface objects, dynamic behavior modeling, and recognition report generation, has high detection accuracy, stable behavior recognition, strong environmental adaptability, and sustainable optimization ability, and is applicable to intelligent monitoring and safety warning scenarios in complex water environments.

[0009] A multi-dimensional fusion recognition method for water surface objects according to an embodiment of the present invention includes the following steps:

[0010] S1. Collect perceptual image data in the water surface environment and preprocess the perceptual image data;

[0011] S2. Generate fused image data based on the preprocessed image data through multi-modal fusion;

[0012] S3. Input the fused image data into the EfficientDet object detection network. The EfficientDet object detection network uses the EfficientNet network as the backbone network to extract feature maps, inputs the feature maps into a bidirectional feature pyramid network for feature fusion to generate fused feature maps, and outputs the class label and spatial position information of the water surface object through a bounding box regression sub-network and an object classification sub-network.

[0013] S4. Based on the feature map extracted by the EfficientNet network, combined with the class label and spatial location information of the water targets, construct a time series input, input it into the Motionformer network, extract the motion trajectory features of the water targets, and generate a dynamic behavior representation;

[0014] S5. Generate an identification report based on the class label, spatial location information, and dynamic behavior representation;

[0015] S6. Compare the identification report with the historical identification data, and perform online update of the network parameters in an incremental learning manner.

[0016] Optionally, the perceived image data includes radar images, visible light images, and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising processing, and normalization processing.

[0017] Optionally, the specific content of S3 includes:

[0018] S31. Perform a tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, C represents the number of modal channels after splicing. Input the input tensor X into the EfficientDet object detection network, and call EfficientNet as the backbone feature extraction network. The EfficientDet object detection network includes an EfficientNet network, a bidirectional feature pyramid network, an object classification sub-network, and a bounding box regression sub-network. The bidirectional feature pyramid network is used for the fusion of features at different scales;

[0019] S32. In the EfficientNet network, with the tensor X as the input, sequentially perform standard convolution, MBConv module concatenation, and spatial downsampling operations to generate feature maps layer by layer. The MBConv module includes depthwise separable convolution, inverted residual structure, residual connection, and activation function:

[0020] ;

[0021] Among them, represents the output feature map of the th layer, represents the output feature map of the th layer. SE represents the channel attention module, which adjusts the weight of each channel through two operations of compression and excitation to enhance the attention to important features. represents the pointwise convolution operation, represents the depthwise separable convolution operation with a scale of ;

[0022] S33. Perform composite scaling configuration on the EfficientNet network. Let the original EfficientNet network parameters be depth, width, and resolution. Define the scaling factors and perform composite scaling calculations for each parameter:

[0023] ;

[0024] Among them, 、 and represent the depth, width, and resolution after composite scaling respectively. d, w, and r represent the depth, width, and resolution of the original EfficientNet network respectively. 、 and represent the scaling coefficients of depth, width, and resolution. represents the scaling factor;

[0025] S34. In the EfficientNet network after composite scaling, extract feature maps of different scales and output the feature map set P in the order of spatial resolution.

[0026] S35. Input the feature map set into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling, and downsampling operations, and calculate the fused feature map in a learnable weighted fusion manner at each fusion node:

[0027] ;

[0028] Among them, represents the fused feature map of the i-th layer. represents the fusion weight of the feature map. represents the j-th input feature map of the i-th layer. represents a constant to prevent division by zero;

[0029] S36. For the corresponding positions of the anchor boxes in each fused feature map , perform feature vector sampling operations, and sample the feature vectors at specific positions from the feature map according to the center positions of the anchor boxes.

[0030] S37. Input the feature vectors at the corresponding positions of the anchor boxes in the fused feature map into the bounding box regression sub-network. In the bounding box regression sub-network, perform three one-dimensional convolution operations and one activation function operation in sequence to generate the intermediate feature representation of the bounding box regression. Continue to input the intermediate feature representation of the bounding box regression into the second and third layers of convolution and then output the bounding box coordinate offset vector:

[0031] ;

[0032] Among them, It represents the bounding box coordinate offset vector of the j-th anchor box in the i-th layer. It represents the horizontal offset ratio of the center point of the predicted box relative to the anchor box. It represents the vertical offset ratio. It represents the logarithmic scale offset of the width. It represents the logarithmic scale offset of the height.

[0033] S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector:

[0034] ;

[0035] Among them, x and y represent the center point coordinates of the predicted bounding box, and w and h represent the width and height of the bounding box. . . and represent the center coordinates and size parameters of the anchor box.

[0036] Perform non-maximum suppression processing on all prediction results, set the class confidence threshold and the intersection over union threshold, and eliminate redundant detection results based on the score sorting and box overlap situation to screen the final target box position and class prediction results.

[0037] S39. Output the final recognition result of each water target, and the final recognition result includes the predicted class label, the center coordinates of the bounding box, the size information, and the confidence score.

[0038] Optionally, the S4 specifically includes:

[0039] S41. Based on the feature map extracted by the EfficientNet network, combine the class label and spatial position information corresponding to each time step, splice the three types of features and perform unified mapping to construct a time series input vector:

[0040] ;

[0041] Among them, represents the input vector at the j-th time step, and represent the linear projection matrix and bias term of the spliced vector, represents the one-dimensional flattened vector of the feature map at the j-th step, represents the feature map extracted by the EfficientNet network, represents the class label, represents the spatial position information, represents the vector splicing operation;

[0042] S42. Add temporal positional encoding to the time series input vector to obtain an embedding vector:

[0043] ;

[0044] Among them, represents the j-th step temporal embedding vector, represents the input vector at the j-th time step, d represents the dimension of the input vector, k represents the dimension index, sin and cos represent the positional encoding functions, represents the positional encoding vector constructed for all dimensions;

[0045] S43. Input the embedding vector sequence into the multi-head attention module in the Motionformer network, extract the context-dependent features at each time step, and calculate the attention-weighted embedding vector. The Motionformer network includes a temporal positional encoding module, a multi-head attention module, a feed-forward neural network module, and a temporal aggregation module:

[0046] ;

[0047] Among them, represents the attention-weighted embedding vector at the j-th step, Softmax represents the normalization operation, Z represents the temporal embedding vector sequence, , and respectively represent the weight matrices of the query, key, and value matrices of the h-th attention head, T represents the matrix transpose operation, represents the dimension of the key matrix, H represents the number of attention heads, and represent the attention projection matrix and the corresponding bias term;

[0048] S44. Input the attention-weighted embedding vector at each time step into the feed-forward neural network to extract the motion trajectory features corresponding to the time step:

[0049] ;

[0050] Among them, represents the motion trajectory feature at the j-th step, represents the dynamic behavior representation at the j-th step, and represent the feed-forward layer weight matrices, and represent the feed-forward layer bias terms;

[0051] S45. Perform temporal aggregation processing on the motion trajectory features of all time steps, and use average pooling to generate the dynamic behavior representation of the target:

[0052] ;

[0053] Among them, B represents the dynamic behavior representation, representing the motion trajectory feature at the j-th step.

[0054] Optionally, the S5 specifically includes:

[0055] S51. Obtain the output category label sequence , the spatial position information sequence , the dynamic behavior representation sequence ;

[0056] S52. Introduce the category confidence weight factor, the spatial position weight factor, and the behavior representation weight factor, define the confidence weighting for each time step based on the feature change rate, and uniformly construct the feature fusion vector. The weight factors are adaptively generated by the exponential function of the feature change between the previous and the next frames:

[0057] ;

[0058] Among them, represents the feature fusion vector, represents the category confidence weight factor, represents the spatial position weight factor, represents the behavior representation weight factor, represents vector concatenation, represents the category label, represents the spatial position information, represents the dynamic behavior representation;

[0059] S53. Perform average aggregation based on the saliency information gating on the feature fusion vector to obtain the fusion representation vector:

[0060] ;

[0061] Among them, represents the fusion representation vector, represents whether the feature change between adjacent frames is lower than the threshold, used to eliminate abnormal frames and enhance stability, represents the feature fusion vector;

[0062] S54. Split the fusion representation vector into three segments, respectively representing the recognition type of the target, the average spatial position, and the dynamic behavior representation, as the basic data of the recognition report;

[0063] S55. Set the confidence inference function to generate the confidence score:

[0064] ;

[0065] Among them, represents the confidence score, represents the Sigmoid activation function, , and represent the weights, represents the class distribution entropy, represents the spatial location similarity, represents the dynamic behavior intensity;

[0066] S56. Output an identification report, where the identification report includes the category, spatial location information, behavior description representation, and confidence score.

[0067] An underwater target multi-dimensional fusion identification system according to an embodiment of the present invention includes:

[0068] An image acquisition module for acquiring perceptual image data in the water surface environment;

[0069] A preprocessing module for preprocessing the perceptual image data;

[0070] A multi-modal fusion module for performing multi-modal fusion based on the preprocessed image data to generate fused image data;

[0071] A target detection module for inputting the fused image data into an EfficientDet target detection network, extracting multi-scale feature maps, and performing scale fusion, upsampling, and downsampling operations based on a bidirectional feature pyramid network, and outputting the category label and spatial location information of the underwater target;

[0072] A trajectory modeling module for constructing a time series input vector based on the feature maps extracted by the EfficientNet network, combining the category label and spatial location information, inputting it into the Motionformer network, extracting the motion trajectory features of the target, and performing temporal aggregation to generate a dynamic behavior representation;

[0073] An identification report generation module for extracting the final identification type, spatial location, dynamic behavior representation, and confidence score based on the category label, spatial location information, and dynamic behavior representation, and outputting an identification report;

[0074] A model adaptive update module for comparing the identification report with historical identification data and online updating the parameters of the EfficientDet network and the Motionformer network based on an incremental learning method.

[0075] The beneficial effects of the present invention are:

[0076] A multi-dimensional fusion recognition method and system for water targets provided by the present invention constructs an end-to-end linked, multi-modal input, time series modeling and recognition enhancement combined fusion recognition framework to address the problems of single perception modality, rough fusion strategy, weak dynamic behavior modeling, and lack of adaptive update ability in existing water surface target recognition methods, improving the recognition accuracy, behavior understanding ability, and system stability of multi-type targets in complex water environments.

[0077] First, the present invention collects multi-modal perception image data including radar images, visible light images, and infrared images, and performs format conversion, spatial registration, time synchronization, denoising, and normalization processing on them, ensuring the consistency and alignment effect of image data between different modalities, providing high-quality input data for subsequent fusion operations, effectively making up for the problem of incomplete information of single-modal images in specific scenarios, and improving the adaptability of the system to complex water surface environments.

[0078] Secondly, in the feature extraction and detection stage, the present invention inputs the preprocessed fusion image data into the EfficientDet target detection network with EfficientNet as the backbone network, performs two-way fusion processing on multi-scale feature maps through the feature pyramid network, and combines the target classification sub-network and the bounding box regression sub-network for fine detection, improving the recognition accuracy of small targets and occluded targets while enhancing the detection efficiency. At the same time, based on the feature sampling at the anchor box position and the bounding box regression coordinate prediction process, the accurate extraction of the target's spatial position information is realized, ensuring the spatial consistency of subsequent trajectory modeling.

[0079] In addition, the present invention introduces the Motionformer network structure, jointly constructs a time series input with the feature map extracted by EfficientNet, the target category label, and the spatial position information, uses the time position encoding, multi-head time attention mechanism, and feed-forward network module to extract the context-dependent features of the target at consecutive time steps, and generates a robust dynamic behavior representation through average pooling operation, realizing the leap from static detection to dynamic understanding. This method can effectively identify the continuous movement features, behavior pattern changes, and potential abnormal states of the target, providing a high-quality basic representation for behavior recognition and safety warning.

[0080] Finally, in terms of the recognition result output, the present invention not only fuses the category information and spatial positions at multiple time steps, but also introduces a confidence weighting mechanism based on the feature change rate, constructs a feature fusion vector and performs significant information gating aggregation, effectively filtering low-confidence frames, enhancing the stability and credibility of the final recognition output. At the same time, a confidence inference function is set, and category distribution entropy, position similarity and behavior intensity indicators are introduced to generate a quantifiable confidence score, so that the recognition report includes not only the target type and position coordinates, but also behavior descriptions and confidence levels, enhancing the interpretability and practicality of the recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0082] Figure 1 is a flowchart of a multi-dimensional fusion recognition method for water targets proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0084] Refer to Figure 1 , a multi-dimensional fusion recognition method for water targets, includes the following steps:

[0085] S1. Collect the perceptual image data in the water surface environment, and preprocess the perceptual image data;

[0086] S2. Generate fused image data based on the preprocessed image data through multi-modal fusion;

[0087] S3. Input the fused image data into the EfficientDet object detection network. The EfficientDet object detection network uses the EfficientNet network as the backbone network to extract feature maps, input the feature maps into the bidirectional feature pyramid network, perform feature fusion to generate fused feature maps, and output the category labels and spatial position information of water targets through the bounding box regression sub-network and the object classification sub-network;

[0088] S4. Based on the feature maps extracted by the EfficientNet network, construct a time series input in combination with the category labels and spatial position information of water targets, input it into the Motionformer network, extract the motion trajectory features of water targets, and generate a dynamic behavior representation;

[0089] S5, generating a recognition report based on the category label, spatial location information and dynamic behavior representation;

[0090] S6. Compare the recognition report with the historical recognition data and use incremental learning to update the network parameters online.

[0091] The present invention realizes high-precision detection, trajectory modeling and dynamic behavior recognition of water targets in complex environments by constructing a recognition process based on the combination of multimodal image fusion and time series modeling. The method integrates spatial information and temporal information, and opens up a complete chain from data acquisition, image preprocessing, multimodal fusion, target detection to behavior analysis and model updating. It has the advantages of strong robustness, stable recognition, strong behavior modeling ability and adaptive optimization, and is suitable for various application scenarios such as water transportation, port monitoring, and maritime security.

[0092] In this implementation, the perceived image data includes radar images, visible light images, and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising, and normalization.

[0093] The present invention introduces three types of multi-source image data, namely radar images, visible light images and infrared images, and combines preprocessing steps such as format conversion, spatial alignment, time synchronization, denoising and normalization to improve the structural consistency and spatiotemporal alignment accuracy between different modal data, lays a high-quality input foundation for subsequent feature extraction and fusion, and enhances the system's adaptability to complex weather, lighting and water surface disturbances.

[0094] In this implementation, S3 specifically includes:

[0095] S31. Perform tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, and C represents the number of modal channels after splicing. The input tensor X is input to the EfficientDet target detection network, and EfficientNet is called as the backbone feature extraction network. The EfficientDet target detection network includes an EfficientNet network, a bidirectional feature pyramid network, a target classification subnetwork, and a bounding box regression subnetwork. The bidirectional feature pyramid network is used for the fusion of features of different scales.

[0096] S32. In the EfficientNet network, with tensor X as input, standard convolution, MBConv module concatenation and spatial downsampling operations are performed in sequence to generate feature maps layer by layer. The MBConv module includes a depthwise separable convolution, an inverted residual structure, a residual connection and an activation function:

[0097] ;

[0098] Among them, represents the output feature map of the -th layer, represents the output feature map of the -th layer, SE represents the channel attention module, which adjusts the weight of each channel through two operations of compression and excitation to enhance the attention to important features. represents the pointwise convolution operation. represents the depthwise separable convolution operation with a scale of ;

[0099] S33. Perform a compound scaling configuration on the EfficientNet network. Let the original EfficientNet network parameters be depth, width, and resolution, define the scaling factors, and perform compound scaling calculations on each parameter:

[0100] ;

[0101] Among them, , and respectively represent the depth, width, and resolution after compound scaling, d, w, and r respectively represent the depth, width, and resolution of the original EfficientNet network, , and represent the scaling coefficients of depth, width, and resolution, represents the scaling factor;

[0102] S34. In the EfficientNet network after compound scaling, extract feature maps of different scales and output the set P of feature maps in sequence according to the spatial resolution;

[0103] S35. Input the set of feature maps into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling, and downsampling operations, and calculate the fused feature map in a learnable weighted fusion manner at each fusion node:

[0104] ;

[0105] Among them, represents the fused feature map of the -th layer, represents the fusion weight of the feature maps, represents the input feature map of the

[0106] S36. For the corresponding positions of the anchor boxes in each fused feature map , perform a feature vector sampling operation, and sample the feature vectors at specific positions from the feature map according to the center positions of the anchor boxes.

[0107] S37. Input the feature vectors at the corresponding positions of the anchor boxes in the fused feature map into the bounding box regression sub-network. In the bounding box regression sub-network, perform three one-dimensional convolutional operations and one activation function operation in sequence to generate an intermediate feature representation for bounding box regression. Then, continue to input the intermediate feature representation for bounding box regression into the second and third layers of convolution and output the bounding box coordinate offset vector:

[0108] ;

[0109] Among them, represents the bounding box coordinate offset vector of the j-th anchor box in the i-th layer, represents the horizontal offset ratio of the center point of the predicted box relative to the anchor box, represents the vertical offset ratio, represents the logarithmic scale offset of the width, represents the logarithmic scale offset of the height;

[0110] S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector:

[0111] ;

[0112] Among them, x and y represent the center point coordinates of the predicted bounding box, and w and h represent the width and height of the bounding box, 、 、 and represent the center coordinates and size parameters of the anchor box;

[0113] Perform non-maximum suppression processing on all prediction results. Set the class confidence threshold and the intersection over union threshold, and eliminate redundant detection results based on the score sorting and box overlap situation to screen the final target box positions and class prediction results;

[0114] S39. Output the final recognition result of each detected object. The final recognition result includes the predicted class label, the center coordinates of the bounding box, the size information, and the confidence score.

[0115] In the object detection stage of the present invention, the EfficientDet object detection network is introduced, and the efficient feature extraction ability of the EfficientNet backbone network and the multi-scale fusion advantage of the bidirectional feature pyramid network are combined. Through the collaborative design of tensor construction, multi-scale feature generation, bounding box regression, and classification sub-networks, accurate detection of small water targets, long-distance targets, and occluded targets is achieved. At the same time, the flexibility of the detection network at different scene resolutions is improved through the compound scaling mechanism, effectively balancing the detection accuracy and computational efficiency, and meeting the application requirements of real-time detection in complex water environments.

[0116] In this embodiment, S4 specifically includes:

[0117] S41. Based on the feature maps extracted by the EfficientNet network, combining the class labels and spatial position information corresponding to each time step, splicing three types of features and then performing a unified mapping to construct a time series input vector:

[0118] ;

[0119] Wherein, represents the input vector at the j-th time step, and represent the weight matrix and bias term of the linear projection of the splicing vector, represents the one-dimensional flattened vector of the feature map at the j-th step, represents the feature map extracted by the EfficientNet network, represents the class label, represents the spatial position information, represents the vector splicing operation;

[0120] S42. Add time position encoding to the time series input vector to obtain an embedding vector:

[0121] ;

[0122] Wherein, represents the time embedding vector at the j-th step, represents the input vector at the j-th time step, d represents the dimension of the input vector, k represents the dimension index, sin and cos represent the position encoding functions, represents the position encoding vector constructed for all dimensions;

[0123] S43. Input the embedding vector sequence into the multi-head attention module in the Motionformer network, extract the context-dependent features of each time step, and calculate the attention-weighted embedding vector. The Motionformer network includes a time position encoding module, a multi-head attention module, a feed-forward neural network module, and a time series aggregation module:

[0124] ;

[0125] Wherein, represents the attention-weighted embedding vector at the j-th step, Softmax represents the normalization operation, Z represents the time embedding vector sequence, , and respectively represent the weight matrices of the query, key, and value matrices of the h-th attention head, and T represents the matrix transpose operation, represents the dimension of the key matrix, H represents the number of attention heads, and represent the attention projection matrix and the corresponding bias term;

[0126] S44. Input the embedding vectors weighted by attention at each time step into the feed-forward neural network to extract the motion trajectory features corresponding to each time step:

[0127] ;

[0128] where, represents the motion trajectory feature at the j-th step, represents the dynamic behavior representation at the j-th step, and represent the weight matrix of the feed-forward layer, and represent the bias term of the feed-forward layer;

[0129] S45. Perform temporal aggregation processing on the motion trajectory features of all time steps, and use average pooling to generate the dynamic behavior representation of the target:

[0130] ;

[0131] where, B represents the dynamic behavior representation, represents the motion trajectory feature at the j-th step.

[0132] In the present invention, by constructing a trajectory modeling network with the Motionformer network as the core, introducing temporal position encoding and multi-head attention mechanism in the time dimension, the extraction of target cross-frame dynamic features and the modeling of temporal context are realized, effectively representing the motion trajectory and behavior features of the target in consecutive time steps. At the same time, by constructing a time series input vector through unified mapping, combining a feed-forward neural network and a pooling aggregation module to generate a dynamic behavior representation, the recognition ability of the system for target abnormal behaviors, interaction behaviors and complex trajectories is improved, providing strong support for generating accurate and interpretable recognition reports.

[0133] In this embodiment, the S5 specifically includes:

[0134] S51. Obtain the output class label sequence , the spatial position information sequence , and the dynamic behavior representation sequence ;

[0135] S52. Introduce the category credibility weight factor, spatial position weight factor, and behavior representation weight factor. Define the confidence weighting for each time step based on the feature change rate, and uniformly construct a feature fusion vector. The weight factors are adaptively generated through the exponential function of the feature change between two consecutive frames:

[0136] ;

[0137] Among them, represents the feature fusion vector, represents the category credibility weight factor, represents the spatial position weight factor, represents the behavior representation weight factor, represents vector concatenation, represents the category label, represents the spatial position information, represents the dynamic behavior representation;

[0138] S53. Perform average aggregation based on the saliency information gating on the feature fusion vector to obtain a fused representation vector:

[0139] ;

[0140] Among them, represents the fused representation vector, represents whether the feature change between adjacent frames is lower than the threshold, which is used to eliminate abnormal frames and enhance stability, represents the feature fusion vector;

[0141] S54. Split the fused representation vector into three segments, which respectively represent the recognition type of the target, the average spatial position, and the dynamic behavior representation, as the basic data of the recognition report;

[0142] S55. Set a confidence inference function to generate a confidence score:

[0143] ;

[0144] Among them, represents the confidence score, represents the Sigmoid activation function, 、 and represent weights, represents the category distribution entropy, represents the spatial position similarity, represents the dynamic behavior intensity;

[0145] S56. Output the recognition report, which includes the category, spatial position information, behavior description representation, and confidence score.

[0146] In the recognition result generation stage, the present invention introduces a saliency information gating and confidence weighted fusion mechanism to deeply fuse class labels, spatial location information, and behavior representations. By means of dynamic feature consistency judgment and confidence inference functions, the credibility and discriminative ability of the recognition results are improved. The generated recognition report not only contains static target attributes but also integrates dynamic behaviors and confidence scores, possessing stronger expressiveness and decision-making value. At the same time, the fused representation vector can be used to guide the incremental update of the model, providing a basis for subsequent online learning and enabling the system to have the ability of continuous learning and self-evolution.

[0147] A multi-dimensional fusion recognition system for water targets, comprising:

[0148] An image acquisition module for acquiring perceptual image data in the water surface environment;

[0149] A preprocessing module for preprocessing the perceptual image data;

[0150] A multi-modal fusion module for performing multi-modal fusion based on the preprocessed image data to generate fused image data;

[0151] A target detection module for inputting the fused image data into the EfficientDet target detection network, extracting multi-scale feature maps, and performing scale fusion, upsampling, and downsampling operations based on the bidirectional feature pyramid network, and outputting the class label and spatial location information of the water target;

[0152] A trajectory modeling module for constructing a time series input vector based on the feature maps extracted by the EfficientNet network, combining the class label and spatial location information, inputting it into the Motionformer network, extracting the motion trajectory features of the target, and performing temporal aggregation to generate a dynamic behavior representation;

[0153] A recognition report generation module for extracting the final recognition type, spatial location, dynamic behavior representation, and confidence score based on the class label, spatial location information, and dynamic behavior representation, and outputting a recognition report;

[0154] A model adaptive update module for comparing the recognition report with historical recognition data and performing online updates on the parameters of the EfficientDet network and the Motionformer network based on the incremental learning method.

[0155] Embodiment 1:

[0156] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the intelligent port surface monitoring system of a coastal city. The average daily ship traffic in the port exceeds 100 times, and it has typical complex environmental characteristics of water targets such as high-frequency operations and concurrent operation of multiple types of ships. The experimental monitoring area is selected at the intersection of the main channel of the port and the berthing area, which is prone to abnormal behaviors such as congestion of surface targets, illegal docking, and reverse crossing, and is a key water area for supervision.

[0157] The port regulator has deployed a group of fixed shore-based cameras that support visible light, infrared and microwave radar imaging respectively, and dynamically adjust the viewing angle through an optoelectronic gimbal. The perceived image data is collected at 1 frame per second for 72 hours, covering different lighting and meteorological conditions such as daytime, nighttime, early morning, rainy days and foggy days. The image acquisition module in the present invention continuously perceives the image data of the area and hands it over to the preprocessing module to perform image format conversion, modal alignment, time synchronization and noise removal operations, providing high-quality data input for subsequent fusion and detection.

[0158] In the system, the three modal images are sent to the multimodal fusion module after preprocessing, and the fused image is then input into the EfficientDet target detection network. After extracting multi-scale spatial features in the backbone network EfficientNet, scale fusion is completed through the bidirectional feature pyramid network. The target classification subnetwork outputs the target category, and the bounding box regression subnetwork outputs the location information. The detected targets include large container ships, small tugboats, unmanned boats, floating rafts and other types of water targets.

[0159] The detection results enter the Motionformer trajectory modeling network, which extracts motion trajectory features through continuous time series input combined with position encoding and multi-head attention mechanism, and then forms dynamic behavior representation through the time series aggregation module. On this basis, the system integrates category labels, spatial position information and dynamic behavior representation, performs saliency screening and confidence score calculation, and finally generates a structured recognition report, including: target type (such as "tugboat"), position coordinates (such as "x=251, y=122, w=41, h=36"), behavior description (such as "continuous retrograde crossing") and confidence score (such as "0.91").

[0160] In order to evaluate the performance of the system of the present invention, a comparative experiment was conducted with the traditional single-modal YOLOv5 detection system using the same data source. The evaluation indicators included target detection accuracy, recall rate, behavior recognition accuracy, report confidence stability, and performance improvement after model update.

[0161] Table 1. Summary of experimental performance comparison between the invented system and the traditional YOLOv5 system

[0162] Project indicators The present invention YOLOv5 Indicator Explanation Average target detection accuracy (%) 92.5 79.8 Average recognition accuracy in all scenarios Target detection accuracy in foggy weather (%) 89.7 71.6 Detection performance in low-visibility environments Abnormal behavior recognition accuracy (%) 90.6 73.1 Can it identify behaviors such as driving in the wrong direction and abnormal parking? Trajectory continuity score (0~1) 0.87 0.69 Smoothness and coherence of behavior representation in time series Report confidence score fluctuation CV value 0.13 0.27 Reports the stability of the confidence score (smaller is more stable) False alarm rate (%) 2.3 6.9 The proportion of false positives for non-existent targets or behaviors Accuracy improvement after incremental model update (%) +7.9 Not supported Adaptive learning performance of the model after 24 hours Anomaly detection response time reduction (seconds) 1.2 Not supported Improvement in anomaly detection speed after model iteration Average recognition report generation time (ms) 162ms 199ms Time consumption from detection to generating a complete report for a single target Nighttime recognition accuracy (%) 93.2 78.4 Target recognition capability in the absence of natural light

[0163] It can be seen from the experimental results that the multi-dimensional fusion recognition system for water targets proposed in the present invention is superior to the traditional YOLOv5 single-modal detection system in multiple key performance indicators.

[0164] In terms of overall recognition accuracy, the average target detection accuracy of the system of the present invention in all scenarios reaches 92.5%, which is 12.7 percentage points higher than the 79.8% of the YOLOv5 system, demonstrating a stronger recognition capability for multiple types of targets on the water surface after integrating multimodal image data and deep detection networks.

[0165] Under extreme weather conditions such as fog, the system of the present invention still maintained a detection accuracy of 89.7%, while the accuracy of the YOLOv5 system dropped to 71.6%, verifying the robustness of multimodal perception input in weak visual environments. More importantly, in terms of behavior recognition, the system of the present invention uses the Motionformer network to model motion trajectories and achieves an abnormal behavior recognition accuracy of 90.6%, which is higher than the 73.1% of the YOLOv5 system, and is more accurate and sensitive in identifying dynamic behaviors such as retrograde and abnormal parking.

[0166] In terms of system stability, the volatility of the recognition report confidence score (CV value) of the present invention is only 0.13, which is much lower than 0.27 of the comparison system, indicating that in different time frames and different target conditions, the recognition confidence output by the present invention is more stable and controllable. The trajectory continuity consistency score reaches 0.87, which is also higher than 0.69 of the YOLOv5 system, indicating that the generated behavior representation has better coherence and credibility in the time dimension.

[0167] In terms of error recognition, the false alarm rate of the system of the present invention is only 2.3%, which is much lower than 6.9% of YOLOv5, showing the superior performance of the fusion feature and confidence screening mechanism in suppressing interference targets and filtering false positives. More importantly, the incremental learning mechanism introduced in the present invention realizes the adaptive update of the model. After 24 hours of deployment and operation, the model recognition accuracy is improved by an average of 7.9%, and the anomaly detection response time is shortened by 1.2 seconds. This enables the system to continuously adapt to the dynamically changing water surface environment and maintain stable and efficient recognition capabilities.

[0168] In addition, from the perspective of overall system efficiency, the average time consumed by the present invention in generating a complete recognition report is 162 milliseconds, which is faster than the 199 milliseconds of the YOLOv5 system, demonstrating the high efficiency of the end-to-end optimization structure in reasoning performance. Especially under night conditions, the accuracy of the system of the present invention is still as high as 93.2%, which is significantly ahead of the 78.4% of YOLOv5, verifying the actual effect of the fusion of infrared and radar modes in nighttime.

[0169] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered by the protection scope of the present invention.

Claims

1. A multi-dimensional fusion recognition method for water targets, characterized in that: The steps include: S1, collecting perception image data in a water surface environment, and preprocessing the perception image data; S2, performing multimodal fusion based on the preprocessed image data to generate fused image data; S3, input the fused image data into the EfficientDet target detection network, the EfficientDet target detection network uses the EfficientNet network as the backbone network to extract feature maps, input the feature maps into the bidirectional feature pyramid network, perform feature fusion to generate fused feature maps, and output the category labels and spatial position information of the water targets through the bounding box regression subnetwork and the target classification subnetwork; S4. Based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information corresponding to each time step, the three types of features are concatenated and mapped uniformly to construct the time series input vector: ; in, represents the input vector at the jth time step, and represents the linear projection matrix and bias term of the stitching vector, represents the one-dimensional flattened vector of the j-th step feature map, represents the feature map extracted by the EfficientNet network, represents the category label, Represents spatial location information, Represents a vector concatenation operation; Input to the Motionformer network to extract the motion trajectory features of the water target and generate dynamic behavior representation; S5, generating a recognition report based on the category label, spatial location information and dynamic behavior representation; S6. Compare the recognition report with the historical recognition data and use incremental learning to update the network parameters online.

2. A multi-dimensional fusion recognition method for water targets according to claim 1, characterized in that: The perceived image data includes radar images, visible light images and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising and normalization.

3. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S3 specifically includes: S31. Perform tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, and C represents the number of modal channels after splicing. The input tensor X is input to the EfficientDet target detection network, and EfficientNet is called as the backbone feature extraction network. The EfficientDet target detection network includes an EfficientNet network, a bidirectional feature pyramid network, a target classification subnetwork, and a bounding box regression subnetwork. The bidirectional feature pyramid network is used for the fusion of features of different scales. S32. In the EfficientNet network, with tensor X as input, standard convolution, MBConv module concatenation and spatial downsampling operations are performed in sequence to generate feature maps layer by layer. The MBConv module includes a depthwise separable convolution, an inverted residual structure, a residual connection and an activation function: ; in, Indicates The layer outputs a feature map, Indicates The layer outputs the feature map, SE represents the channel attention module, which adjusts the weight of each channel through the two operations of compression and excitation to enhance the attention to important features. represents the point-by-point convolution operation, The scale is Depthwise separable convolution operation; S33. Perform compound scaling configuration on the EfficientNet network. Set the original EfficientNet network parameters as depth, width and resolution, define the scaling factor, and perform compound scaling calculation on each parameter: ; in, , and They represent the depth, width and resolution after compound scaling, d, w and r represent the depth, width and resolution of the original EfficientNet network, respectively. , and represents the scaling factor for depth, width, and resolution, represents the scaling factor; S34, in the EfficientNet network after compound scaling, extract feature maps of different scales, and output the feature map set P in sequence according to the spatial resolution; S35, input the feature map set into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling and downsampling operations, and calculate the fused feature map using a learnable weighted fusion method at each fusion node: ; in, represents the fusion feature map of the i-th layer, represents the fusion weight of the feature map, represents the jth input feature map of the i-th layer, Represents a constant that prevents division by zero; S36. For each fusion feature map At the corresponding position of the anchor frame in the image, a feature vector sampling operation is performed, and a feature vector corresponding to the position of the anchor frame is sampled from the feature map according to the center position of the anchor frame; S37, input the feature vector of the corresponding position of the anchor box in the fusion feature map to the bounding box regression sub-network, perform three one-dimensional convolution operations and one activation function operation in the bounding box regression sub-network in sequence, generate a bounding box regression intermediate feature representation, and continue to input the bounding box regression intermediate feature representation into the second and third convolution layers to output the bounding box coordinate offset vector: ; in, represents the bounding box coordinate offset vector of the jth anchor box at the i-th layer, Indicates the lateral offset ratio of the predicted box relative to the center point of the anchor box. Indicates the longitudinal offset ratio, represents the logarithmic scale shift of the width, represents the logarithmic scale shift of the height; S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector: ; Among them, x and y represent the coordinates of the center point of the predicted bounding box, w and h represent the width and height of the bounding box, , , and Represents the center coordinates and size parameters of the anchor box; Perform non-maximum suppression processing on all prediction results, set the category confidence threshold and intersection-over-union ratio threshold, remove redundant detection results based on score sorting and box overlap, and filter the final target box position and category prediction results; S39. Output the final recognition result of each water target, wherein the final recognition result includes the predicted category label, the center coordinates of the bounding box, the size information and the confidence score.

4. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S4 further comprises: Add the time position encoding to the time series input vector to get the embedding vector: ; in, represents the j-th time embedding vector, represents the input vector of the jth time step, d represents the input vector dimension, k represents the dimension index, sin and cos represent the position encoding function, Represents the position encoding vector constructed for all dimensions; The embedding vector sequence is input into the multi-head attention module in the Motionformer network, the context-dependent features of each time step are extracted, and the attention-weighted embedding vector is calculated. The Motionformer network includes a temporal position encoding module, a multi-head attention module, a feedforward neural network module, and a temporal aggregation module: ; in, represents the weighted embedding vector of the jth step, Softmax represents the normalization operation, Z represents the temporal embedding vector sequence, , and denote the weight matrices of the query, key, and value matrices of the h-th attention head, respectively. T denotes the matrix transpose operation. represents the dimension of the key matrix, H represents the number of attention heads, and Represents the attention projection matrix and the corresponding bias term; The weighted embedding vector of attention at each time step Input the feedforward neural network to extract the motion trajectory features of the corresponding time step: ; in, represents the motion trajectory characteristics of the jth step, represents the dynamic behavior of the j-th step, and represents the feedforward layer weight matrix, and represents the feedforward layer bias term; The motion trajectory features of all time steps are aggregated in time series, and the dynamic behavior representation of the target is generated by average pooling: ; Among them, B represents the dynamic behavior representation, Represents the motion trajectory characteristics of the jth step.

5. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S5 specifically includes: S51. Get the output category label sequence , spatial position information sequence , Dynamic Behavior Representation Sequence ; S52, introduce category credibility weight factor, spatial position weight factor, behavior representation weight factor, define confidence weighting of each time step based on feature change rate, and uniformly construct feature fusion vector. The weight factor is adaptively generated by exponential function of feature change between previous and next two frames: ; in, represents the feature fusion vector, represents the category credibility weight factor, represents the spatial position weight factor, represents the behavior weight factor, represents vector concatenation, represents the category label, Represents spatial location information, Represents dynamic behavior representation; S53, performing average aggregation based on saliency information gating on the feature fusion vector to obtain a fusion representation vector: ; in, represents the fused representation vector, Indicates whether the feature change between adjacent frames is below the threshold, which is used to remove abnormal frames and enhance stability. represents the feature fusion vector; S54, splitting the fusion representation vector into three segments, which respectively represent the recognition type, average spatial position and dynamic behavior representation of the target, as basic data for the recognition report; S55. Set the confidence inference function and generate the confidence score: ; in, represents the confidence score, represents the Sigmoid activation function, , and represents the weight, represents the category distribution entropy, Indicates the spatial location similarity, Indicates the intensity of dynamic behavior; S56. Output a recognition report, wherein the recognition report includes a category, spatial location information, a behavior description representation, and a confidence score.

6. A multi-dimensional fusion recognition system for water targets, which executes the multi-dimensional fusion recognition method for water targets according to any one of claims 1 to 5, characterized in that: include: An image acquisition module, used to collect perception image data in a water surface environment; A preprocessing module, used for preprocessing the perceived image data; A multimodal fusion module, used for performing multimodal fusion based on the preprocessed image data to generate fused image data; The target detection module is used to input the fused image data into the EfficientDet target detection network, extract multi-scale feature maps, and perform scale fusion, upsampling and downsampling operations based on the bidirectional feature pyramid network to output the category label and spatial location information of the water target; The trajectory modeling module is used to construct a time series input vector based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information, and input it into the Motionformer network to extract the motion trajectory features of the target, perform time series aggregation, and generate dynamic behavior representation; The recognition report generation module is used to extract the final recognition type, spatial position, dynamic behavior representation and confidence score based on the category label, spatial position information and dynamic behavior representation, and output the recognition report; The model adaptive update module is used to compare the recognition report with the historical recognition data and update the parameters of the EfficientDet network and the Motionformer network online based on incremental learning.

Citation Information

Patent Citations

  • Automatic ship berthing and leaving method based on video monitoring and image recognition

    CN119200588A