Multi-dimensional fusion recognition method and system for water target

By combining the multi-dimensional fusion recognition method of EfficientDet and Motionformer networks, the shortcomings of the existing technology in multimodal perception and dynamic trajectory modeling are solved, and high-precision identification and dynamic behavior understanding of water targets are achieved, which is suitable for intelligent monitoring and safety warning in complex water environments.

CN119992275AActive Publication Date: 2025-05-13ANHUI GUANGCHENG TECH CO LTD

Patent Information

Application Number
CN202510473530.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing water target recognition technology has shortcomings in multimodal perception, space-time feature fusion, dynamic trajectory modeling, identification report generation and model adaptive update, and it is difficult to meet the needs of high-precision identification and dynamic behavior understanding in complex water environments.

Method used

A multi-dimensional fusion recognition method for water targets is proposed. By integrating the EfficientDet target detection network and the Motionformer network, a recognition framework that coordinates multimodal perception and timing modeling is built to achieve accurate detection of water targets, dynamic behavior modeling and identification report generation.

Benefits of technology

It improves the identification accuracy, behavioral understanding ability and system stability of multiple types of targets in complex water environments, and has the ability to have high detection accuracy, stable behavioral recognition, strong environmental adaptability and sustainable optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992275A_ABST
    Figure CN119992275A_ABST
Patent Text Reader

Abstract

The invention discloses an overwater target multi-dimensional fusion identification method and system, and the method comprises the following steps: S1, collecting perception image data in a water surface environment, and carrying out the preprocessing of the perception image data; s2, carrying out multi-modal fusion; s3, inputting the fused image data into an OfficientDet target detection network, extracting a feature map, and outputting a category label and spatial position information of a water target through a bounding box regression sub-network and a target classification sub-network; s4, constructing time sequence input, inputting the time sequence input to a Motionform network, and extracting motion trail features of the water target; s5, generating an identification report; and S6, performing online updating on the network parameters by adopting an incremental learning mode. According to the method, the OfficitDet network and the Motionform network are fused, accurate detection and dynamic recognition of the water target are achieved, and the method has the advantages of being high in precision, high in stability and capable of achieving self-adaptive optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target recognition technology, and in particular to a method and system for multi-dimensional fusion recognition of water targets. Background Art

[0002] With the continuous development of artificial intelligence, remote sensing image analysis and edge computing technology, water target detection and recognition based on computer vision has been widely used in intelligent shipping, maritime patrol, water monitoring, ecological protection and other fields. By deploying shore-based or ship-borne visual perception equipment, collecting water surface image data and performing target recognition and behavior analysis, it has become an important technical path to achieve intelligent water surface supervision and dynamic early warning. However, due to the complex water target scenes, diverse imaging modalities, and frequent changes in target states, traditional recognition methods have many shortcomings in robustness, real-time and generalization capabilities.

[0003] Most of the current mainstream water target recognition technologies rely on single-modality image data for processing, such as the YOLO series, Faster R-CNN, RetinaNet and other target detection methods based on visible light images. These methods have high detection accuracy under ideal lighting conditions, but their performance is significantly reduced at night, in fog, strong reflections or occlusion conditions, making it difficult to meet the needs of all-weather monitoring. In order to improve the reliability of recognition, some studies have introduced multi-source data such as infrared images and radar images, but the common practice is simple fusion or independent processing, which fails to fully explore the complementarity between different modalities. In addition, most traditional methods lack the ability to continuously model target behavior, and often only detect based on single frames or short time frames, and cannot identify the dynamic changes of targets in time series, which limits the effective identification of abnormal behaviors and high-risk events.

[0004] In terms of model structure, existing methods generally use two-dimensional convolutional neural networks to extract static image features, which makes it difficult to simultaneously model the relationship between spatial structure and temporal evolution. Some improved methods attempt to introduce temporal neural networks (such as LSTM, GRU) or 3D convolution modules to enhance temporal modeling capabilities, but there are problems such as large computational complexity, difficulty in training, and insufficient target motion modeling. In addition, there is a lack of a unified end-to-end framework to integrate image fusion, target detection, trajectory modeling, and behavior recognition, resulting in independent operation of each module and inconsistent feature expression, which affects the overall recognition performance.

[0005] In terms of fusion, although some studies have explored deep fusion methods of multimodal images, most of them focus on specific modal pairs (such as infrared and visible light) or simply splice them at the feature layer, and fail to achieve deep collaborative expression of cross-modal features. At the same time, existing multimodal fusion methods generally ignore the impact of image quality differences and spatial registration errors between modalities, which can easily lead to redundant or distorted fusion information, thereby affecting subsequent recognition effects.

[0006] In terms of target detection, EfficientDet, as an efficient target detection framework, has excellent detection accuracy and computational efficiency and has gradually been applied to surface target detection tasks. However, in complex water environments, it is difficult to support higher-level behavior analysis and risk warnings relying solely on target frame coordinates and category labels. Especially in scenarios such as multi-target interaction, occlusion crossing, and drastic changes in moving speed, traditional detection results are difficult to provide complete target status information and lack a deep understanding of target trajectories and behavior patterns.

[0007] In summary, the existing water target recognition technology has deficiencies to varying degrees in multimodal perception, space-time feature fusion, dynamic trajectory modeling, recognition report generation and model adaptive updating. It lacks a unified and efficient multi-dimensional fusion recognition framework and is difficult to meet the comprehensive needs of accurate target recognition and dynamic behavior understanding in complex water environments. Summary of the invention

[0008] One purpose of the present invention is to propose a multi-dimensional fusion recognition method and system for water targets. The present invention integrates the EfficientDet target detection network and the Motionformer network to construct a recognition framework that collaborates with multimodal perception and time series modeling, thereby realizing accurate detection of water targets, dynamic behavior modeling, and generation of recognition reports. The method has high detection accuracy, stable behavior recognition, strong environmental adaptability, and sustainable optimization capabilities, and is suitable for intelligent monitoring and safety warning scenarios in complex water environments.

[0009] A multi-dimensional fusion recognition method for water targets according to an embodiment of the present invention comprises the following steps: S1, collecting perception image data in a water surface environment, and preprocessing the perception image data; S2, performing multimodal fusion based on the preprocessed image data to generate fused image data; S3, input the fused image data into the EfficientDet target detection network, the EfficientDet target detection network uses the EfficientNet network as the backbone network to extract feature maps, input the feature maps into the bidirectional feature pyramid network, perform feature fusion to generate fused feature maps, and output the category labels and spatial position information of the water targets through the bounding box regression subnetwork and the target classification subnetwork; S4, based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information of the water target to construct a time series input, input it into the Motionformer network, extract the motion trajectory features of the water target, and generate a dynamic behavior representation; S5, generating a recognition report based on the category label, spatial location information and dynamic behavior representation; S6. Compare the recognition report with the historical recognition data and use incremental learning to update the network parameters online.

[0010] Optionally, the perceived image data includes radar images, visible light images and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising and normalization.

[0011] Optionally, the S3 specifically includes: S31. Perform tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, and C represents the number of modal channels after splicing. The input tensor X is input to the EfficientDet target detection network, and EfficientNet is called as the backbone feature extraction network. The EfficientDet target detection network includes an EfficientNet network, a bidirectional feature pyramid network, a target classification subnetwork, and a bounding box regression subnetwork. The bidirectional feature pyramid network is used for the fusion of features of different scales. S32. In the EfficientNet network, with tensor X as input, standard convolution, MBConv module concatenation and spatial downsampling operations are performed in sequence to generate feature maps layer by layer. The MBConv module includes a depthwise separable convolution, an inverted residual structure, a residual connection and an activation function: ; in, Indicates The layer outputs a feature map, Indicates The layer outputs the feature map, SE represents the channel attention module, which adjusts the weight of each channel through the two operations of compression and excitation to enhance the attention to important features. represents the point-by-point convolution operation, The scale is Depthwise separable convolution operation; S33. Perform compound scaling configuration on the EfficientNet network. Set the original EfficientNet network parameters as depth, width and resolution, define the scaling factor, and perform compound scaling calculation on each parameter: ; in, , and They represent the depth, width and resolution after compound scaling, d, w and r represent the depth, width and resolution of the original EfficientNet network, respectively. , and represents the scaling factor for depth, width, and resolution, represents the scaling factor; S34, in the EfficientNet network after compound scaling, extract feature maps of different scales, and output the feature map set P in sequence according to the spatial resolution; S35, input the feature map set into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling and downsampling operations, and calculate the fused feature map using a learnable weighted fusion method at each fusion node: ; in, represents the fusion feature map of the i-th layer, represents the fusion weight of the feature map, represents the jth input feature map of the i-th layer, Represents a constant that prevents division by zero; S36. For each fusion feature map The feature vector sampling operation is performed at the corresponding position of the anchor frame in the feature map, and the feature vector at a specific position is sampled from the feature map according to the center position of the anchor frame; S37, input the feature vector of the corresponding position of the anchor box in the fusion feature map to the bounding box regression sub-network, perform three one-dimensional convolution operations and one activation function operation in the bounding box regression sub-network in sequence, generate a bounding box regression intermediate feature representation, and continue to input the bounding box regression intermediate feature representation into the second and third convolution layers to output the bounding box coordinate offset vector: ; in, represents the bounding box coordinate offset vector of the jth anchor box at the i-th layer, Indicates the lateral offset ratio of the predicted box relative to the center point of the anchor box. Indicates the longitudinal offset ratio, represents the logarithmic scale shift of the width, represents the logarithmic scale shift of the height; S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector: ; Among them, x and y represent the coordinates of the center point of the predicted bounding box, w and h represent the width and height of the bounding box, , , and Represents the center coordinates and size parameters of the anchor box; Perform non-maximum suppression processing on all prediction results, set the category confidence threshold and intersection-over-union ratio threshold, remove redundant detection results based on score sorting and box overlap, and filter the final target box position and category prediction results; S39. Output the final recognition result of each water target, wherein the final recognition result includes the predicted category label, the center coordinates of the bounding box, the size information and the confidence score.

[0012] Optionally, the S4 specifically includes: S41. Based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information corresponding to each time step, the three types of features are concatenated and mapped uniformly to construct the time series input vector: ; in, represents the input vector at the jth time step, and represents the linear projection matrix and bias term of the stitching vector, represents the one-dimensional flattened vector of the feature map at step j, represents the feature map extracted by the EfficientNet network, represents the category label, Represents spatial location information, Represents a vector concatenation operation; S42. Add the time position code to the time series input vector to obtain the embedding vector: ; in, represents the j-th time embedding vector, represents the input vector of the jth time step, d represents the input vector dimension, k represents the dimension index, sin and cos represent the position encoding function, Represents the position encoding vector constructed for all dimensions; S43, input the embedding vector sequence into the multi-head attention module in the Motionformer network, extract the context-dependent features of each time step, and calculate the attention-weighted embedding vector, wherein the Motionformer network includes a time position encoding module, a multi-head attention module, a feedforward neural network module, and a temporal aggregation module: ; in, represents the weighted embedding vector of the jth step, Softmax represents the normalization operation, Z represents the temporal embedding vector sequence, , and denote the weight matrices of the query, key, and value matrices of the h-th attention head, respectively. T denotes the matrix transpose operation. represents the dimension of the key matrix, H represents the number of attention heads, and Represents the attention projection matrix and the corresponding bias term; S44, embedding vector after attention weighting at each time step Input the feedforward neural network to extract the motion trajectory features of the corresponding time step: ; in, represents the motion trajectory characteristics of the jth step, represents the dynamic behavior of the j-th step, and represents the feedforward layer weight matrix, and represents the feedforward layer bias term; S45. Perform temporal aggregation processing on the motion trajectory features of all time steps, and use average pooling to generate the dynamic behavior representation of the target: ; Among them, B represents the dynamic behavior representation, Represents the motion trajectory characteristics of the jth step.

[0013] Optionally, the S5 specifically includes: S51. Get the output category label sequence , spatial position information sequence , Dynamic Behavior Representation Sequence ; S52, introduce category credibility weight factor, spatial position weight factor, behavior representation weight factor, define confidence weighting of each time step based on feature change rate, and uniformly construct feature fusion vector. The weight factor is adaptively generated by exponential function of feature change between previous and next two frames: ; in, represents the feature fusion vector, represents the category credibility weight factor, represents the spatial position weight factor, represents the behavior weight factor, represents vector concatenation, represents the category label, Represents spatial location information, Represents dynamic behavior representation; S53, performing average aggregation based on saliency information gating on the feature fusion vector to obtain a fusion representation vector: ; in, represents the fused representation vector, Indicates whether the feature change between adjacent frames is below the threshold, which is used to remove abnormal frames and enhance stability. represents the feature fusion vector; S54, splitting the fusion representation vector into three segments, which respectively represent the recognition type, average spatial position and dynamic behavior representation of the target, as basic data for the recognition report; S55. Set the confidence inference function and generate the confidence score: ; in, represents the confidence score, represents the Sigmoid activation function, , and represents the weight, represents the category distribution entropy, Indicates the spatial location similarity, Indicates the intensity of dynamic behavior; S56. Output a recognition report, wherein the recognition report includes a category, spatial location information, a behavior description representation, and a confidence score.

[0014] A multi-dimensional fusion recognition system for water targets according to an embodiment of the present invention includes: An image acquisition module, used to collect perception image data in a water surface environment; A preprocessing module, used for preprocessing the perceived image data; A multimodal fusion module, used for performing multimodal fusion based on the preprocessed image data to generate fused image data; The target detection module is used to input the fused image data into the EfficientDet target detection network, extract multi-scale feature maps, perform scale fusion, upsampling and downsampling operations based on the bidirectional feature pyramid network, and output the category label and spatial location information of the water target; The trajectory modeling module is used to construct a time series input vector based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information, and input it into the Motionformer network to extract the motion trajectory features of the target, perform time series aggregation, and generate dynamic behavior representation; The recognition report generation module is used to extract the final recognition type, spatial position, dynamic behavior representation and confidence score based on the category label, spatial position information and dynamic behavior representation, and output the recognition report; The model adaptive update module is used to compare the recognition report with the historical recognition data and update the parameters of the EfficientDet network and the Motionformer network online based on incremental learning.

[0015] The beneficial effects of the present invention are: The present invention provides a multi-dimensional fusion recognition method and system for water targets. To address the problems of existing surface target recognition methods, such as single perception modality, rough fusion strategy, weak dynamic behavior modeling, and lack of adaptive update capability of the system, a fusion recognition framework combining end-to-end linkage, multi-modal input, time series modeling and recognition enhancement is constructed, which improves the recognition accuracy, behavior understanding capability and system stability of multiple types of targets in complex water environments.

[0016] First of all, the present invention collects multimodal perception image data including radar images, visible light images and infrared images, and performs format conversion, spatial registration, time synchronization, denoising and normalization processing on them, thereby ensuring the consistency and alignment of image data between different modalities, providing high-quality input data for subsequent fusion operations, effectively making up for the problem of incomplete information of single modality images in specific scenarios, and improving the system's adaptability to complex water surface environments.

[0017] Secondly, in the feature extraction and detection stage, the present invention inputs the preprocessed fused image data into the EfficientDet target detection network with EfficientNet as the backbone network, performs bidirectional fusion processing on the multi-scale feature maps through the feature pyramid network, and combines the target classification subnetwork and the bounding box regression subnetwork for fine detection, thereby improving the recognition accuracy of small targets and occluded targets while improving the detection efficiency. At the same time, the feature sampling based on the anchor box position and the bounding box regression coordinate prediction process realizes the accurate extraction of the target spatial position information, ensuring the spatial consistency of subsequent trajectory modeling.

[0018] In addition, the present invention introduces the Motionformer network structure, jointly constructs a time series input with the feature map extracted by EfficientNet, the target category label and the spatial position information, and uses the time position encoding, the multi-head temporal attention mechanism and the feedforward network module to extract the context-dependent features of the target in continuous time steps, and generates a robust dynamic behavior representation through the average pooling operation, thus achieving a leap from static detection to dynamic understanding. This method can effectively identify the continuous movement characteristics, behavior pattern changes and potential abnormal states of the target, and provide a high-quality basic representation for behavior recognition and safety warning.

[0019] Finally, in terms of the output of recognition results, the present invention not only integrates the category information and spatial position under multiple time steps, but also introduces a confidence weighting mechanism based on the feature change rate, constructs a feature fusion vector and performs gated aggregation of saliency information, effectively filters low-confidence frames, and improves the stability and credibility of the final recognition output. At the same time, a confidence inference function is set, and category distribution entropy, position similarity and behavior intensity indicators are introduced to generate a quantifiable confidence score, so that the recognition report includes not only the target type and location coordinates, but also the behavior description and confidence level, which enhances the interpretability and practicality of the recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 The present invention provides a flowchart of a multi-dimensional fusion recognition method for water targets. DETAILED DESCRIPTION

[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0022] refer to Figure 1 , a multi-dimensional fusion recognition method for water targets, comprising the following steps: S1, collecting perception image data in a water surface environment, and preprocessing the perception image data; S2, performing multimodal fusion based on the preprocessed image data to generate fused image data; S3, input the fused image data into the EfficientDet target detection network, the EfficientDet target detection network uses the EfficientNet network as the backbone network to extract feature maps, input the feature maps into the bidirectional feature pyramid network, perform feature fusion to generate fused feature maps, and output the category labels and spatial position information of the water targets through the bounding box regression subnetwork and the target classification subnetwork; S4, based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information of the water target to construct a time series input, input it into the Motionformer network, extract the motion trajectory features of the water target, and generate a dynamic behavior representation; S5, generating a recognition report based on the category label, spatial location information and dynamic behavior representation; S6. Compare the recognition report with the historical recognition data and use incremental learning to update the network parameters online.

[0023] The present invention realizes high-precision detection, trajectory modeling and dynamic behavior recognition of water targets in complex environments by constructing a recognition process based on the combination of multimodal image fusion and time series modeling. The method integrates spatial information and temporal information, and opens up a complete chain from data acquisition, image preprocessing, multimodal fusion, target detection to behavior analysis and model updating. It has the advantages of strong robustness, stable recognition, strong behavior modeling ability and adaptive optimization, and is suitable for various application scenarios such as water transportation, port monitoring, and maritime security.

[0024] In this implementation, the perceived image data includes radar images, visible light images, and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising, and normalization.

[0025] The present invention introduces three types of multi-source image data, namely radar images, visible light images and infrared images, and combines preprocessing steps such as format conversion, spatial alignment, time synchronization, denoising and normalization to improve the structural consistency and spatiotemporal alignment accuracy between different modal data, lays a high-quality input foundation for subsequent feature extraction and fusion, and enhances the system's adaptability to complex weather, lighting and water surface disturbances.

[0026] In this implementation, S3 specifically includes: S31. Perform tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, and C represents the number of modal channels after splicing. The input tensor X is input to the EfficientDet target detection network, and EfficientNet is called as the backbone feature extraction network. The EfficientDet target detection network includes an EfficientNet network, a bidirectional feature pyramid network, a target classification subnetwork, and a bounding box regression subnetwork. The bidirectional feature pyramid network is used for the fusion of features of different scales. S32. In the EfficientNet network, with tensor X as input, standard convolution, MBConv module concatenation and spatial downsampling operations are performed in sequence to generate feature maps layer by layer. The MBConv module includes a depthwise separable convolution, an inverted residual structure, a residual connection and an activation function: ; in, Indicates The layer outputs a feature map, Indicates The layer outputs the feature map, SE represents the channel attention module, which adjusts the weight of each channel through the two operations of compression and excitation to enhance the attention to important features. represents the point-by-point convolution operation, The scale is Depthwise separable convolution operation; S33. Perform compound scaling configuration on the EfficientNet network. Set the original EfficientNet network parameters as depth, width and resolution, define the scaling factor, and perform compound scaling calculation on each parameter: ; in, , and They represent the depth, width and resolution after compound scaling, d, w and r represent the depth, width and resolution of the original EfficientNet network, respectively. , and represents the scaling factor for depth, width, and resolution, represents the scaling factor; S34, in the EfficientNet network after compound scaling, extract feature maps of different scales, and output the feature map set P in sequence according to the spatial resolution; S35, input the feature map set into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling and downsampling operations, and calculate the fused feature map using a learnable weighted fusion method at each fusion node: ; in, represents the fusion feature map of the i-th layer, represents the fusion weight of the feature map, represents the jth input feature map of the i-th layer, Represents a constant that prevents division by zero; S36. For each fused feature map The feature vector sampling operation is performed at the corresponding position of the anchor frame in the feature map, and the feature vector at a specific position is sampled from the feature map according to the center position of the anchor frame; S37, input the feature vector of the corresponding position of the anchor box in the fusion feature map to the bounding box regression sub-network, perform three one-dimensional convolution operations and one activation function operation in the bounding box regression sub-network in sequence, generate a bounding box regression intermediate feature representation, and continue to input the bounding box regression intermediate feature representation into the second and third convolution layers to output the bounding box coordinate offset vector: ; in, represents the bounding box coordinate offset vector of the jth anchor box at the i-th layer, Indicates the lateral offset ratio of the predicted box relative to the center point of the anchor box. Indicates the longitudinal offset ratio, represents the logarithmic scale shift of the width, represents the logarithmic scale shift of the height; S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector: ; Among them, x and y represent the coordinates of the center point of the predicted bounding box, w and h represent the width and height of the bounding box, , , and Represents the center coordinates and size parameters of the anchor box; Perform non-maximum suppression processing on all prediction results, set the category confidence threshold and intersection-over-union ratio threshold, remove redundant detection results based on score sorting and box overlap, and filter the final target box position and category prediction results; S39. Output the final recognition result of each detected target, wherein the final recognition result includes the predicted category label, the center coordinates of the bounding box, the size information and the confidence score.

[0027] The present invention introduces the EfficientDet target detection network in the target detection stage, and combines the efficient feature extraction capability of the EfficientNet backbone network with the multi-scale fusion advantages of the bidirectional feature pyramid network. Through tensor construction, multi-scale feature generation, bounding box regression and classification sub-network collaborative design, it realizes accurate detection of small targets, long-distance targets and occluded targets on the water. At the same time, the flexibility of the detection network under different scene resolutions is improved through the composite scaling mechanism, effectively taking into account the detection accuracy and computational efficiency, and meeting the application requirements of real-time detection in complex water environments.

[0028] In this implementation, S4 specifically includes: S41. Based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information corresponding to each time step, the three types of features are concatenated and mapped uniformly to construct the time series input vector: ; in, represents the input vector at the jth time step, and represents the linear projection matrix and bias term of the stitching vector, represents the one-dimensional flattened vector of the j-th step feature map, represents the feature map extracted by the EfficientNet network, represents the category label, Represents spatial location information, Represents a vector concatenation operation; S42. Add the time position code to the time series input vector to obtain the embedding vector: ; in, represents the j-th time embedding vector, represents the input vector of the jth time step, d represents the input vector dimension, k represents the dimension index, sin and cos represent the position encoding function, Represents the position encoding vector constructed for all dimensions; S43, input the embedding vector sequence into the multi-head attention module in the Motionformer network, extract the context-dependent features of each time step, and calculate the attention-weighted embedding vector, wherein the Motionformer network includes a time position encoding module, a multi-head attention module, a feedforward neural network module, and a temporal aggregation module: ; in, represents the weighted embedding vector of the jth step, Softmax represents the normalization operation, Z represents the temporal embedding vector sequence, , and denote the weight matrices of the query, key, and value matrices of the h-th attention head, respectively. T denotes the matrix transpose operation. represents the dimension of the key matrix, H represents the number of attention heads, and Represents the attention projection matrix and the corresponding bias term; S44, embedding vector after attention weighting at each time step Input the feedforward neural network to extract the motion trajectory features of the corresponding time step: ; in, represents the motion trajectory characteristics of the jth step, represents the dynamic behavior of the j-th step, and represents the feedforward layer weight matrix, and represents the feedforward layer bias term; S45. Perform temporal aggregation processing on the motion trajectory features of all time steps, and use average pooling to generate the dynamic behavior representation of the target: ; Among them, B represents the dynamic behavior representation, Represents the motion trajectory characteristics of the jth step.

[0029] The present invention constructs a trajectory modeling network with Motionformer network as the core, introduces time position encoding and multi-head attention mechanism in the time dimension, realizes the extraction of cross-frame dynamic features of the target and the modeling of temporal context, and effectively characterizes the motion trajectory and behavior characteristics of the target in continuous time steps. At the same time, the time series input vector is constructed through unified mapping, and the dynamic behavior representation is generated by combining the feedforward neural network and the pooling aggregation module, which improves the system's recognition ability of abnormal behavior, interactive behavior and complex trajectory of the target, and provides strong support for generating accurate and explainable recognition reports.

[0030] In this implementation manner, S5 specifically includes: S51. Get the output category label sequence , spatial position information sequence , Dynamic Behavior Representation Sequence ; S52, introduce category credibility weight factor, spatial position weight factor, behavior representation weight factor, define confidence weighting of each time step based on feature change rate, and uniformly construct feature fusion vector. The weight factor is adaptively generated by exponential function of feature change between previous and next two frames: ; in, represents the feature fusion vector, represents the category credibility weight factor, represents the spatial position weight factor, represents the behavior weight factor, represents vector concatenation, represents the category label, Represents spatial location information, Represents dynamic behavior representation; S53, performing average aggregation based on saliency information gating on the feature fusion vector to obtain a fusion representation vector: ; in, represents the fused representation vector, Indicates whether the feature change between adjacent frames is below the threshold, which is used to remove abnormal frames and enhance stability. represents the feature fusion vector; S54, splitting the fusion representation vector into three segments, which respectively represent the recognition type, average spatial position and dynamic behavior representation of the target, as basic data for the recognition report; S55. Set the confidence inference function and generate the confidence score: ; in, represents the confidence score, represents the Sigmoid activation function, , and represents the weight, represents the category distribution entropy, Indicates the spatial location similarity, Indicates the intensity of dynamic behavior;

[0031] S56. Output a recognition report, wherein the recognition report includes a category, spatial location information, a behavior description representation, and a confidence score.

[0032] The present invention introduces saliency information gating and confidence weighted fusion mechanism in the recognition result generation stage, deeply integrates category labels, spatial location information and behavior representation, and improves the credibility and discrimination ability of recognition results through dynamic feature consistency judgment and confidence reasoning function. The generated recognition report not only contains static target attributes, but also integrates dynamic behaviors and confidence scores, with stronger expressiveness and decision-making value. At the same time, the fused representation vector can be used to guide the incremental update of the model, providing a basis for subsequent online learning, so that the system has the ability of continuous learning and self-evolution.

[0033] A multi-dimensional fusion recognition system for water targets, comprising: An image acquisition module, used to collect perception image data in a water surface environment; A preprocessing module, used for preprocessing the perceived image data; A multimodal fusion module, used for performing multimodal fusion based on the preprocessed image data to generate fused image data; The target detection module is used to input the fused image data into the EfficientDet target detection network, extract multi-scale feature maps, and perform scale fusion, upsampling and downsampling operations based on the bidirectional feature pyramid network to output the category label and spatial location information of the water target; The trajectory modeling module is used to construct a time series input vector based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information, and input it into the Motionformer network to extract the motion trajectory features of the target, perform time series aggregation, and generate dynamic behavior representation; The recognition report generation module is used to extract the final recognition type, spatial position, dynamic behavior representation and confidence score based on the category label, spatial position information and dynamic behavior representation, and output the recognition report; The model adaptive update module is used to compare the recognition report with the historical recognition data and update the parameters of the EfficientDet network and the Motionformer network online based on incremental learning.

[0034] Embodiment 1: In order to verify the feasibility of the present invention in implementation, the present invention is applied to the intelligent port surface monitoring system of a coastal city. The average daily ship traffic in the port exceeds 100 times, and it has typical complex environmental characteristics of water targets such as high-frequency operations and concurrent operation of multiple types of ships. The experimental monitoring area is selected at the intersection of the main channel of the port and the berthing area, which is prone to abnormal behaviors such as congestion of surface targets, illegal docking, and reverse crossing, and is a key water area for supervision.

[0035] The port regulator has deployed a group of fixed shore-based cameras that support visible light, infrared and microwave radar imaging respectively, and dynamically adjust the viewing angle through an optoelectronic gimbal. The perceived image data is collected at 1 frame per second for 72 hours, covering different lighting and meteorological conditions such as daytime, nighttime, early morning, rainy days and foggy days. The image acquisition module in the present invention continuously perceives the image data of the area and hands it over to the preprocessing module to perform image format conversion, modal alignment, time synchronization and noise removal operations, providing high-quality data input for subsequent fusion and detection.

[0036] In the system, the three modal images are sent to the multimodal fusion module after preprocessing, and the fused image is then input into the EfficientDet target detection network. After extracting multi-scale spatial features in the backbone network EfficientNet, scale fusion is completed through the bidirectional feature pyramid network. The target classification subnetwork outputs the target category, and the bounding box regression subnetwork outputs the location information. The detected targets include large container ships, small tugboats, unmanned boats, floating rafts and other types of water targets.

[0037] The detection results enter the Motionformer trajectory modeling network, which extracts motion trajectory features through continuous time series input combined with position encoding and multi-head attention mechanism, and then forms dynamic behavior representation through the time series aggregation module. On this basis, the system integrates category labels, spatial position information and dynamic behavior representation, performs saliency screening and confidence score calculation, and finally generates a structured recognition report, including: target type (such as "tugboat"), position coordinates (such as "x=251, y=122, w=41, h=36"), behavior description (such as "continuous retrograde crossing") and confidence score (such as "0.91").

[0038] In order to evaluate the performance of the system of the present invention, a comparative experiment was conducted with the traditional single-modal YOLOv5 detection system using the same data source. The evaluation indicators included target detection accuracy, recall rate, behavior recognition accuracy, report confidence stability, and performance improvement after model update.

[0039] Table 1. Summary of experimental performance comparison between the invented system and the traditional YOLOv5 system Project indicators The present invention YOLOv5 Indicator Explanation Average target detection accuracy (%) 92.5 79.8 Average recognition accuracy in all scenarios Target detection accuracy in foggy weather (%) 89.7 71.6 Detection performance in low-visibility environments Abnormal behavior recognition accuracy (%) 90.6 73.1 Can it identify behaviors such as driving in the wrong direction and abnormal parking? Trajectory continuity score (0~1) 0.87 0.69 Smoothness and coherence of behavior representation in time series Report confidence score fluctuation CV value 0.13 0.27 Reports the stability of the confidence score (smaller is more stable) False alarm rate (%) 2.3 6.9 The proportion of false positives for non-existent targets or behaviors Accuracy improvement after incremental model update (%) +7.9 Not supported Adaptive learning performance of the model after 24 hours Anomaly detection response time reduction (seconds) 1.2 Not supported Improvement in anomaly detection speed after model iteration Average recognition report generation time (ms) 162ms 199ms Time consumption from detection to generating a complete report for a single target Nighttime recognition accuracy (%) 93.2 78.4 Target recognition capability in the absence of natural light ; It can be seen from the experimental results that the multi-dimensional fusion recognition system for water targets proposed in the present invention is superior to the traditional YOLOv5 single-modal detection system in multiple key performance indicators.

[0040] In terms of overall recognition accuracy, the average target detection accuracy of the system of the present invention in all scenarios reaches 92.5%, which is 12.7 percentage points higher than the 79.8% of the YOLOv5 system, demonstrating a stronger recognition capability for multiple types of targets on the water surface after integrating multimodal image data and deep detection networks.

[0041] Under extreme weather conditions such as fog, the system of the present invention still maintained a detection accuracy of 89.7%, while the accuracy of the YOLOv5 system dropped to 71.6%, verifying the robustness of multimodal perception input in weak visual environments. More importantly, in terms of behavior recognition, the system of the present invention uses the Motionformer network to model motion trajectories and achieves an abnormal behavior recognition accuracy of 90.6%, which is higher than the 73.1% of the YOLOv5 system, and is more accurate and sensitive in identifying dynamic behaviors such as retrograde and abnormal parking.

[0042] In terms of system stability, the volatility of the recognition report confidence score (CV value) of the present invention is only 0.13, which is much lower than 0.27 of the comparison system, indicating that in different time frames and different target conditions, the recognition confidence output by the present invention is more stable and controllable. The trajectory continuity consistency score reaches 0.87, which is also higher than 0.69 of the YOLOv5 system, indicating that the generated behavior representation has better coherence and credibility in the time dimension.

[0043] In terms of error recognition, the false alarm rate of the system of the present invention is only 2.3%, which is much lower than 6.9% of YOLOv5, showing the superior performance of the fusion feature and confidence screening mechanism in suppressing interference targets and filtering false positives. More importantly, the incremental learning mechanism introduced in the present invention realizes the adaptive update of the model. After 24 hours of deployment and operation, the model recognition accuracy is improved by an average of 7.9%, and the anomaly detection response time is shortened by 1.2 seconds. This enables the system to continuously adapt to the dynamically changing water surface environment and maintain stable and efficient recognition capabilities.

[0044] In addition, from the perspective of overall system efficiency, the average time consumed by the present invention in generating a complete recognition report is 162 milliseconds, which is faster than the 199 milliseconds of the YOLOv5 system, demonstrating the high efficiency of the end-to-end optimization structure in reasoning performance. Especially under night conditions, the accuracy of the system of the present invention is still as high as 93.2%, which is significantly ahead of the 78.4% of YOLOv5, verifying the actual effect of the fusion of infrared and radar modes in nighttime.

[0045] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A multi-dimensional fusion recognition method for water targets, characterized in that: The steps include: S1, collecting perception image data in a water surface environment, and preprocessing the perception image data; S2, performing multimodal fusion based on the preprocessed image data to generate fused image data; S3, input the fused image data into the EfficientDet target detection network, the EfficientDet target detection network uses the EfficientNet network as the backbone network to extract feature maps, input the feature maps into the bidirectional feature pyramid network, perform feature fusion to generate fused feature maps, and output the category labels and spatial position information of the water targets through the bounding box regression subnetwork and the target classification subnetwork; S4, based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information of the water target to construct a time series input, input it into the Motionformer network, extract the motion trajectory features of the water target, and generate a dynamic behavior representation; S5, generating a recognition report based on the category label, spatial location information and dynamic behavior representation; S6. Compare the recognition report with the historical recognition data and use incremental learning to update the network parameters online.

2. A multi-dimensional fusion recognition method for water targets according to claim 1, characterized in that: The perceived image data includes radar images, visible light images and infrared images, and the preprocessing includes format conversion, spatial registration, time synchronization, denoising and normalization.

3. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S3 specifically includes: S31. Perform tensor format conversion operation on the fused image data to construct an input tensor , where H represents the image height, W represents the image width, and C represents the number of modal channels after splicing. The input tensor X is input to the EfficientDet target detection network, and EfficientNet is called as the backbone feature extraction network. The EfficientDet target detection network includes an EfficientNet network, a bidirectional feature pyramid network, a target classification subnetwork, and a bounding box regression subnetwork. The bidirectional feature pyramid network is used for the fusion of features of different scales. S32. In the EfficientNet network, with tensor X as input, standard convolution, MBConv module concatenation and spatial downsampling operations are performed in sequence to generate feature maps layer by layer. The MBConv module includes a depthwise separable convolution, an inverted residual structure, a residual connection and an activation function: ; in, Indicates The layer outputs a feature map, Indicates The layer outputs the feature map, SE represents the channel attention module, which adjusts the weight of each channel through the two operations of compression and excitation to enhance the attention to important features. represents the point-by-point convolution operation, The scale is Depthwise separable convolution operation; S33. Perform compound scaling configuration on the EfficientNet network. Set the original EfficientNet network parameters as depth, width and resolution, define the scaling factor, and perform compound scaling calculation on each parameter: ; in, , and They represent the depth, width and resolution after compound scaling, d, w and r represent the depth, width and resolution of the original EfficientNet network, respectively. , and represents the scaling factor for depth, width, and resolution, represents the scaling factor; S34, in the EfficientNet network after compound scaling, extract feature maps of different scales, and output the feature map set P in sequence according to the spatial resolution; S35, input the feature map set into the bidirectional feature pyramid network, perform bidirectional connection, scale fusion, upsampling and downsampling operations, and calculate the fused feature map using a learnable weighted fusion method at each fusion node: ; in, represents the fusion feature map of the i-th layer, represents the fusion weight of the feature map, represents the jth input feature map of the i-th layer, Represents a constant that prevents division by zero; S36. For each fusion feature map The feature vector sampling operation is performed at the corresponding position of the anchor frame in the feature map, and the feature vector at a specific position is sampled from the feature map according to the center position of the anchor frame; S37, input the feature vector of the corresponding position of the anchor box in the fusion feature map to the bounding box regression sub-network, perform three one-dimensional convolution operations and one activation function operation in the bounding box regression sub-network in sequence, generate a bounding box regression intermediate feature representation, and continue to input the bounding box regression intermediate feature representation into the second and third convolution layers to output the bounding box coordinate offset vector: ; in, represents the bounding box coordinate offset vector of the jth anchor box at the i-th layer, Indicates the lateral offset ratio of the predicted box relative to the center point of the anchor box. Indicates the longitudinal offset ratio, represents the logarithmic scale shift of the width, represents the logarithmic scale shift of the height; S38. Call the initial parameters of the anchor box and perform the bounding box position decoding operation in combination with the bounding box coordinate offset vector: ; Among them, x and y represent the coordinates of the center point of the predicted bounding box, w and h represent the width and height of the bounding box, , , and Represents the center coordinates and size parameters of the anchor box; Perform non-maximum suppression processing on all prediction results, set the category confidence threshold and intersection-over-union ratio threshold, remove redundant detection results based on score sorting and box overlap, and filter the final target box position and category prediction results; S39. Output the final recognition result of each water target, wherein the final recognition result includes the predicted category label, the center coordinates of the bounding box, the size information and the confidence score.

4. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S4 specifically includes: S41. Based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information corresponding to each time step, the three types of features are concatenated and mapped uniformly to construct the time series input vector: ; in, represents the input vector at the jth time step, and represents the linear projection matrix and bias term of the stitching vector, represents the one-dimensional flattened vector of the j-th step feature map, represents the feature map extracted by the EfficientNet network, represents the category label, Represents spatial location information, Represents a vector concatenation operation; S42. Add the time position code to the time series input vector to obtain the embedding vector: ; in, represents the j-th time embedding vector, represents the input vector of the jth time step, d represents the input vector dimension, k represents the dimension index, sin and cos represent the position encoding function, Represents the position encoding vector constructed for all dimensions; S43, input the embedding vector sequence into the multi-head attention module in the Motionformer network, extract the context-dependent features of each time step, and calculate the attention-weighted embedding vector, wherein the Motionformer network includes a time position encoding module, a multi-head attention module, a feedforward neural network module, and a temporal aggregation module: ; in, represents the weighted embedding vector of the jth step, Softmax represents the normalization operation, Z represents the temporal embedding vector sequence, , and denote the weight matrices of the query, key, and value matrices of the h-th attention head, respectively. T denotes the matrix transpose operation. represents the dimension of the key matrix, H represents the number of attention heads, and Represents the attention projection matrix and the corresponding bias term; S44, embedding vector after attention weighting at each time step Input the feedforward neural network to extract the motion trajectory features of the corresponding time step: ; in, represents the motion trajectory characteristics of the jth step, represents the dynamic behavior of the j-th step, and represents the feedforward layer weight matrix, and represents the feedforward layer bias term; S45. Perform temporal aggregation processing on the motion trajectory features of all time steps, and use average pooling to generate the dynamic behavior representation of the target: ; Among them, B represents the dynamic behavior representation, Represents the motion trajectory characteristics of the jth step.

5. The multi-dimensional fusion recognition method for water targets according to claim 1 is characterized in that: The S5 specifically includes: S51. Get the output category label sequence , spatial position information sequence , Dynamic Behavior Representation Sequence ; S52, introduce category credibility weight factor, spatial position weight factor, behavior representation weight factor, define confidence weighting of each time step based on feature change rate, and uniformly construct feature fusion vector. The weight factor is adaptively generated by exponential function of feature change between previous and next two frames: ; in, represents the feature fusion vector, represents the category credibility weight factor, represents the spatial position weight factor, represents the behavior weight factor, represents vector concatenation, represents the category label, Represents spatial location information, Represents dynamic behavior representation; S53, performing average aggregation based on saliency information gating on the feature fusion vector to obtain a fusion representation vector: ; in, represents the fused representation vector, Indicates whether the feature change between adjacent frames is below the threshold, which is used to remove abnormal frames and enhance stability. represents the feature fusion vector; S54, splitting the fusion representation vector into three segments, which respectively represent the recognition type, average spatial position and dynamic behavior representation of the target, as basic data for the recognition report; S55. Set the confidence inference function and generate the confidence score: ; in, represents the confidence score, represents the Sigmoid activation function, , and represents the weight, represents the category distribution entropy, Indicates the spatial location similarity, Indicates the intensity of dynamic behavior; S56. Output a recognition report, wherein the recognition report includes a category, spatial location information, a behavior description representation, and a confidence score.

6. A multi-dimensional fusion recognition system for water targets, which executes the multi-dimensional fusion recognition method for water targets according to any one of claims 1 to 5, characterized in that: include: An image acquisition module, used to collect perception image data in a water surface environment; A preprocessing module, used for preprocessing the perceived image data; A multimodal fusion module, used for performing multimodal fusion based on the preprocessed image data to generate fused image data; The target detection module is used to input the fused image data into the EfficientDet target detection network, extract multi-scale feature maps, and perform scale fusion, upsampling and downsampling operations based on the bidirectional feature pyramid network to output the category label and spatial location information of the water target; The trajectory modeling module is used to construct a time series input vector based on the feature map extracted by the EfficientNet network, combined with the category label and spatial position information, and input it into the Motionformer network to extract the motion trajectory features of the target, perform time series aggregation, and generate dynamic behavior representation; The recognition report generation module is used to extract the final recognition type, spatial position, dynamic behavior representation and confidence score based on the category label, spatial position information and dynamic behavior representation, and output the recognition report; The model adaptive update module is used to compare the recognition report with the historical recognition data and update the parameters of the EfficientDet network and the Motionformer network online based on incremental learning.

Citation Information

Patent Citations

  • Urban river water multi-target detection method and system based on DCBFFNet

    CN114973054A

  • Water surface target detection method and system based on improved Officientdet

    CN116311092A

  • Lung CT image aging evaluation method and system based on improved CNN

    CN118537694A

  • Automatic ship berthing and leaving method based on video monitoring and image recognition

    CN119200588A

  • Ship overwater multi-target detection method

    CN119206640A

Cited By

  • Water surface target height calculation method, system and device and computer readable storage medium

    CN120672824A

  • Language learning dynamic resource configuration and interaction system based on Internet platform

    CN120928957A

  • Target identification method based on monitoring video

    CN121053606A

  • Medical report generation method, device and equipment based on multi-branch feature fusion

    CN121122553A

  • Fishing boat operation mode identification method based on multiple modes

    CN121259759A