Intelligent ship detection system based on convolutional neural network and radar signal processing

The intelligent ship detection system, which combines convolutional neural networks and radar signal processing, solves the problems of detection accuracy and environmental adaptability of traditional methods in complex maritime scenarios. It achieves high-precision 3D ship reconstruction and intelligent risk warning, thereby improving the intelligence and automation level of water traffic management.

CN120953931APending Publication Date: 2025-11-14JIANGSU HUASHUN INTELLIGENT TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510983052.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional methods struggle to detect and identify ships in complex maritime scenarios with multiple targets and dynamic backgrounds. They have poor environmental adaptability, cannot perform high-level semantic scene analysis and behavior understanding, and are costly and inefficient.

Method used

A ship intelligent detection system based on convolutional neural networks and radar signal processing is adopted, including ship 3D perception modeling, visual feature hierarchical fusion, target detection and anomaly detection modules. Through multimodal data fusion, spatiotemporal registration and feature encoder, high-precision 3D spatial reconstruction and abnormal behavior monitoring are achieved.

Benefits of technology

It significantly improves the accuracy and anti-interference capability of three-dimensional ship detection in complex environments, enhances the robustness and detection accuracy of the system, realizes interpretable identification and intelligent risk warning of ships, and improves the level of maritime traffic safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953931A_ABST
    Figure CN120953931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and radar perception, in particular to a ship intelligent detection system based on a convolutional neural network and radar signal processing, which comprises a ship three-dimensional perception modeling module, a visual feature hierarchical fusion module, a target detection module and an anomaly detection module. The system emphatically utilizes deep learning methods such as a convolutional neural network and the like to realize three-dimensional space modeling and visual feature extraction of a ship by a radar in a water area environment. And a layered adaptive fusion and mutual information enhancement mechanism is adopted. The end-to-end detection model introduces a spatial hierarchy weighting strategy, so that the object detection accuracy and interpretability under the conditions of multi-target density, shielding and dynamic change in a complex water scene are improved. The system realizes continuous tracking, anomaly detection and risk early warning of ship navigation behaviors based on space-time dynamic modeling. The whole scheme has high precision, strong robustness and adaptive ability, and can meet the requirements of ship detection and intelligent management and control in complex water area environments such as smart ports and water traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and radar perception technology, specifically to an intelligent ship detection system based on convolutional neural networks and radar signal processing. Background Technology

[0002] In the fields of intelligent ship monitoring, maritime traffic control, and port automation management, achieving intelligent identification and comprehensive understanding of multiple targets and events in maritime scenarios has become a core support for ensuring safety and improving management efficiency. Existing technologies typically rely on image or video data acquired from surveillance cameras and employ manual feature extraction and rule-based classification to perform preliminary target detection and differentiation. These solutions are mostly applied to traditional ship-background segmentation, simple motion analysis, and limited forms of target tracking. However, with the increasing complexity of maritime scenarios and the strong interference from waves, lighting, and reflections in the aquatic environment, traditional algorithms struggle to achieve scene understanding in multi-target, multi-dynamic contexts. They exhibit weak adaptability to environmental changes, achieving only limited object segmentation and low-level classification, and are unable to perform high-level semantic scene analysis or behavior-based event understanding.

[0003] Traditional methods rely heavily on manual features and are inadequate in dynamically representing ship behavior when faced with multimodal heterogeneous data. They also fall short in ship detection and recognition in complex water environments such as wet and foggy conditions. Even with efforts to improve single-point accuracy by increasing algorithmic layers or investing more hardware resources, these methods are costly and inefficient to deploy and maintain, and still fail to meet the comprehensive requirements of object detection, intelligent event recognition, and risk warning in dynamic and complex scenarios. Currently, radar-based ship detection technology has significant limitations in environmental adaptability, event recognition, and automatic analysis of complex scenes, and cannot effectively meet the ship detection needs of intelligent and automated maritime scenarios.

[0004] To address this, a ship intelligent detection system based on convolutional neural networks and radar signal processing is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent ship detection system based on convolutional neural networks and radar signal processing to solve the problems mentioned in the background art.

[0006] The intelligent ship detection system based on convolutional neural networks and radar signal processing provided by this invention specifically includes the following modules:

[0007] Ship 3D Perception Modeling Module: Acquires radar signals and spatiotemporal attribute data, reconstructs the ship's 3D spatial perception model through spatial perception imaging, and performs spatiotemporal alignment with spatiotemporal attribute data using spatiotemporal registration, and obtains multimodal spatial perception data through data fusion;

[0008] The visual feature hierarchical fusion module employs a feature encoder that combines hierarchical adaptive fusion with mutual information enhancement to output spatial, frequency, and temporal visual features from multimodal spatial perception data and perform feature fusion. Through mutual information enhancement and redundancy suppression, multimodal fused visual features are obtained.

[0009] Target detection module: Constructs an end-to-end interpretable ship detection and prediction model, performs spatial hierarchical weighting and annotation on multimodal fusion visual features, and outputs ship attribute interpretation and 3D spatial distribution prediction results;

[0010] Anomaly Detection Module: Input the ship attribute interpretation and 3D spatial distribution prediction results, monitor the spatial behavior evolution process based on the visual spatiotemporal dynamic modeling mechanism, and output anomaly suppression and behavior prediction information.

[0011] Preferably, the ship 3D perception modeling module includes:

[0012] Through a multi-channel radar signal synchronous acquisition unit, high-resolution radar echo data and spatiotemporal attribute data are collected for different meteorological and water surface conditions in the water environment. Through adaptive beamforming and pulse compression processing, spatial perception imaging is performed, and three-dimensional spatial perception imaging data is output.

[0013] Preferably, the ship's three-dimensional spatial perception model specifically includes:

[0014] Inputting 3D spatial perception imaging data, and performing high-dimensional data unfolding based on azimuth, pitch, distance, and polarization spatial dimensions, generates multidimensional tensor data that combines spatial geometry and physical properties; based on the multidimensional tensor data, constructing a 3D spatial feature representation, extracting the 3D spatial distribution features of the ship target, and outputting a 3D spatial perception model of the ship, a 3D point cloud representation, a voxel feature representation, and a 3D feature dataset, wherein the 3D feature data set includes the target's spatial location, volume, and structural type.

[0015] Preferably, the spatiotemporal registration specifically includes:

[0016] Spatial coordinate registration is performed on 3D feature data groups from different time batches, and temporal synchronization is achieved through timestamp alignment. A unified spatial coordinate system and time axis reference are established. The ship's 3D spatial perception model, 3D point cloud representation, voxel feature representation, and 3D feature dataset are received through a multimodal data fusion method. Feature pairing is performed, and feature fusion is achieved through feature splicing and weighted averaging to obtain a 3D information body containing multidimensional visual semantics, spatial structural elements, and temporal change characteristics. Multimodal spatial perception data is then output.

[0017] Preferably, the visual feature hierarchical fusion module specifically includes:

[0018] The hierarchical adaptive fusion and mutual information enhancement collaborative feature encoder receives multimodal spatial perception data and generates heterogeneous spatial feature groups in spatial, frequency, and temporal dimensions, outputting three types of 3D scene feature data: spatial structure information, frequency distribution information, and temporal change information. The hierarchical adaptive fusion, through multi-level 3D feature processing, aggregates the three types of 3D scene feature data in layers according to spatial attributes, structural attributes, and dynamic attributes, extracts multi-level spatial information fusion features, and outputs 3D information fusion feature data containing 3D attribute layer labels, spatial distribution features, and scene semantics.

[0019] Preferably, the mutual information enhancement collaboration specifically includes:

[0020] The mutual information enhancement method receives hierarchically output 3D information fusion feature data through a collaborative feature learning unit, completes the correlation encoding between features of different modalities based on the mutual information measurement method, optimizes the feature information distribution through a redundancy suppression strategy, and outputs multimodal fusion visual feature data.

[0021] Preferably, the target detection module specifically includes:

[0022] An end-to-end interpretable ship detection and prediction model is constructed. Multimodal fusion visual feature data is input, and feature weights are assigned to spatial feature levels. Feature annotations are performed based on spatial scale, spatial location, and semantic hierarchy. Combined with signal response distribution, and according to ship detection and prediction requirements, ship attribute features are annotated through weight distribution and causal response relationships to obtain multidimensional detection result data containing spatial distribution, key attributes, and prediction information. The multidimensional detection result data is then processed by an attribute interpretation unit to parse the attribute feature annotations, classify and label the ship's type, structural parameters, and current dynamic state, and output ship attribute interpretations. Through 3D coordinate mapping and spatial structure discrimination mechanisms, combined with the spatial distribution, key attributes, and prediction information in the detection result data, the 3D spatial distribution prediction of each ship in the scene is performed, and the 3D spatial distribution prediction result is output.

[0023] Preferably, the anomaly detection module specifically includes:

[0024] Through the spatiotemporal dynamic modeling unit, the interpretation of ship attributes and the prediction results of three-dimensional spatial distribution are input. Based on the visual spatiotemporal dynamic modeling mechanism, the spatiotemporal evolution trajectory of the ship target in the water scene is obtained. Using a multi-time series data analysis method, behavioral sequence change features are extracted from the spatiotemporal evolution trajectory. Referring to the normal state model, a discrimination threshold is set, and target data that differs from the normal spatiotemporal evolution features are labeled. Target data with abnormal spatial location, motion state, and behavioral patterns are aggregated into abnormal behavior data. Through the abnormal feature analysis unit, features are summarized and classified according to abnormal attribute type, spatiotemporal distribution pattern, and behavioral evolution trend. An suppression scoring and risk classification mechanism is used for related abnormal targets to output abnormal detection information. Through the abnormal behavior change sequence, combined with the behavioral evolution law, the corresponding behavior prediction information is output.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] 1. Regarding the accuracy and anti-interference capability of ship 3D detection, by using a radar signal synchronous acquisition unit, combined with adaptive beamforming and pulse compression processing, high-resolution radar echo data under different weather and water surface conditions are spatially perceived and imaged. By adopting unified registration of spatial coordinates and time axis and multimodal data fusion, the accuracy of ship 3D spatial reconstruction under complex environments is significantly improved. This effectively overcomes the problems of low detection accuracy and susceptibility to environmental clutter interference in traditional methods under adverse weather and variable water surface conditions, and realizes high-precision, spatiotemporally aligned ship 3D modeling and attribute extraction.

[0027] 2. Regarding multi-source heterogeneous data fusion and system robustness, a feature encoder with a hierarchical adaptive fusion structure is used to extract and fuse multimodal spatial sensing data in the spatial, frequency, and temporal domains. Combined with mutual information measurement and redundancy suppression strategies, the feature information distribution is optimized, which significantly improves the system's ability to collaboratively process sensor data in multi-source heterogeneous and complex environments. It effectively suppresses the impact of single-point sensor failure or local interference on the detection results, ensuring the robustness and global stability of detection and recognition.

[0028] 3. In terms of ship 3D recognition and attribute interpretability, through an end-to-end ship detection and prediction model, based on multi-level features in the spatial, frequency, and temporal domains, and employing spatial hierarchical weighting and multi-dimensional feature annotation mechanisms, combined with signal response distribution and causal response relationships, it can not only accurately output the 3D position, volume, and structural type of each ship in the scene, but also realize interpretable ship categories, structural parameters, and dynamic feature labels, providing a rich, accurate, and structured information foundation for intelligent control and automated decision-making.

[0029] 4. In terms of abnormal ship behavior detection and intelligent risk warning, through the visual spatiotemporal dynamic modeling unit, multi-time series data analysis is integrated to extract the three-dimensional trajectory, dynamic voxel feature flow and behavioral sequence change features of the target. Combined with the normal state model and abnormal feature analysis and risk classification mechanism, it can proactively and accurately identify the spatial anomalies, dangerous behaviors and risk conditions of ships, realize the proactive discovery, timely classification and early warning intervention of water traffic incidents, and significantly improve the level and response efficiency of water safety management. Attached Figure Description

[0030] Figure 1 A schematic diagram of the structure of a ship intelligent detection system based on convolutional neural networks and radar signal processing provided in an embodiment of the present invention;

[0031] Figure 2 This is a schematic diagram of the structure of a ship 3D perception modeling module provided in an embodiment of the present invention;

[0032] Figure 3 This is a schematic diagram of the visual feature layering fusion module structure provided in an embodiment of the present invention;

[0033] Figure 4 This is a schematic diagram of the target detection module structure provided in an embodiment of the present invention;

[0034] Figure 5 This is a schematic diagram of the anomaly detection module structure provided in an embodiment of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Example 1

[0037] This embodiment takes an automated port area in the coastal waters of location A as the application scenario, and specifically describes the implementation process and actual performance of the system to address the needs of three-dimensional detection, attribute recognition and abnormal behavior monitoring of multi-target ships in complex water environments.

[0038] Reference Figure 1 The diagram below illustrates the structure of a ship intelligent detection system based on convolutional neural networks and radar signal processing, as provided in an embodiment of the present invention. The system includes:

[0039] Ship 3D perception modeling module, such as Figure 2This is a schematic diagram of the ship 3D perception modeling module provided in this embodiment of the invention. The radar signal synchronous acquisition unit theoretically deploys four sets of S-band multi-channel radar equipment in port area A. Each radar has eight receiving channels, achieving 360-degree omnidirectional coverage. Each radar acquisition module is equipped with a high-precision constant-temperature clock reference to achieve nanosecond-level synchronization, and uses a high-speed single-mode fiber optic channel to synchronously transmit all radar echo signals to the central processing server in real time. The equipment parameters refer to existing high-performance radar specifications. Each radar is designed with a maximum sampling rate of 200 MSa / s, a pulse width of 1 microsecond, a range resolution of 1.5 meters, and a velocity measurement accuracy of 0.5 meters per second. The theoretical data volume is derived from batch acquisition simulation. Under continuous 72-hour operation, the average effective data frames per hour are 16,000 frames, and each frame is approximately 140MB after lossless compression. To adapt to different meteorological conditions (such as sunny days, rain, fog, and wind changes) and variable water surface states (such as calm, waves, and swells) in the maritime environment, this embodiment fully considers the impact of environmental diversity on radar echo signal acquisition and processing during the design and deployment of the radar signal synchronous acquisition unit. Specifically, this includes a radar acquisition system that can dynamically adjust key indicators such as sampling rate, pulse width, and beamforming parameters based on real-time meteorological data and water surface conditions. For example, in typical port environment conditions such as wind speeds of 2 to 5, visibility of 200 to 1000 meters, and wave heights of 0.2 to 1.5 meters, the system can automatically optimize radar parameters to ensure the stability and reliability of high-resolution data acquisition. The actual data acquisition process is adaptable to typical port environments, such as wind speeds of 2 to 5 and visibility of 200 to 1000 meters. The radar echo data output by this unit serves as the basis for subsequent spatial imaging and feature modeling, ensuring data synchronization and high resolution, and providing reliable raw information for downstream spatial reconstruction. The beamforming and pulse compression processing unit inputs radar echo data and uses adaptive bandpass filtering and digital beamforming to suppress out-of-band noise and clutter. The main lobe width is automatically adjusted to 3 to 5 degrees to improve signal resolution under complex ship distribution conditions. In the linear frequency modulated pulse compression process, algorithms such as Fast Fourier Transform are used to process each channel signal, enhancing range image resolution and effectively suppressing sidelobe and main lobe noise. The main lobe signal-to-noise ratio is improved from 13 dB to over 31 dB, laying the foundation for high-precision imaging and 3D modeling.

[0040] This unit achieves high-resolution 3D spatial perception imaging of ships and surrounding targets in a maritime environment through synchronous acquisition of multi-channel radar signals, adaptive beamforming, and pulse compression processing. The output data possesses accurate spatial positioning, motion attributes, and structural features, providing high-quality, standardized 3D perception data support for subsequent applications such as target detection, behavior analysis, and intelligent control. It adapts to the actual needs of complex port areas and effectively overcomes the spatial perception blind spots of traditional methods under conditions of multiple targets, weak signals, strong background clutter, and limited field of view.

[0041] Furthermore, after performing high-dimensional unfolding on the spatial perception imaging data, the 3D spatial modeling unit first generates multi-dimensional tensor data that integrates spatial geometry, reflectivity, and polarization features based on parameters of multiple dimensions such as azimuth, pitch, distance, and polarization. Utilizing a multi-resolution modeling mechanism, the system can simultaneously output global coarse-grained and locally high-density 3D point cloud representations, thereby achieving multi-level spatial structure characterization of the entire ship and key components, outputting 3D point cloud representations and voxel feature representations. Combining deep semantic segmentation networks and attribute annotation methods, automatic component segmentation and structural semantic annotation are performed on the 3D point cloud and voxel data, forming hierarchical feature labels with spatial location, structural type, and functional semantics. Within each spatial unit of the 3D data, in addition to geometric features, physical attribute indicators such as reflectivity and confidence are embedded to quantify the reliability of each spatial feature. For the continuous evolution of target behavior, the system continuously acquires 3D point cloud data and voxel features of each target at fixed time intervals (e.g., 1 frame per second). Each frame of data includes the target's spatial coordinates, voxel intensity, reflectivity, velocity vector, and other attributes. By assigning a unique identifier to each target, the system can correlate data of the same target across different time frames to form a complete temporal point cloud sequence. Subsequently, the system employs temporal registration algorithms such as Kalman filtering or Multiple Hypothesis Tracking (MHT) to smooth the trajectory and suppress noise of the target's point cloud center and voxel features, generating a continuous motion trajectory of the target in 3D space. Simultaneously, the system stacks the voxel features of each target across consecutive time frames to form a dynamic voxel feature stream, reflecting the changing trends of the target's structure, attitude, and velocity over time. For multiple targets, the system uses a global data association algorithm to analyze the spatial distance, relative velocity, and interaction events between targets in real time, and automatically labels abnormal behaviors such as sudden acceleration, abnormal turning, and prolonged stillness. Finally, the system outputs the spatiotemporal continuous point cloud trajectory, dynamic voxel feature stream, and behavioral labels for each target, providing a structured and traceable data foundation for subsequent behavior pattern recognition, anomaly detection, and risk prediction. All output feature representations with multi-level, semantically enhanced, and spatiotemporally dynamic characteristics are uniformly encoded into 3D feature data sets in a structured manner. This dataset supports intelligent indexing and efficient retrieval, and can be expanded to integrate other modal labels, providing multidimensional, rich and efficient data support for subsequent detection, recognition, intelligent control and management of massive scene data.

[0042] This unit inputs and performs high-dimensional unfolding of spatial perception imaging data to generate a multidimensional tensor that integrates spatial geometry and physical properties. Employing a 3D-ResNet18 neural network and AIS data calibration, it significantly improves the accuracy of spatial modeling of ship targets detected by radar signals in complex and dynamic port environments, achieving precise representation of the target's 3D point cloud. The output includes a high-density point cloud, a voxel feature model, and a set of 3D feature data, accurately describing the target's spatial location, volume, and structural type, providing a standardized 3D data foundation for subsequent feature registration and multi-source fusion.

[0043] Furthermore, the spatiotemporal registration and multimodal fusion unit takes 3D point cloud representation, voxel feature representation, and 3D feature data sets as inputs, and combines multi-source spatiotemporal attribute data such as time and coordinate axes. Through a spatial coordinate transformation network, it achieves millisecond-level synchronization and timestamp normalization of station point clouds, establishing a unified spatial and temporal reference. The system automatically extracts multiple anchor point targets for cross-coordinate correction, and the maximum theoretical error of single-frame multi-station registration is controlled within 1.5 meters. In the multimodal fusion stage, algorithms are used for automatic matching and weighted stitching of heterogeneous features such as spatial, semantic, and behavioral data to realize the expression of multimodal 3D information volumes.

[0044] The multimodal spatial perception data output by this unit possesses complete spatial structural continuity and uniformity of attribute and temporal variations, greatly enhancing the ability to stably detect and comprehensively correlate targets in nighttime, rainy, foggy, and obstructed environments. This point cloud is aligned with multi-source spatiotemporal attribute data, strongly supporting the accuracy and adaptability of downstream target recognition and behavior modeling.

[0045] Furthermore, the visual feature hierarchical fusion module, such as Figure 3This is a schematic diagram of the visual feature hierarchical fusion module structure provided in an embodiment of the present invention. It includes a hierarchical adaptive fusion structural feature encoder, a multi-level 3D feature processing strategy, and a mutual information enhancement collaborative feature learning mechanism. First, the hierarchical adaptive fusion structural feature encoder receives input from multimodal spatial perception data and extracts fusion features at the spatial, frequency, and temporal domains. The spatial domain encoder captures the spatial distribution and structural attributes of the 3D scene; the frequency domain encoder analyzes the frequency distribution characteristics of radar echoes and multimodal signals; and the temporal domain encoder mines the behavioral evolution and dynamic changes of the target. Each level of the feature encoder dynamically adjusts the fusion weights based on the statistical distribution and spatial correlation of the input data to achieve adaptive hierarchical integration of multi-source heterogeneous features. The multi-level 3D feature processing strategy is based on the encoder output above. It performs hierarchical processing and strategic convergence of spatial structure information, frequency distribution information and temporal change information. The specific process is as follows: At the spatial level, the 3D spatial distribution position and geometric features of the target are extracted. At the frequency level, the high-dimensional spectral response of different targets and backgrounds is analyzed to improve the distinction between moving targets and static / clutter. At the temporal level, the motion trajectory and behavior pattern of the target in the scene are tracked.

[0046] Through a feature encoder that combines hierarchical adaptive fusion and mutual information enhancement, multimodal spatial perception data is received, and heterogeneous spatial feature groups are sequentially extracted and generated hierarchically in the spatial, frequency, and temporal domains. After multi-level 3D feature processing, the three types of scene feature data—spatial structure, frequency distribution, and temporal variation—are hierarchically aggregated to output multi-dimensional information fusion feature data with 3D attribute hierarchical labels, accurate spatial distribution, and scene semantic annotations, providing comprehensive data support for subsequent detection.

[0047] Furthermore, the mutual information-enhanced collaborative feature learning mechanism improves the effectiveness of feature representation based on fusion. This mechanism uses a mutual information-driven feature redundancy suppression network, with mutual information as the core indicator, to automatically suppress redundant features and improve feature representation efficiency. It receives the aforementioned multi-dimensional information fusion feature data and automatically calculates the mutual information and redundancy between feature groups in different spatial, frequency, and temporal domains. Based on a theoretical model, the system assigns priority weights to data blocks with high mutual information and low redundancy between spatial, structural, and dynamic attributes, adaptively adjusting the feature representation distribution. For each pair of feature groups (e.g., spatial and frequency domains, spatial and temporal domains, frequency and temporal domains), its joint probability distribution P(x, y) and marginal probability distributions P(x) and P(y) are calculated. Using the mutual information formula (X; Y) = ∑∑P(x, y)log[P(x, y) / (P(x)P(y))], the mutual information between each feature group is calculated. All feature groups are combined pairwise to obtain three sets of mutual information indicators: spatial-frequency, spatial-temporal, and frequency-temporal. The system further performs redundancy analysis within each feature group. The theoretical value of redundancy can be defined using information entropy. Specifically, the entropy of each feature in the feature group and their joint entropy are calculated first. Then, the proportion of redundant information within the feature group is quantified by comparing the relationship between the joint entropy and the sum of the entropies of each feature. The redundancy of each layer before fusion is derived to be 21%, followed by priority weight allocation and feature optimization. The system assigns priority weights to feature groups based on the principle of high mutual information and low redundancy. Specifically, for feature groups with mutual information higher than a set threshold (e.g., 0.7), their weight in the feature fusion output is increased. For feature groups with redundancy higher than a set threshold (e.g., 0.2), an encoder optimization method is used for dimensionality reduction and redundancy suppression. The overall feature expression structure is optimized by adaptively adjusting the feature representation distribution. The optimization results are output. After mutual information enhancement and encoder optimization, the system recalculates the redundancy and mutual information index of the fused features. The derivation results show that the redundancy is reduced from 21% to 8%, and the overall mutual information index is improved by approximately 12%. The output is multimodal fused visual feature data with low redundancy and strong complementarity.

[0048] This mechanism significantly improves the discriminative power and discrimination effect of fused features in typical complex scenarios, providing higher-quality visual feature input for target detection and behavior analysis modules and enhancing the overall robustness of the system. It employs a mutual information metric to encode the correlation between features of different modalities and combines redundancy suppression strategies to optimize feature distribution. The system adaptively adjusts feature weights, significantly reducing redundancy and improving mutual information metrics, ultimately outputting multimodal fused visual feature data with low redundancy and strong complementarity, providing high-quality input for subsequent target detection and behavior analysis modules.

[0049] Furthermore, the target detection module, such as Figure 4This is a schematic diagram of the target detection module structure provided in an embodiment of the present invention. It includes an end-to-end ship interpretability detection and prediction structure, a key feature annotation and hierarchical weighting unit, and a detection output and attribute annotation unit. This module takes the multimodal fused visual feature data output by the visual feature hierarchical fusion module as input, and sequentially completes the automatic detection, attribute annotation, and three-dimensional spatial distribution prediction of the ship target.

[0050] The end-to-end ship interpretability detection and prediction model first receives multimodal fused visual feature data. Based on the YOLO series detection networks, this model introduces a multimodal feature fusion layer in the backbone network, concatenating or weighting feature tensors from different sources along the channel dimension to form a unified high-dimensional feature representation. The model further introduces a spatial hierarchy weighting mechanism: for features at different spatial scales (e.g., large, medium, and small targets) and different semantic levels (e.g., edges, textures, structures), learnable weight parameters are set, and features from each layer are grouped and weighted for fusion. Before the detection head at each scale, a spatial attention module and a channel attention module are added to dynamically adjust the response intensity of each spatial region and feature channel, thereby improving the ability to distinguish dense targets, occluded targets, and targets in complex backgrounds. During the model training phase, a multi-task loss function is used to jointly optimize the target detection box regression, category classification, and attribute interpretation branches. The attribute interpretation branch outputs labels including ship type, structural parameters, and dynamic state. The model can output the spatial location, confidence score, attribute label, and 3D spatial distribution prediction results for each detected target. Through the aforementioned spatial hierarchical weighting and multimodal fusion mechanisms, the model not only improves detection accuracy but also enhances the interpretability of detection results, providing a high-quality and structured data foundation for subsequent ship attribute annotation and spatial distribution prediction. After fusing multimodal visual features, convolution, pooling, upsampling, and fully connected operations are performed sequentially. The system introduces a spatial hierarchical weighting mechanism to hierarchically group input features based on spatial scale, spatial location, and semantic hierarchy, and dynamically adjusts the weight allocation of multiple feature groups according to the signal response distribution theory model. Theoretically, this processing method can improve the network's generalization ability in detection scenarios with multiple overlapping targets, occlusion, and complex background interference, and is expected to improve overall detection accuracy. To verify the effectiveness of the spatial hierarchical weighting mechanism, the system sets up a control network and an improved network through theoretical derivation. The control network adopts the traditional YOLO structure, while the improved network introduces a spatial hierarchical weighting model, with other parameters remaining the same. Through theoretical modeling, the processing outputs of different detection modules on simulation data are calculated. Using the mean Average Precision (mAP) formula, the model is expected to achieve higher target detection probabilities under conditions of dense multi-target activity, occlusion, and background interference. Based on the extrapolation, the spatially weighted YOLO module improves the overall detection accuracy on the simulation dataset by approximately 11% compared to the traditional YOLO structure. The unit outputs multi-dimensional detection results, including target bounding boxes, confidence scores, and spatial distribution information, providing a foundation for subsequent attribute annotation and spatial prediction.

[0051] The key feature annotation and hierarchical weighting unit further processes the multidimensional detection results data output by the end-to-end detection structure. For each detected ship target in a frame, the system automatically extracts spatial geometric coordinates (X, Y, Z three-dimensional position), contour dimensions (e.g., ship length 110 meters, ship width 14 meters, error not exceeding ±2%), sailing speed (e.g., 6.1 knots, accuracy 0.1 knots), and heading azimuth (0-359 degrees, accuracy better than 1 degree). Each type of feature is modeled hierarchically by an independent parameter network and assigned normalized weights. By fusing feature responses through weighted layers, the ability to distinguish target attributes in complex scenes is improved. All hierarchical annotations are synchronously written to the task index library, facilitating subsequent tracing, visualization, and behavioral modeling. The hierarchical weighting mechanism of this unit can theoretically improve the accuracy and consistency of attribute annotation. In situations with multiple dense targets or blurred local signals, it can effectively reduce the probability of false detection and missed detection. The detection output and attribute annotation unit integrates all key features and automatically generates the detection results and full attribute labels for each ship. The detection results include essential information such as spatial location, shape and size, speed, and target orientation. The attribute interpretation and processing unit analyzes the attribute feature annotations in the detection results to classify and generate labels for ship type, structural parameters, and current dynamic state. Leveraging 3D coordinate mapping and spatial structure discrimination mechanisms, and combining the spatial distribution, key attributes, and prediction information in the detection results data, the system can predict the 3D spatial distribution of each ship in the scene, outputting ship attribute interpretations and 3D spatial distribution prediction results. Furthermore, the output unit can be expanded to receive AIS ship information and integrate multi-source labels to meet the complex operations and safety management needs of smart ports.

[0052] In summary, the target detection module, through the collaborative work of end-to-end ship interpretability detection and prediction structures, key feature annotation and hierarchical weighting units, and detection output and attribute annotation units, efficiently detects, annotates attributes, and predicts the 3D spatial distribution of multimodal fusion visual features. This module not only improves detection accuracy and attribute interpretation capabilities in complex scenarios but also provides high-quality structured data support for subsequent abnormal behavior analysis and intelligent decision-making, significantly enhancing the system's intelligent detection and management capabilities in actual port and waterway environments.

[0053] Furthermore, the anomaly detection module, such as Figure 5 This is a schematic diagram of the anomaly detection module provided in an embodiment of the present invention. It includes a spatiotemporal dynamic modeling unit, an anomaly behavior discrimination unit, and an anomaly data output and early warning unit. This module takes the ship's attribute interpretation and three-dimensional spatial distribution prediction results as input, and monitors and provides anomaly warnings for the entire process of the ship's spatiotemporal behavior evolution in the aquatic environment.

[0054] The spatiotemporal dynamic modeling unit is primarily used to continuously receive ship attribute interpretations and 3D spatial distribution prediction results. This modeling method relies on high-frequency data acquisition, continuously recording multiple dimensions of attributes such as spatial coordinates, speed, and heading for each target ship in real time. The system employs dynamic sliding windows and recursive temporal modeling algorithms. By seamlessly stitching and dynamically updating spatiotemporal data on the target ship's motion trajectory and behavioral evolution throughout its entire lifecycle, it achieves continuous modeling and real-time evolution tracking of target behavior. This spatiotemporal dynamic continuous modeling method can automatically integrate multi-source data from different time periods and spatial locations, accurately reconstructing the ship's complete motion and behavioral change process, providing highly timely and complete spatiotemporal information support for anomaly detection, behavior prediction, and intelligent decision-making. This modeling mechanism can dynamically capture subtle changes in target behavior, detect early signs of abnormal behavior in a timely manner, and make forward-looking predictions of abnormal trends. Compared with traditional segmented modeling methods, spatiotemporal dynamic modeling can more accurately reconstruct the spatiotemporal evolution trajectory of targets in complex aquatic environments. It also improves the system's sensitivity to abnormal behavior detection and response speed, laying a solid data foundation and technical support for subsequent anomaly identification, risk classification, and behavior prediction.

[0055] The abnormal behavior discrimination unit takes the target's spatiotemporal trajectory and behavioral change characteristics generated by the spatiotemporal dynamic modeling unit as input, and performs real-time comparison and discrimination in conjunction with the visual spatiotemporal dynamic modeling mechanism. The system first automatically sets discrimination thresholds based on the normal behavior model. Commonly used discrimination items include algorithm parameters such as spatial offset, velocity mutation amplitude, rotational angular velocity, and continuous behavior duration. For example, for a large cargo ship, if the system continuously detects a change in heading angle exceeding 30 degrees within 10 seconds and a trajectory deviation distance reaching 40 meters, far exceeding the normal fluctuation parameters for that segment, it can trigger abnormal approach and abnormal deviation behavior labeling. All target data judged as abnormal obtains detailed abnormal attribute labels, including abnormality type, spatiotemporal distribution pattern, and abnormal behavior evolution trend indicators. Simultaneously, the abnormal behavior discrimination unit constructs a risk classification and multi-level response linkage mechanism to classify abnormal events by risk and link different response strategies. For each abnormal target, a suppression scoring and risk classification process is triggered, dynamically assigning high-risk, moderate-risk, and low-risk levels. This unit effectively summarizes various abnormal events, significantly reducing the false positive and false negative rates, and greatly improving the system's response accuracy to abnormal targets.

[0056] The abnormal data output and early warning unit is responsible for data encapsulation and output management of the classification results and risk ratings of the abnormal behavior discrimination unit. The system automatically generates an abnormal behavior data packet containing the target number, abnormal type, abnormal location, navigation attributes, and time information, and pushes it to the shore-based management platform or dispatching background within 30 seconds through the internal management network to achieve rapid early warning of abnormal behavior and scenario linkage. At the same time, the module combines the abnormal behavior change sequence and behavior evolution law to predict the subsequent abnormal behavior state, outputs behavior prediction information, and provides prediction support for operation and maintenance control and safety response. All output abnormal information, including alarm signals, risk levels, and prediction data, is encoded in the port management standard format to ensure easy parsing and disposal by on-site personnel and automated platforms. This unit supports abnormal suppression and dynamic early warning processing, and realizes the efficient control and risk controllability of multiple types of abnormal ships in complex water scenarios.

[0057] In summary, based on the interpretation of ship attributes and the prediction results of three-dimensional space distribution, the abnormal detection module combines the visual spatio-temporal dynamic modeling mechanism, and realizes the highly controllable abnormal detection and real-time early warning of ship targets in water scenarios through three sub-units: spatio-temporal dynamic modeling, abnormal behavior discrimination, and abnormal data output and early warning. This mechanism can shorten the abnormal recognition response time to within 30 seconds, greatly improve the shore-based system's monitoring, grading, and proactive early warning capabilities for abnormal events, and effectively ensure the intelligent and safe operation of ports and waters.

[0058] The intelligent ship detection system based on convolutional neural network and radar signal processing provided by the present invention has core functions such as multi-dimensional data acquisition, three-dimensional space modeling, hierarchical feature fusion, end-to-end target detection, and abnormal behavior analysis. The system reconstructs multi-channel radar signals into a high-precision three-dimensional space perception model, significantly improving the space positioning and attribute recognition capabilities of multiple target ships in complex water environments. By using a feature encoder with hierarchical adaptive fusion and mutual information enhancement collaboration, the efficient fusion and redundancy suppression of multi-modal features such as space, frequency, and time series are realized, enhancing the system's adaptability to dynamic scenarios and multi-target overlaps. The end-to-end interpretable detection and prediction model can accurately identify ship targets, interpret attributes, and predict three-dimensional space distribution, and combine the visual spatio-temporal dynamic modeling mechanism to realize real-time monitoring and intelligent early warning of abnormal behaviors. This system effectively solves the problems of low detection accuracy, incomplete attribute recognition, and lagging abnormal response in traditional methods in scenarios with dense multi-targets, complex environments, and strong signal interference, significantly improving the intelligent and automated levels of water traffic safety management, and providing solid technical support for the efficient control and event early warning of typical waters such as ports and waterways.

[0059] Embodiment 2

[0060] This embodiment uses an automated port area in the coastal waters of location A as an application scenario. Focusing on multi-target vessels in complex water environments, it primarily explains the structural composition and main functions of the ship 3D perception and modeling module in a ship intelligent detection system based on convolutional neural networks and radar signal processing, illustrating how it enables the detection of multiple targets in complex water environments. (See attached...) Figure 2 As shown, the ship 3D perception modeling module includes a radar signal synchronous acquisition unit, a beamforming and pulse compression processing unit, a 3D spatial modeling unit, and a spatiotemporal registration and multimodal fusion unit. These units are sequentially connected to form a complete data processing chain, specifically designed for fine 3D modeling of ship targets in complex water environments.

[0061] The main function of the radar signal synchronization acquisition unit is to receive and synchronously acquire multi-channel raw radar echo signals covering a designated water area in real time. The system is configured with multiple shipborne or shore-based radar devices, each equipped with a multi-channel receiving array to fully cover the detection range. All radar devices achieve a unified time reference through an external high-precision timing module, ensuring synchronous acquisition of multi-source radar data at the microsecond level. The raw echo signals undergo preliminary local noise filtering and labeling, and the synchronized radar data is output in a standard format as the sole input data source for subsequent processing stages.

[0062] The beamforming and pulse compression processing unit is used to improve spatial resolution against complex aquatic backgrounds. First, the input is the raw radar echo data output from the radar signal synchronous acquisition unit. For multi-channel signals, the unit employs a digital beamforming algorithm to focus energy in the target direction. By adjusting the array weighting, it enhances the target echo response and suppresses sidelobe interference. The main lobe width is precisely set between 3 and 5 degrees as required. The filtered signal is then fed into the pulse compression module, where a linear frequency modulated pulse compression algorithm and a fast Fourier transform are used to process each signal, thereby improving range resolution. The output is high-resolution spatial perception imaging data, laying the foundation for deeper 3D reconstruction.

[0063] After receiving the aforementioned spatial perception imaging data, the 3D spatial modeling unit tensors and expands the signal data based on coordinates in multiple dimensions, including azimuth, pitch, distance, and polarization. This module utilizes a 3D convolutional neural network to perform convolutional processing on the multidimensional spatial data, achieving automated feature extraction and multi-target spatial structure modeling. Through end-to-end learning modeling, it can automatically separate multiple point cloud clusters, identify the physical properties of each spatial voxel, and output a 3D point cloud model along with feature data involving coordinates, dimensions, and physical properties, laying a standard data foundation for multi-source fusion.

[0064] The spatiotemporal registration and multimodal fusion unit takes into account the point cloud model and 3D feature data output by the 3D spatial modeling unit, while simultaneously invoking multi-source spatiotemporal attributes (such as GPS timing and AIS tracks). In terms of steps, a spatial transformation network is first used to align the coordinates of data from multiple stations, and dynamic frame-by-frame synchronization is achieved through timestamp normalization to establish a unified spatial and temporal reference. Further, other perceptual data, such as video streams and optical images, are integrated, and a feature-weighted integration method is used to fuse the 3D spatial structure with multimodal attributes, forming a unified multimodal spatial perceptual data volume. This output dataset integrates 3D coordinates, structural semantics, and temporal attributes, providing a highly redundant and attribute-rich original foundation for downstream ship detection, recognition, and intelligent analysis modules.

[0065] Example 3

[0066] This embodiment uses an automated port area in the coastal waters of location A as an application scenario. It describes the structural composition and main functional implementation of the visual feature layered fusion module in the intelligent ship detection system, as shown in the attached diagram. Figure 3 As shown, the visual feature hierarchical fusion module includes a hierarchical adaptive fusion structural feature encoder, a multi-level 3D feature processing strategy, and a mutual information enhancement collaborative feature learning mechanism. This module is mainly used to encode features from multiple visual sources and perform hierarchical fusion to improve the perception and recognition capabilities of ship targets in multi-target environments.

[0067] The hierarchical adaptive fusion structural feature encoder is responsible for receiving multimodal spatial perception data volumes output from the 3D perception modeling module. By constructing multi-layered feature channels, the encoder performs feature grouping on the input data at multiple levels, including spatial, scale, and semantic dimensions. The encoding process employs multi-resolution convolution, spatial attention mechanisms, and adaptive weight adjustment methods, enabling the encoder to flexibly adapt to changing lighting, reflections, and occlusion entities in aquatic scenes. It outputs hierarchical high-dimensional feature vectors, with all hierarchical features serving as inputs for downstream multi-level processing, ensuring a strict one-to-one correspondence between data streams.

[0068] The multi-level 3D feature processing strategy takes the hierarchical high-dimensional feature vectors generated by the hierarchical adaptive fusion structural feature encoder as input. The strategy utilizes cascaded 3D convolution, spatial feature map fusion, and inter-layer feature relationship reconstruction to progressively extract deep semantic information and spatial structural characteristics at each stage. Furthermore, by fusing physical attributes, it achieves differentiation between multiple target ships. During feature processing, the system employs cross-layer information flow and residual connection mechanisms to effectively suppress spatial information loss caused by the aliasing of multiple target distributions. Finally, it outputs a multi-scale 3D semantic feature set, which is provided to the subsequent mutual information enhancement module to ensure that the output features of each frame strictly correspond to the input structure.

[0069] The mutual information-enhanced collaborative feature learning mechanism takes the multi-scale 3D semantic feature set output by a multi-level 3D feature processing strategy as input, and combines it with front-end radar data, AIS-assisted information, and target historical motion trajectory information. An interactive attention mechanism is used for feature alignment and mutual information gain. This mechanism mainly establishes multi-dimensional collaborative associations between features, improving information sharing and semantic collaboration between features of different modalities and layers through a multi-branch Transformer structure and graph neural network model. During this process, the mutual information of each feature subspace is dynamically estimated, improving the spatial separability and global consistency of weak signal targets under complex conditions such as high occlusion, strong reflection, and clutter in the scene. The output is highly redundant 3D semantic feature data enhanced with mutual information, providing standard input for back-end target detection and intelligent recognition.

[0070] Example 4

[0071] This embodiment uses an automated port area in the coastal waters of location A as an application scenario to illustrate the target detection module in the intelligent ship detection system, combined with... Figure 4 The structure shown is described in detail, outlining its constituent units and functional implementation process. The target detection module includes an end-to-end ship interpretability detection and prediction model, a key feature annotation and hierarchical weighting unit, and a detection output and attribute annotation unit. These units collaborate to efficiently detect, interpret features, and output attributes for multiple target ships in the waterway.

[0072] The end-to-end ship interpretability detection and prediction model receives 3D semantic feature data from the visual feature hierarchical fusion module. This structure employs an end-to-end detection framework based on deep convolutional neural networks, incorporating an attention mechanism and a spatial feature enhancement module to extract global and local information from the input features. By introducing interpretability analysis methods, the system can explicitly label key regions and discrimination criteria in the detection results, thereby improving the transparency and traceability of the detection process. The output of this structure covers the spatial location, confidence score, and interpretability heatmap of the ship target, providing foundational data for subsequent feature annotation and attribute weighting.

[0073] The key feature annotation and hierarchical weighting unit takes the spatial location, confidence score, and interpretability heatmap of the ship target output by the end-to-end ship interpretability detection and prediction structure as input. This unit uses a multi-layer feature fusion algorithm to automatically identify key structural features of the ship target, such as the bow, stern, deck, and mast, and combines spatial distribution and semantic information to perform hierarchical weighting. A hierarchical feature weight allocation mechanism quantifies the criticality of features at different levels, highlighting the key feature regions that have the greatest impact on the detection results. The output is a ship target feature set with hierarchical weights and structural annotations, ensuring that each detected target has detailed structural explanations and feature distribution information.

[0074] The detection output and attribute labeling unit receives the ship target feature set output by the key feature labeling and hierarchical weighting unit. Based on a multi-attribute classification network, this unit performs attribute recognition and labeling for each detected target, encompassing ship type, size estimation, heading, speed, cargo status, etc. Relying on the joint output of attribute labels and spatial location, the system can generate a complete attribute description and spatial labeling information for each detected target. The final output is a structured ship detection result set, containing target spatial coordinates, key feature distribution, attribute labels, and confidence levels, providing standardized input for downstream behavior analysis and intelligent decision-making modules.

[0075] Example 5

[0076] In this embodiment, the anomaly detection module in the intelligent ship detection system is described in detail with reference to the structure shown in the attached figure. The anomaly detection module includes a spatiotemporal dynamic modeling unit, an abnormal behavior discrimination unit, and an abnormal data output and early warning unit. The various units work together to carry out dynamic modeling, intelligent discrimination, and real-time early warning of abnormal behavior of ships in the water.

[0077] The spatiotemporal dynamic modeling unit is responsible for receiving structured data such as the ship's spatial position, trajectory, and attribute labels output by the target detection module. This unit employs modeling methods based on spatiotemporal graph convolutional networks or long short-term memory networks to model the ship's dynamic characteristics, including its spatial distribution, speed changes, and heading adjustments over continuous time intervals. By jointly analyzing historical trajectories and the current state, it constructs a baseline model of the ship's spatiotemporal behavior. This unit outputs a sequence of spatiotemporal behavioral features for each target, providing input for the subsequent abnormal behavior discrimination unit.

[0078] The abnormal behavior discrimination unit takes the spatiotemporal behavior feature sequence output by the spatiotemporal dynamic modeling unit as input, and combines it with a pre-defined normal behavior pattern library. It employs clustering analysis, probabilistic statistics, or deep anomaly detection algorithms to discriminate ship behavior. This unit can identify various typical abnormal maritime behaviors, including abnormal course deviations, abnormal speed changes, illegal berthing, reverse navigation, and prolonged stays. During the discrimination process, the system calculates an anomaly probability score for each behavioral event and outputs the anomaly type, anomaly confidence level, and corresponding target identification information.

[0079] The abnormal data output and early warning unit receives the abnormality type, confidence level, and target identification information output by the abnormal behavior discrimination unit. This unit is responsible for structuring and outputting abnormal events in real time, and automatically triggering the early warning mechanism based on the abnormality level. The early warning mechanism includes various methods such as local audible and visual alarms, remote information push, and abnormal event log recording. The output content includes the spatial location of the abnormal target, the abnormality type, the occurrence time, and the confidence level, which facilitates subsequent manual intervention or automatic decision-making system processing.

[0080] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A ship intelligent detection system based on convolutional neural networks and radar signal processing, characterized in that, It includes the following modules: Ship 3D Perception Modeling Module: Acquires radar signals and spatiotemporal attribute data, reconstructs the ship's 3D spatial perception model through spatial perception imaging, and performs spatiotemporal alignment with spatiotemporal attribute data using spatiotemporal registration, and obtains multimodal spatial perception data through data fusion; The visual feature hierarchical fusion module employs a feature encoder that combines hierarchical adaptive fusion with mutual information enhancement to output spatial, frequency, and temporal visual features from multimodal spatial perception data and perform feature fusion. Through mutual information enhancement and redundancy suppression, multimodal fused visual features are obtained. Target detection module: Constructs an end-to-end interpretable ship detection and prediction model, performs spatial hierarchical weighting and annotation on multimodal fusion visual features, and outputs ship attribute interpretation and 3D spatial distribution prediction results; Anomaly Detection Module: Input the ship attribute interpretation and 3D spatial distribution prediction results, monitor the spatial behavior evolution process based on the visual spatiotemporal dynamic modeling mechanism, and output anomaly detection and behavior prediction information.

2. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The ship 3D perception and modeling module includes: Through a multi-channel radar signal synchronous acquisition unit, high-resolution radar echo data and spatiotemporal attribute data are collected for different meteorological and water surface conditions in the water environment. Through adaptive beamforming and pulse compression processing, spatial perception imaging is performed, and three-dimensional spatial perception imaging data is output.

3. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The ship's three-dimensional spatial perception model specifically includes: Inputting 3D spatial perception imaging data, and performing high-dimensional data unfolding based on azimuth, pitch, distance, and polarization spatial dimensions, generates multidimensional tensor data that combines spatial geometry and physical properties; based on the multidimensional tensor data, constructing a 3D spatial feature representation, extracting the 3D spatial distribution features of the ship target, and outputting a 3D spatial perception model of the ship, a 3D point cloud representation, a voxel feature representation, and a 3D feature dataset, wherein the 3D feature dataset includes the target's spatial location, volume, and structural type.

4. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The spatiotemporal registration specifically includes: Spatial coordinate registration is performed on 3D feature data groups from different time batches, and temporal synchronization is achieved through timestamp alignment. A unified spatial coordinate system and time axis reference are established. The ship's 3D spatial perception model, 3D point cloud representation, voxel feature representation, and 3D feature dataset are received through a multimodal data fusion method. Feature pairing is performed, and feature fusion is achieved through feature splicing and weighted averaging to obtain a 3D information body containing multidimensional visual semantics, spatial structural elements, and temporal change characteristics. Multimodal spatial perception data is then output.

5. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The visual feature hierarchical fusion module specifically includes: The hierarchical adaptive fusion and mutual information enhancement collaborative feature encoder receives multimodal spatial perception data and generates heterogeneous spatial feature groups in spatial, frequency, and temporal dimensions, outputting three types of 3D scene feature data: spatial structure information, frequency distribution information, and temporal variation information. The hierarchical adaptive fusion, through multi-level 3D feature processing, aggregates the three types of 3D scene feature data in layers according to spatial attributes, structural attributes, and dynamic attributes, extracts multi-level spatial information fusion features, and outputs 3D information fusion feature data containing 3D attribute layer labels, spatial distribution features, and scene semantics.

6. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 5, characterized in that, The mutual information enhancement collaboration specifically includes: The mutual information enhancement method receives hierarchically output 3D information fusion feature data through a collaborative feature learning unit, completes the correlation encoding between features of different modalities based on the mutual information measurement method, optimizes the feature information distribution through a redundancy suppression strategy, and outputs multimodal fusion visual feature data.

7. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The target detection module specifically includes: An end-to-end interpretable ship detection and prediction model is constructed. Multimodal fusion visual feature data is input, and feature weights are assigned to spatial feature levels. Feature annotations are performed based on spatial scale, spatial location, and semantic hierarchy. Combined with signal response distribution, and according to ship detection and prediction requirements, ship attribute features are annotated through weight distribution and causal response relationships to obtain multidimensional detection result data containing spatial distribution, key attributes, and prediction information. The multidimensional detection result data is then processed by an attribute interpretation unit to parse the attribute feature annotations, classify and label the ship's type, structural parameters, and current dynamic state, and output ship attribute interpretations. Through 3D coordinate mapping and spatial structure discrimination mechanisms, combined with the spatial distribution, key attributes, and prediction information in the multidimensional detection result data, the 3D spatial distribution prediction of each ship in the scene is performed, and the 3D spatial distribution prediction results are output.

8. The intelligent ship detection system based on convolutional neural networks and radar signal processing according to claim 1, characterized in that, The anomaly detection module specifically includes: Through the spatiotemporal dynamic modeling unit, the interpretation of ship attributes and the prediction results of three-dimensional spatial distribution are input. Based on the visual spatiotemporal dynamic modeling mechanism, the spatiotemporal evolution trajectory of the ship target in the water scene is obtained. Using a multi-time series data analysis method, behavioral sequence change features are extracted from the spatiotemporal evolution trajectory. Referring to the normal state model, a discrimination threshold is set, and target data that differs from the normal spatiotemporal evolution features are labeled. Target data with abnormal spatial location, motion state, and behavioral patterns are aggregated into abnormal behavior data. Through the abnormal feature analysis unit, features are summarized and classified according to abnormal attribute type, spatiotemporal distribution pattern, and behavioral evolution trend. An suppression scoring and risk classification mechanism is used for related abnormal targets to output abnormal detection information. Through the abnormal behavior change sequence, combined with the behavioral evolution law, the corresponding behavior prediction information is output.

Citation Information

Cited By

  • Multi-mode and reinforcement learning combined offshore overboard person searching method and system

    CN121459396A

  • Underwater target multi-sensor fusion magnetic field measurement system in complex magnetic environment

    CN121765643A

  • Underwater target multi-sensor fusion magnetic field measurement system in complex magnetic environment

    CN121765643B

  • Smart reservoir unattended operation and maintenance scheduling method based on unmanned ship

    CN122067200A