Dynamic target detection and tracking methods for unmanned surface vessel remote sensing images

By combining adaptive sensor fusion, deep learning, and multi-task networks, the problem of diversity and occlusion in dynamic target detection in unmanned surface vessel remote sensing images is solved, enabling real-time and accurate target detection and tracking in complex maritime environments and meeting the needs of autonomous navigation.

CN117218380BActive Publication Date: 2025-12-02CHINA WATER RESOURCES PEARL RIVER PLANNING SURVERYING & DESIGNING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311069337.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-23
Publication Date
2025-12-02
Estimated Expiration
2043-08-23

AI Technical Summary

Technical Problem

Dynamic target detection in unmanned surface vessel (USV) remote sensing images faces challenges in practical applications, including unstable data quality, insufficient diversity, target occlusion and deformation, target diversity and size variations, insufficient algorithm versatility, and difficulties in data source fusion. These issues make it difficult to meet the requirements for real-time performance and accuracy.

Method used

By employing adaptive sensor fusion technology, spatiotemporal fusion technology of deep learning, unsupervised anomaly detection technology, water surface reflection elimination technology, adaptive occlusion processing strategy, multi-task deep network, wave adaptive correction, edge computing and model compression technology, and unsupervised or semi-supervised model training technology, combined with various algorithms and technical means, we can achieve accurate detection and tracking of dynamic targets.

Benefits of technology

It improves the accuracy and stability of target detection and tracking in complex maritime environments, meets real-time requirements, reduces computing and storage requirements, lowers operational complexity, and automates autonomous navigation and target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218380B_ABST
    Figure CN117218380B_ABST
Patent Text Reader

Abstract

This invention relates to a dynamic target detection and tracking method for unmanned surface vessel (USV) remote sensing images. It includes the following steps: a) employing adaptive sensor fusion technology to integrate radar, optical, and infrared sensors for all-weather, multimodal target detection; b) using deep learning spatiotemporal fusion technology, employing 3D convolutional neural networks or LSTM to capture the spatiotemporal features of dynamic targets; c) using unsupervised anomaly detection technology to train models for common maritime scenes to identify and track anomalous targets; d) implementing water surface reflection reduction technology to reduce the impact of water surface reflection and solar flare; e) applying an adaptive occlusion handling strategy, using a predictive model to maintain target tracking even when the target is occluded by waves or other objects; f) utilizing a multi-task deep network for target detection, classification, velocity estimation, and orientation prediction; g) implementing wave adaptive correction, using machine learning to correct image distortions caused by water surface waves in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for dynamic target detection and tracking of remote sensing images from unmanned surface vessels. Background Technology

[0002] Currently, dynamic target detection using unmanned surface vessel (USV) remote sensing images faces several drawbacks and limitations in practical applications. Some common problems include: Insufficient data quality and diversity: USV remote sensing images are affected by natural factors such as weather and lighting, leading to unstable image quality, which can affect the accuracy of target detection and tracking. Furthermore, insufficient diversity of image samples collected from different regions and at different times may result in insufficient model generalization ability, making it difficult to adapt to different environmental changes. Target occlusion and deformation: Targets in USV remote sensing images may be occluded or deformed by other objects, water ripples, etc., causing changes in the target's appearance. This can make it difficult for target detection and tracking algorithms to accurately track the target's trajectory, especially when the target is partially occluded. Target diversity and scale variation: Targets in USV remote sensing images vary significantly in type and scale, potentially including small, low-contrast targets that are difficult to detect, or large targets that cause information overload within the area. Traditional target detection and tracking methods may struggle to handle this diversity and scale variation. In some application scenarios, such as maritime traffic control, real-time target detection and tracking are required. However, some efficient target detection and tracking methods may sacrifice some accuracy, while some accurate methods may have high computational complexity, failing to meet real-time requirements. Annotation of unmanned surface vessel (USV) remote sensing images is relatively complex, requiring manual labeling of target locations and categories, and due to image quality issues, labeling errors or inconsistencies may occur. This can affect the accuracy and stability of the trained model. Dynamic target detection and tracking methods developed for USV remote sensing images may be difficult to generalize to other fields because differences in image characteristics and application scenarios lead to insufficient algorithm versatility. USV remote sensing images typically need to be fused with other data sources (such as radar data, AIS data, etc.) to obtain more comprehensive target information. However, differences in data format and quality between different data sources can make the fusion process difficult, affecting the accuracy of the integrated information. Summary of the Invention

[0003] The purpose of this invention is to provide a method for dynamic target detection and tracking of unmanned surface vessel remote sensing images, thereby addressing some of the shortcomings and deficiencies pointed out in the background art.

[0004] The technical solution adopted by this invention to solve the above-mentioned technical problems includes the following steps:

[0005] a. Adaptive sensor fusion technology is used to integrate radar, optical and infrared sensors for all-weather, multi-modal target detection;

[0006] b. Using spatiotemporal fusion technology of deep learning, 3D convolutional neural networks or LSTM are used to capture the spatiotemporal features of dynamic targets;

[0007] c. Use unsupervised anomaly detection techniques to train models for common maritime scenarios in order to identify and track anomalous targets;

[0008] d. Implement water surface reflection elimination technology to reduce the impact of water surface reflection and solar scintillation;

[0009] e. Apply an adaptive occlusion handling strategy and use a predictive model to maintain target tracking, even if the target is occluded by waves or other objects.

[0010] f. Utilize multi-task deep networks for target detection, classification, velocity estimation, and orientation prediction;

[0011] g. Implement wave adaptive correction and use machine learning to correct image deformation caused by water surface waves in real time;

[0012] h. Adopt edge computing and model compression technologies to ensure real-time operation on the edge devices of unmanned vessels;

[0013] i. Use unsupervised or semi-supervised model training techniques, combining a small amount of labeled data with a large amount of unlabeled data.

[0014] Furthermore, the adaptive sensor fusion technology employs the following method: acquiring and integrating multimodal data from radar, optical, and infrared sensors; directly fusing the raw data from each sensor at the data layer to create a unified data framework; extracting features from the data of each sensor independently at the feature layer and integrating these features to provide comprehensive information for target recognition; and making decisions independently for each sensor at the decision layer and integrating these decisions to form a final target decision.

[0015] The system dynamically adjusts the weight of each sensor in the fusion process based on its performance and reliability in specific contexts; it achieves context awareness, automatically identifies and understands the current external environment, and adjusts the fusion strategy accordingly; it is equipped with a feedback mechanism that allows the system to automatically adjust its fusion strategy when the fusion decision is verified as an error; it utilizes advanced algorithms, including Kalman filters, particle filters, or neural networks, to process and integrate multimodal data; and it ensures that the system continuously updates its fusion model and strategy as new data is collected through real-time updates and online learning.

[0016] Furthermore, the spatiotemporal fusion technology utilizes a 3D convolutional network to simultaneously process the spatial and temporal dimensions of video data, where a three-dimensional convolutional kernel is used to capture dynamic information; it employs a dual-stream network structure, with one branch independently processing spatial information and the other branch processing temporal information, and the two streams are integrated in a subsequent stage; it integrates an LSTM or GRU structure, first extracting spatial features through a CNN, and then inputting these features into an LSTM or GRU to capture the continuity of the time series;

[0017] Adopting a spatiotemporal graph structure, this approach treats objects in the video as nodes and analyzes their relative positions and interactions in space and time using graph neural networks. It achieves feature fusion at different levels of the model, fusing dynamic features from low to high levels from the primary to the advanced layers. Combined with an attention mechanism, the model can focus on key spatiotemporal regions in the video during decision-making, assigning different weights to different parts of the content. A multi-scale fusion strategy is implemented, capturing and integrating features at different time scales to obtain richer spatiotemporal information.

[0018] Furthermore, the unsupervised anomaly detection technique employs: statistical methods, including Z-score and IQR, to evaluate the distribution of the data and identify data points that deviate from the mean; and integrates machine learning models, including One-Class SVM and Isolation Forest, to learn the normal boundaries of the data and identify data points that deviate from these boundaries; based on the above, density-based methods, including DBSCAN and LOF, are further employed to mark anomalies by comparing the density differences between data points and their neighbors.

[0019] When the above anomaly detection techniques fail to identify anomalies: Initiate integrated deep learning methods, especially autoencoders and deep class-I SVMs, to utilize neural network techniques to identify anomalous data points; or introduce time-series anomaly detection techniques, including LSTM and ARIMA; or adopt ensemble methods, including Feature Bagging and Isolation Anomaly Ensemble, to combine the strengths of multiple detection models for more accurate anomaly identification; or perform feature engineering, including using PCA, T-SNE, or UMAP techniques, to enhance data features and improve the accuracy of anomaly detection.

[0020] Furthermore, the water surface reflection elimination technology employs the following method:

[0021] When starting the method, polarization filtering is first used to filter out most of the scattered light from the water surface, providing a basic image for subsequent steps; on this basis, multi-wavelength and multi-angle observation techniques are used to capture images, ensuring that data obtained from multiple perspectives can fill in details that may have been lost due to the initial polarization filtering.

[0022] Next, deep learning, especially convolutional neural networks, was trained to recognize water surface reflection patterns that still exist in these underlying images; in order to provide training data for these networks, high dynamic range (HDR) imaging and multi-exposure techniques were used to capture images under various brightness conditions to ensure that the deep models are effective under various lighting conditions.

[0023] Meanwhile, spectral analysis is used to monitor reflection at specific wavelengths; combined with real-time environmental data, including wind speed and direction, the filtering and deep learning model parameters are dynamically adjusted according to specific environmental conditions; through adaptive local adjustments, the system performs detailed corrections in local areas, further ensuring the quality and accuracy of the image.

[0024] Furthermore, the implementation steps of the adaptive occlusion processing strategy are as follows:

[0025] S1. Dynamic background modeling technology is adopted to establish a stable background reference, laying the foundation for subsequent occlusion detection and target segmentation;

[0026] S2. Real-time occlusion detection is performed based on the established background model to identify foreground targets that do not match the background and potential occlusion situations;

[0027] S3. Evaluate the level of occlusion detected, distinguish between partially occluded and completely occluded targets, and select different processing strategies accordingly;

[0028] S4. For partially occluded targets, use a continuously updated appearance template for tracking;

[0029] S5. For a completely occluded target, use a prediction model based on historical trajectories to predict the target's movement trend after encountering occlusion;

[0030] S6. Combine the output of the prediction model and template matching to determine the most likely location of the target after occlusion;

[0031] S7. Through a feedback loop, the template is dynamically updated and the prediction model is adjusted based on the reappearance of the target or the disappearance of occlusion, so as to improve the accuracy of future occlusion processing.

[0032] Furthermore, the multi-task deep network method for target detection, classification, velocity estimation, and orientation prediction employs the following steps: performing data augmentation and normalization processing on the input image, including random cropping, rotation, scaling, color distortion, and uniform normalization of pixel values; and using multi-layer convolutional layers and depthwise separable convolutions to extract spatial features.

[0033] Based on the extracted features, a branching structure is used to perform the following tasks:

[0034] S1. Target detection using specific convolutional and upsampling layers;

[0035] S2. Target classification is performed using a fully connected layer and a softmax activation function;

[0036] S3. Use regression analysis to estimate the velocity of the detected target;

[0037] S4. Use angle regression to predict the target's direction of movement;

[0038] After completing the above tasks, various anchor box mechanisms are used to cover targets of different sizes and shapes, and positive and negative sample balance is achieved through hard negative sample mining. The losses of each task are integrated using a joint loss function, including cross-entropy loss and smoothing L1 loss. After obtaining the preliminary results, non-maximum suppression is performed to eliminate redundant detection boxes, and speed and direction are corrected based on historical data.

[0039] Furthermore, the process of performing wave adaptive correction on the input image affected by water surface fluctuations involves training the image using data collected under various wave conditions; capturing and correcting image deformations using a deep convolutional neural network or a deformable convolutional network; and selecting perceptual loss or structural similarity index as an optimized loss function to ensure the quality of the corrected image.

[0040] The edge computing and model compression technologies mentioned above are used to ensure real-time operation on the edge devices of unmanned vessels, and are combined with model pruning, model quantization, model distillation and specialized edge computing frameworks, including TensorRT, to optimize and deploy models.

[0041] The unsupervised or semi-supervised model training technique involves using an autoencoder to learn the internal structure of the data; then using generative adversarial networks (GANs) to assist training and generate synthetic samples; employing semi-supervised learning strategies, including MixMatch or Pseudo-Labeling, to combine a small amount of labeled data with a large amount of unlabeled data; performing data augmentation on the unlabeled data to generate new training samples and ensuring that the model's output remains consistent across different augmented versions; and using a self-learning strategy to use the model's output as pseudo-labels for the unlabeled data for further training.

[0042] Furthermore, the specific implementation process of using statistical methods and ensemble machine learning models is as follows: First, calculate the Z-score of the target attributes captured from the remote sensing image, such as pixel intensity, size, shape, and motion speed, using the formula: Furthermore, its interquartile range (IQR) is calculated using the formula IQR = Q3 - Q1;

[0043] If the absolute value of the obtained Z-score exceeds a predetermined threshold or the target attribute exceeds the boundary calculated based on IQR, the target is marked as a preliminary anomaly.

[0044] Secondly, these identified target features are input into a pre-trained One-Class SVM or IsolationForest model. Both models have been trained on historical data to learn the normal distribution of the data and identify outlier data points that deviate from this distribution.

[0045] Model judgment based on One-Class SVM relies on the following formula: The weights w, slack variables ξ, and the offset ρ of the decision function have all been determined during the previous model training.

[0046] If a target is flagged as an anomaly by a machine learning model, it is further labeled and analyzed to determine its nature;

[0047] Furthermore, the specific process of feature engineering using the PCA, T-SNE, and UMAP technologies is as follows:

[0048] S1. Using PCA, first calculate the covariance matrix of the data, then find the eigenvectors and eigenvalues ​​of the covariance matrix, and sort them according to the size of the eigenvalues; select the first k principal components for data dimensionality reduction.

[0049] PCA(X)=X×eigenvectors(Cov(X))

[0050] Where X is the original data, and Cov(X) is the covariance matrix of X;

[0051] S2. Using T-SNE technology, based on the similarity between two points in the original space and the low-dimensional space, the Kullback-Leibler divergence between them is minimized to reduce the dimensionality of the data;

[0052]

[0053] Where, p ij and q ij These represent the similarity of two points in the original space and the lower-dimensional space, respectively.

[0054] S3. Using UMAP technology, a topological representation of the original space is constructed, and then the best representation in the low-dimensional space is found; UMAP aims to minimize the change in distance and preserve the topological structure of the data;

[0055]

[0056] Where f is the mapping function, w ijIt is the weight, which reflects the similarity between point i and point j.

[0057] The beneficial effects of this invention are as follows: By combining multiple technologies and algorithms, such as feature engineering, machine learning models, and adaptive sensor fusion, this method can more accurately detect and track dynamic targets, even in complex maritime environments. This method exhibits excellent robustness to common interference factors such as environmental noise, water surface reflection, solar flare, and occlusion, ensuring stability and reliability under various complex conditions. Through edge computing and model compression techniques, this method can run in real-time on the edge devices of unmanned surface vessels, meeting the needs of real-time navigation and monitoring. Utilizing model compression and optimization techniques, this method reduces the demand for computing resources and storage while maintaining high accuracy, extending the operational time of unmanned surface vessels.

[0058] By combining adaptive techniques and machine learning, this method can automatically adjust to different marine environments and target types, meeting the needs of various application scenarios. By integrating multiple anomaly detection technologies, such as Z-score, IQR, and machine learning models, this method reduces false alarms and false negatives, improving the overall system performance. The application of this method reduces the need for human intervention, making autonomous navigation and target tracking of unmanned vessels more automated, and reducing operational complexity and risk. Attached Figure Description

[0059] Figure 1 This is a flowchart of the dynamic target detection and tracking method for unmanned surface vessel remote sensing images according to the present invention. Detailed Implementation

[0060] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0061] Example: A dynamic target detection and tracking method for unmanned surface vessel (USV) remote sensing images combines multiple advanced sensors, algorithms, and computing technologies to achieve accurate and real-time detection and tracking of dynamic targets at sea. First, adaptive sensor fusion technology is employed, integrating radar, optical, and infrared sensors to ensure effective target detection under varying weather and lighting conditions. Second, spatiotemporal fusion techniques using deep learning, such as 3D convolutional neural networks or LSTM, can capture the spatiotemporal features of the target. LSTM, in particular, has the following formula: f t =σ(W f ·[h t-1 ,x t ]+b f ), where f t Represents the forget gate, h t-1 It is the hidden state of the previous time step, x t σ is the current input, and σ is the Sigmoid activation function.

[0062] To handle common maritime scenarios, this method employs unsupervised anomaly detection techniques for model training, enabling the effective detection and tracking of even uncommon targets. Furthermore, by implementing water surface reflection reduction techniques, the effects of water surface reflection and solar flare are reduced, improving image quality and target detection accuracy.

[0063] Considering that targets may be obscured by waves or other objects in a marine environment, this method employs an adaptive occlusion handling strategy and uses a predictive model to maintain target tracking, ensuring continuous target tracking even when occlusion occurs. Furthermore, the multi-task deep network not only detects targets but also classifies them, estimates their velocity, and predicts their direction of motion.

[0064] Furthermore, since water surface fluctuations can cause image distortion, this method also implements wave-adaptive correction, using machine learning to correct the image in real time to ensure image stability. To meet the requirements of real-time processing, this method also adopts edge computing and model compression techniques to ensure that the algorithm can run in real time on the edge devices of the unmanned vessel.

[0065] Finally, considering that acquiring labeled data is usually costly and time-consuming, this method also incorporates unsupervised or semi-supervised model training techniques, so that effective model training and optimization can be carried out even with only a small amount of labeled data and a large amount of unlabeled data.

[0066] Example 1: The unmanned vessel in this case is navigating in a certain sea area with the goal of tracking and identifying a distant dynamic object, which may be a ship, a floating object or other maritime target.

[0067] S1. Multimodal data acquisition and integration:

[0068] The radar sensor detected the target at a distance of 3,000 meters and a speed of 20 knots.

[0069] An optical sensor captures an image of the target, and image analysis leads to the conclusion that the target may be a speedboat.

[0070] Infrared sensors captured the target's heat distribution, indicating that the target's engine was operating and had a moderate heat output.

[0071] S2. Data layer fusion: Integrating radar, optical, and infrared data into a unified data framework to obtain comprehensive target information.

[0072] S3. Feature Layer Fusion:

[0073] Extract speed and range features from radar data.

[0074] The morphological characteristics of a target, such as its size, color, and shape, can be obtained from optical data.

[0075] The target's temperature characteristics, such as engine operating status and heat output, are obtained from infrared data.

[0076] S4. Integration of Decision-Making Levels:

[0077] The radar data indicates that the target is moving.

[0078] The target was determined to be a ship based on optical data.

[0079] Based on infrared data, it was determined that the target's engine was operational.

[0080] Final decision: The target is a ship that is moving.

[0081] S5. Weight Adjustment: Assuming it is a cloudy day, which may affect the recognition performance of the optical sensor, the weights of the radar and infrared sensors are relatively increased.

[0082] S6. Context Awareness: Because it is a cloudy day, the system knows that the optical sensor may not be so reliable, so it automatically adjusts the strategy according to the current environmental context.

[0083] S7. Feedback Mechanism: If the system subsequently discovers an error in its decision-making, such as identifying a speedboat as a fishing boat when it is actually a fishing boat, the system will automatically adjust its strategy to improve the accuracy of the next decision.

[0084] S8. Advanced Algorithm Applications:

[0085] Kalman filters are used to process radar data to more accurately estimate the target's actual position and velocity.

[0086] By utilizing neural networks and combining optical and infrared data, the accuracy of target type identification can be further improved.

[0087] S9. Real-time updates and online learning: When a new similar target enters the sensor's detection range, the system will use previous data and experience to make a faster and more accurate judgment on the new target.

[0088] Example 2: The mission of an unmanned vessel is to monitor and analyze the movement trajectory of fish schools in a complex marine environment. The size, speed and direction of the fish schools are changing, and fluctuations in the sea surface and other large marine organisms may interfere with the observation.

[0089] S1.3D convolutional network processing:

[0090] This embodiment has a video sequence containing 100 frames, each frame being 3840x2160 pixels in size.

[0091] The three-dimensional convolution kernel is set to [3,3,3], which means that it will process three consecutive frames of image data at the same time and perform convolution within a 3x3 pixel area in space.

[0092] This means that, in addition to the current location of the captured fish, this embodiment can also obtain their short-term movement trends.

[0093] S2. Two-stream network structure:

[0094] The spatial branch uses 2D convolution to extract the morphological features of the fish in each frame.

[0095] The time branching process handles the differences between consecutive frames and can identify the instantaneous movement characteristics of the fish.

[0096] The outputs of the two branches will be merged in a fully connected layer to form a complete description of the fish.

[0097] S3. Integrated LSTM or GRU architecture:

[0098] In the captured video sequences, schools of fish may suddenly change direction as they search for food or flee from predators.

[0099] The LSTM structure can capture these sudden changes in behavior because it remembers the previous state and combines it with the current input.

[0100] At each time step, the LSTM receives the features extracted by the CNN from the current frame and combines them with the previous internal state to make a decision.

[0101] S4. Spatiotemporal diagram structure:

[0102] Assuming that 100 fish are observed in a certain frame in this embodiment, this embodiment can construct a graph structure with 100 nodes.

[0103] If two fish are close enough, create an edge between them. The weight of the edge can be based on their distance, speed difference, and direction difference.

[0104] The graph neural network will perform iterative calculations based on this to identify the core members and peripheral members of the fish group.

[0105] S5. Feature-level fusion:

[0106] At a basic level, features such as the fish's scales, eyes, and tail may be captured.

[0107] As the network expands, it becomes possible to capture more macroscopic features, such as the overall shape, density, and movement trends of fish schools.

[0108] S6. Attention Mechanism:

[0109] In a complex ocean context, this embodiment aims to enable the model to automatically focus on those parts most relevant to the task, such as the leader fish in a school of fish.

[0110] Attention weights will be applied to these key areas, making the model more accurate in making decisions.

[0111] S7. Multi-scale fusion strategy:

[0112] On a short timescale, this embodiment may focus on minute changes within the school of fish, such as a fish suddenly accelerating or turning.

[0113] On a long-term scale, this embodiment focuses on the macroscopic movement of fish schools, such as the migration routes they follow.

[0114] These two types of information are combined to help this embodiment to more comprehensively understand the behavioral characteristics of the fish school.

[0115] Example 3: An unmanned surface vessel is collecting physical property data of seawater in a certain sea area, such as temperature, salinity, and current velocity. This data changes over time, and this example requires the detection of potential anomalies.

[0116] Sample data:

[0117] Data from sea area A over 100 consecutive days showed an average temperature of 15℃, a salinity of 35‰, and a current velocity of 0.9m / s.

[0118] However, on the 101st day, the data suddenly changed to: temperature 45℃, salinity 70‰, and flow rate 5.5m / s.

[0119] S1. Statistical Methods:

[0120] When using the Z-score method: For temperature: Z = (45-15) / standard deviation. If the Z-value is much greater than 2 or much less than 2, then the day may be abnormal. Similarly, this embodiment can also calculate the Z-score for salinity and flow rate.

[0121] When using the IQR method: the difference between the first quartile (Q1) and the third quartile (Q3) is the IQR. For temperature: if 45°C exceeds Q3 + 1.5IQR or falls below Q1 - 1.5IQR, it is considered an anomaly.

[0122] S2. Machine Learning Models: OneClassSVM: This model attempts to find the normal boundaries of the data. After training on 100 days of data, if a data point on the 101st day falls outside this boundary, it is marked as an anomaly. IsolationForest: This is a tree-structured model designed to "isolate" outliers. More anomalous data points can be isolated earlier in the decision tree.

[0123] S3. Density-based methods: DBSCAN: This method defines anomalies by finding regions of low density. Because the density difference between the data point on day 101 and other points is significant, it will be considered an anomaly. LOF (Local Outlier Factor): This measures the density deviation of a data point from its neighbors. A larger LOF value indicates a greater degree of anomaly.

[0124] S4. Deep Learning Methods: Autoencoder: This is a neural network structure that attempts to reconstruct its input. It is trained on normal data, but the reconstruction error increases significantly when encountering anomalous data. Deep OneClassSVM: Combines deep learning feature extraction with OneClassSVM to achieve deeper anomaly detection.

[0125] S5. Time Series Anomaly Detection: LSTM (Long Short-Term Memory) networks can capture patterns and associations in time series data. If the predicted value on day 101 differs significantly from the actual value, it may be an anomaly. ARIMA (Arithmetic and Randomization Analysis) is a statistical model used for time series forecasting. Similar to LSTM, it can also be used to identify anomalies.

[0126] S6. Ensemble methods, such as FeatureBagging, use multiple sub-models (e.g., multiple IsolationForests) and then average or vote on their outputs to obtain more robust anomaly detection. IsolationAnomalyEnsemble, on the other hand, combines the strengths of multiple anomaly detection models to generate a comprehensive anomaly score for each data point.

[0127] S7. Feature engineering, PCA, and principal component analysis can reduce the dimensionality of data and highlight major outlier patterns; while TSNE & UMAP: both of these techniques are dimensionality reduction techniques used for data visualization. They maintain the distance between data points in high-dimensional space, and outliers are often more apparent in low-dimensional projections.

[0128] Example 4: An unmanned vessel is patrolling a vast sea area and needs to accurately detect and track dynamic targets, such as small boats, floating objects or other unknown targets, under varying light and sea conditions.

[0129] When unmanned surface vessels (USVs) patrol vast ocean areas, accurate dynamic target detection and tracking are crucial in the face of ever-changing lighting and sea conditions. In this mission, the polarizing filter on the camera plays a key role, effectively filtering out most of the scattered light and laying a solid foundation for subsequent steps. To further improve image clarity, the USV achieves multi-angle observation through a rotating camera array, and utilizes cameras operating within different wavelength ranges to ensure high-quality images are acquired under various lighting and weather conditions.

[0130] The acquired images are fed into a pre-trained convolutional neural network (CNN) for preliminary analysis to identify potential dynamic targets. Simultaneously, a spectral analysis module extracts spectral features from the images, creating a unique spectral "fingerprint" for each target detected by the CNN. When these "fingerprints" match known spectral data, this embodiment can further determine the nature and type of the target.

[0131] To improve detection accuracy, the unmanned surface vessel is also equipped with other sensors, such as weather sensors and magnetometers, which provide real-time data on wind speed, wind direction, sea conditions, and more. This environmental data, combined with the output of the CNN and spectral analysis results, allows the system to provide more accurate target predictions. For example, by understanding wind speed and direction, the algorithm can predict the movement trend of floating objects, while understanding lighting conditions helps the system fine-tune image processing parameters, ensuring images remain clear in any environment.

[0132] Once the target is identified, the system automatically activates adaptive local adjustment, magnifying and enhancing the image to ensure every detail of the target is clearly visible. Furthermore, a dynamic target tracking algorithm is activated, adjusting the camera's focus and angle to ensure the target remains centered in the field of view regardless of its movement. In addition, the unmanned surface vessel continuously optimizes its tracking strategy based on spectral analysis and real-time environmental data, ensuring the target is never lost under any circumstances.

[0133] Example 5: Adaptive occlusion handling strategy is a difficult problem in dynamic target tracking; this example discusses the implementation and technical details of each step in detail.

[0134] S1: Dynamic Background Modeling

[0135] In this step, a technique such as a Gaussian Mixture Model (GMM) is used to build a background model from consecutive video frames. This model captures stable parts of the scene, such as walls, buildings, and roads, while ignoring moving objects such as cars and pedestrians. The stability of the background model is central to the entire strategy because it provides a reference standard for occlusion detection.

[0136] S2: Real-time occlusion detection. Whenever a new video frame arrives, it is compared with the background model. Regions that do not conform to the background model are marked as foreground or potential occlusion areas. This is typically achieved through frame differencing or background subtraction.

[0137] S3: Evaluate the occlusion level. The identified foreground region is further analyzed. Features such as area, shape, and texture are used to evaluate the degree of occlusion. For example, if a known object overlaps with another object by only a small portion, it will be considered "partially occluded."

[0138] S4: Tracking partially occluded targets. The core here is the appearance template, a visual representation of the target that is dynamically updated over time. Template matching technology searches for regions similar to the template in each new video frame, maintaining tracking even when the target is partially occluded.

[0139] S5: Handling completely occluded targets. When a target is completely occluded, the strategy shifts to historical trajectory data. For example, if a pedestrian was previously walking north, they may still be moving in that direction. Based on this, techniques such as Kalman filtering or particle filtering are used to predict the target's future position.

[0140] S6: Determine the most probable location. This is a data fusion step that combines the outputs of S4 and S5. The location given by the predictive model and the location identified by template matching will be considered together. The decision process may give higher weight to the predictive model, especially when there is severe occlusion.

[0141] S7: Feedback Loop, a crucial step in ensuring policy adaptability. When the target reappears or the occlusion disappears, the appearance template is updated, and the parameters of the prediction model are fine-tuned accordingly to ensure the system can react more accurately to similar occlusion situations in the future.

[0142] Example 6: An unmanned surface vessel is conducting a remote sensing mission in a complex marine environment.

[0143] Data augmentation and normalization:

[0144] In-depth analysis: In the complex marine environment of unmanned vessels, lighting conditions, the effects of ripples, nearby targets such as buoys or other vessels, and distant targets such as port facilities or coastlines all have different impacts on images. Therefore, random cropping can simulate different field of view, rotation can simulate the ship's swaying, and color distortion can simulate different weather and time conditions. All of these contribute to greater robustness of the model.

[0145] Spatial feature extraction:

[0146] In-depth analysis: By utilizing multiple convolutional layers, this embodiment can first capture low-level features, such as ripples on the sea surface or the edges of a ship. As the network deepens, this embodiment can acquire high-level features, such as the overall structure of a ship or specific patterns on a buoy. These features then provide the foundation for a multi-task network.

[0147] Object detection:

[0148] In-depth analysis: In unmanned surface vessel (USV) scenarios, targets can be static or dynamic. For static targets (such as buoys or anchored vessels), the model needs to determine whether they blend into the background (sea surface). For dynamic targets (such as moving vessels), the model also needs to capture their velocity and direction, which lays the foundation for velocity estimation and direction prediction.

[0149] Target Classification:

[0150] In-depth analysis: After detecting a target, the model needs to determine what it is. Why is this important? Because different targets may require different obstacle avoidance strategies. For example, for fast-moving dynamic targets (such as speedboats), unmanned surface vessels may need to plan their paths in advance.

[0151] Velocity estimation and direction prediction:

[0152] In-depth analysis: These two tasks are interdependent. Once a target is detected, its dynamic characteristics need to be assessed. For example, if a large cargo ship is increasing its speed and heading towards the unmanned vessel, this could indicate a potential collision risk.

[0153] Anchor frame mechanism and positive / negative sample balance:

[0154] In-depth analysis: Unmanned surface vessels (USVs) need to identify targets of various sizes and shapes, from large distant ships to small nearby buoys. The anchor-frame mechanism ensures this diversity, while hard-to-bearer sample mining ensures that the model can handle situations that are easily overlooked or misidentified during training.

[0155] Joint loss function:

[0156] In-depth analysis: Multi-task learning requires optimizing multiple objectives, but the optimization of each objective should not hinder the optimization of others. This embodiment ensures that all tasks are appropriately considered through a joint loss function.

[0157] Nonmaximum suppression and correction:

[0158] In-depth analysis: Unmanned surface vessels operate in fast-paced, dynamic environments, thus requiring rapid and accurate target identification. Non-maximum suppression ensures that this embodiment focuses only on the most relevant targets, while corrections based on historical data improve the model's real-time performance.

[0159] Example 7: In remote sensing applications of unmanned surface vessels (USVs), achieving effective target detection and tracking requires a series of complex steps addressing specific problems and technologies. To better cope with the effects of sea surface reflection and waves, a wave-adaptive correction method is first employed. By collecting data under various wave conditions, a deep convolutional neural network or deformable convolutional network learns image distortion and corrects it. Let the original image be I. original The deformed image is I distorted The goal of a neural network is to find a transformation function F such that F(I) distorted )≈I original To ensure the quality of the corrected image, the structural similarity index SSIM is used as the loss function, which is defined as:

[0160] Where, μ x and μ y σ is the mean of the two images. xy These are the covariances, and c1 and c2 are constants.

[0161] To ensure the real-time performance of edge devices on unmanned surface vessels, edge computing and model compression techniques were adopted. The model weights (W) can be reduced to a smaller set (Wi) through pruning strategies, such as threshold pruning. ′ In model quantization, weight values ​​are mapped from a continuous space to a discrete space, such as quantization to 8-bit integers: W q =round(W / q) max *255) where q max It is the maximum value of the weight.

[0162] Finally, unsupervised or semi-supervised model training techniques allow models to achieve good performance even with limited labeled data. Using autoencoders to learn the internal structure of the data and generating synthetic samples through generative adversarial networks (GANs) enables the model to learn more robust features. In semi-supervised learning strategies, such as MixMatch, labeled data D_l and unlabeled data (D_l) are utilized. u A new sample set (D) is obtained through data augmentation. aug Then, the model is used for training. Data augmentation can be viewed as a function A:D aug =A(D) l D u The model's output is used as a pseudo-label, and the model is further improved through a series of self-learning strategies.

[0163] For example, suppose that in a remote sensing mission, an unmanned surface vessel (USV) captures an image under conditions of large fluctuations, and the target in the image is distorted. Using the aforementioned method, a neural network corrects this distortion and performs target detection, velocity estimation, and orientation prediction in real time on the USV's edge device.

[0164] Assuming this embodiment captures an image (I) from the sea surface distorted ) and the original undisturbed image (I original The structural similarity index (SSIM) between them is 0.7 (the value range is [-1,1], and the closer the value is to 1, the better).

[0165] Calculate SSIM using the formula above:

[0166] SSIM(I original ,I distorted )

[0167] If (μ) x =100,μ y =95),(σ xy Substituting (c1 = 0.01, c2 = 0.03) into the above formula, this embodiment can obtain the SSIM value. If this value is close to 0.7, it indicates that the model has achieved good results in correcting fluctuations.

[0168] Suppose this embodiment has a model with weights of [0.5, -0.2, 0.01, 0.05]. To simplify the problem, this embodiment chooses 0.05 as the pruning threshold. Then, all weights with absolute values ​​less than 0.05 will be set to 0, resulting in a new weight set of 0.5, -0.2, 0, 0.

[0169] For model quantization, assuming the maximum weight value is 0.5, this embodiment can use formula W. q =round(W / q) max The weights are quantized using round(0.5 / 0.5*255) = 255. Therefore, a weight of 0.5 will be mapped to round(0.5 / 0.5*255) = 255.

[0170] Assuming this embodiment has 10 labeled data points (D l ) and 100 unlabeled data points (D u During the data augmentation phase, this embodiment randomly selects data from (D) l ) and (D u Five points were selected from the dataset, and random cropping and rotation were applied to generate 10 new sample points, resulting in the dataset (D). aug ).

[0171] Example 8: The actual process of using statistical methods and ensemble machine learning models lies in the calculation of the correlation function. The following example will be explained with actual numerical values:

[0172] Calculation of Z-score and IQR:

[0173] The Z-score is the difference between a data point and its mean, relative to the standard deviation. It can reflect whether the data point is an outlier. Assume the sample values ​​for the target attribute (e.g., pixel intensity) in the example are: 50, 55, 53, 49, 400, where 400 might be an anomalous pixel intensity.

[0174] average value

[0175] Standard deviation:

[0176] For a pixel intensity of 400, its Z-score is:

[0177]

[0178] To calculate the IQR, the sample data must first be sorted and the first quartile (Q1) and third quartile (Q3) must be found. For the sample data in this example, after sorting, the values ​​are 49, 50, 53, 55, 400. Therefore, Q1 = 50 and Q3 = 55, so the IQR is:

[0179] IQR = 55 - 50 = 5

[0180] Using OneClassSVM for anomaly detection:

[0181] The goal of OneClassSVM is to find the hyperplane that best approximates the majority of data points and to treat data points deviating from this hyperplane as outliers. In the formula, ρho is the offset of the decision function, xii is the slack variable representing the distance of the i-th data point to the correct class, and w is the normal vector of the hyperplane. Essentially, the formula balances two objectives: making the hyperplane as close as possible to the majority of data points while allowing some points to deviate from it.

[0182] Suppose that the example obtained the following parameter values ​​during training on five data points: ρ = 0.5, w = [0.3, 0.4], ξ = [0.1, 0.2, 0.1, 0.1, 0.3] (this is for simplification; in practice, the dimensions of these parameters are related to the dimensions of the features). For ν = 0.2 (representing the expected proportion of outliers) and n = 5, the example can calculate the above formula:

[0183]

[0184] The goal of the above optimization problem is to minimize this value in order to determine the optimal w and ρ.

[0185] Ultimately, if a target is considered an anomaly on one side of the decision plane, it will be further labeled and analyzed. This strategy can help unmanned surface vessels better identify anomalies in remote sensing images, enabling appropriate countermeasures.

[0186] Example 9: An unmanned surface vessel (USV) uses sensors to collect a series of multidimensional data, such as the vessel's speed, direction, water temperature, and wind speed. To effectively analyze and visualize this data, this example requires dimensionality reduction.

[0187] 1. PCA (Principal Component Analysis):

[0188] PCA can help implementations identify key trends in data, such as the movement patterns of unmanned vessels.

[0189] Example: Suppose the implementation plan collected the following two-dimensional data, representing ship speed and wind speed, respectively:

[0190]

[0191] After calculating the mean, the data points are centered. Then, the covariance matrix of Cov(X) is calculated. Its eigenvalues ​​and eigenvectors are calculated and sorted. These vectors indicate which directions of ship speed and wind speed changes are most significant.

[0192] 2.TSNE (t-distributed stochastic neighbor embedding):

[0193] When analyzing multidimensional sensor data from unmanned surface vessels, t-SNE can help implementations effectively visualize this data in two-dimensional space while preserving the structure between the data.

[0194] Example: Suppose that in the original space, the similarity (p) between the data from two sensors is... ij =0.8), while in the dimensionality-reduced space, the similarity (q) ij =0.5). Using the formula above, the example yields:

[0195]

[0196] The above can help the embodiments better understand the similarities between sensor data.

[0197] 3.UMAP (Uniform Manifold Approximation and Projection):

[0198] UMAP can help implementations reduce the dimensionality of unmanned surface vessel data more quickly while preserving its topology.

[0199] Assuming that in a specific sensor reading, the embodiment has a mapping function value f(u) for two high-dimensional data points. i ) = 0.6 and f(u j The weight w between the two points is 0.4. ij =0.7. Then calculate:

[0200]

[0201] In this way, the implementation can preserve the important structure in the sensor data.

[0202] In summary, by utilizing dimensionality reduction techniques, the implementation examples can effectively analyze and visualize sensor data from unmanned surface vessels (USVs) to better understand their motion and changes in the surrounding environment.

[0203] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for dynamic target detection and tracking in unmanned surface vessel remote sensing images, the method comprising the following steps: a. Adaptive sensor fusion technology is used to integrate radar, optical and infrared sensors for all-weather, multi-modal target detection; b. Using spatiotemporal fusion technology of deep learning, 3D convolutional neural networks or LSTM are used to capture the spatiotemporal features of dynamic targets; c. Use unsupervised anomaly detection techniques to train models for common maritime scenarios in order to identify and track anomalous targets; d. Implement water surface reflection elimination technology to reduce the impact of water surface reflection and solar scintillation; e. Apply an adaptive occlusion handling strategy and use a predictive model to maintain target tracking, even if the target is occluded by waves or other objects. f. Utilize multi-task deep networks for target detection, classification, velocity estimation, and orientation prediction; g. Implement wave adaptive correction and use machine learning to correct image deformation caused by water surface waves in real time; h. Adopt edge computing and model compression technologies to ensure real-time operation on the edge devices of unmanned vessels; i. Use unsupervised or semi-supervised model training techniques, combining a small amount of labeled data with a large amount of unlabeled data; The adaptive sensor fusion technology described above employs the following method: acquiring and integrating multimodal data from radar, optical, and infrared sensors; The raw data from various sensors are directly fused at the data layer to create a unified data framework; At the feature layer, features are extracted from the data of each sensor independently and these features are integrated to provide comprehensive information for target recognition; at the decision layer, decisions are made independently for each sensor and these decisions are integrated to form a final target decision. The weight of each sensor in the fusion process is dynamically adjusted based on its performance and reliability in a specific context. It achieves context awareness, automatically identifies and understands the current external environment, and adjusts the fusion strategy accordingly; it is equipped with a feedback mechanism that allows the system to automatically adjust its fusion strategy when the fusion decision is verified as an error. Advanced algorithms, including Kalman filters, particle filters, or neural networks, are used to process and integrate multimodal data; real-time updates and online learning ensure that the system continuously updates its fusion model and strategy as new data is collected. The spatiotemporal fusion technology described above utilizes a 3D convolutional network to simultaneously process the spatial and temporal dimensions of video data, employing 3D convolutional kernels to capture dynamic information; it adopts a dual-stream network structure, with one branch independently processing spatial information and the other processing temporal information, and the two streams are integrated in a subsequent stage; it integrates an LSTM or GRU structure, first extracting spatial features through a CNN, and then inputting these features into an LSTM or GRU to capture the continuity of the time series; It adopts a spatiotemporal graph structure, treats objects in the video as nodes, and analyzes the relative positions and interactions of objects in space and time through graph neural networks; it achieves feature fusion at different levels of the model, fusing dynamic features from low level to high level from the primary layer to the advanced layer respectively; By incorporating an attention mechanism, the model can focus on key spatiotemporal regions in the video when making decisions, and assign different weights to each part of the content. Implement a multi-scale fusion strategy to capture and integrate features at different time scales to obtain richer spatiotemporal information.

2. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 1, characterized in that... The unsupervised anomaly detection technique employs the following methods: statistical methods, including Z-score and IQR, to evaluate the distribution of the data and identify data points that deviate from the mean; and integrating machine learning models, including One-Class SVM and IsolationForest, to learn the normal boundaries of the data and identify data points that deviate from these boundaries; further, density-based methods, including DBSCAN and LOF, are used to mark anomalies by comparing the density differences between data points and their neighbors. When the above anomaly detection techniques fail to identify anomalies: an integrated deep learning method is initiated, which is an autoencoder and a deep class-1 SVM, using neural network technology to identify anomalous data points; Alternatively, time-series anomaly detection techniques, including LSTM and ARIMA, can be introduced; or ensemble methods, including Feature Bagging and Isolation Anomaly Ensemble, can be adopted to combine the strengths of multiple detection models for more accurate anomaly identification; or feature engineering can be performed, including using PCA, T-SNE, or UMAP techniques, to enhance data features and improve the accuracy of anomaly detection.

3. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 1, characterized in that... The water surface reflection elimination technology employs the following method: When starting the method, polarization filtering is first used to filter out most of the scattered light from the water surface, providing a basic image for subsequent steps; on this basis, multi-wavelength and multi-angle observation techniques are used to capture images, ensuring that data obtained from multiple perspectives can fill in details that may have been lost due to the initial polarization filtering. Next, deep learning, specifically convolutional neural networks, is trained to recognize water surface reflection patterns that still exist in these underlying images. To provide training data for these networks, high dynamic range (HDR) imaging and multi-exposure techniques are used to capture images under various brightness conditions, ensuring that the deep models are effective under all lighting conditions. Meanwhile, spectral analysis is used to monitor reflection at specific wavelengths; combined with real-time environmental data, including wind speed and direction, the filtering and deep learning model parameters are dynamically adjusted according to specific environmental conditions; through adaptive local adjustments, the system performs detailed corrections in local areas, further ensuring the quality and accuracy of the image.

4. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 1, characterized in that... The implementation steps of the adaptive occlusion handling strategy are as follows: S1. Dynamic background modeling technology is adopted to establish a stable background reference, laying the foundation for subsequent occlusion detection and target segmentation; S2. Real-time occlusion detection is performed based on the established background model to identify foreground targets that do not match the background and potential occlusion situations; S3. Evaluate the level of occlusion detected, distinguish between partially occluded and completely occluded targets, and select different processing strategies accordingly; S4. For partially occluded targets, use a continuously updated appearance template for tracking; S5. For a completely occluded target, use a prediction model based on historical trajectories to predict the target's movement trend after encountering occlusion; S6. Combine the output of the prediction model and template matching to determine the most likely location of the target after occlusion; S7. Through a feedback loop, the template is dynamically updated and the prediction model is adjusted based on the reappearance of the target or the disappearance of occlusion, so as to improve the accuracy of future occlusion processing.

5. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 1, characterized in that... The multi-task deep network method for target detection, classification, velocity estimation, and orientation prediction employs the following steps: performing data augmentation and normalization on the input image, including random cropping, rotation, scaling, color distortion, and unified normalization of pixel values; and extracting spatial features using multi-layer convolutional layers and depthwise separable convolutions. Based on the extracted features, a branching structure is used to perform the following tasks: S1. Target detection using specific convolutional and upsampling layers; S2. Target classification is performed using a fully connected layer and a softmax activation function; S3. Use regression analysis to estimate the velocity of the detected target; S4. Use angle regression to predict the target's direction of movement; After completing the above tasks, various anchor box mechanisms are used to cover targets of different sizes and shapes, and positive and negative sample balance is achieved through hard negative sample mining. The losses of each task are integrated using a joint loss function, including cross-entropy loss and smoothing L1 loss. After obtaining the preliminary results, non-maximum suppression is performed to eliminate redundant detection boxes, and speed and direction are corrected based on historical data.

6. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 1, characterized in that: in, The aforementioned wave-adaptive correction of input images affected by water surface fluctuations is performed by training on data collected under various wave conditions; deep convolutional neural networks or deformable convolutional networks are used to capture and correct image deformations; and perceptual loss or structural similarity index is selected as the optimized loss function to ensure the quality of the corrected image. The edge computing and model compression technologies mentioned above are used to ensure real-time operation on the edge devices of unmanned vessels, and are combined with model pruning, model quantization, model distillation and specialized edge computing frameworks, including TensorRT, to optimize and deploy models. The unsupervised or semi-supervised model training technique involves using an autoencoder to learn the internal structure of the data; then using generative adversarial networks (GANs) to assist training and generate synthetic samples; employing semi-supervised learning strategies, including MixMatch or Pseudo-Labeling, to combine a small amount of labeled data with a large amount of unlabeled data; performing data augmentation on the unlabeled data to generate new training samples and ensuring that the model's output remains consistent across different augmented versions; and using a self-learning strategy to use the model's output as pseudo-labels for the unlabeled data for further training.

7. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 2, characterized in that... The specific implementation process of using statistical methods and ensemble machine learning models is as follows: First, calculate the Z-score of the target captured from the remote sensing image, including pixel intensity, size, shape, and motion speed. The formula is: Furthermore, its interquartile range (IQR) is calculated using the formula IQR = Q3 - Q1; If the absolute value of the obtained Z-score exceeds a predetermined threshold or the target attribute exceeds the boundary calculated based on IQR, the target is marked as a preliminary anomaly. Secondly, these identified target features are input into a pre-trained One-Class SVM or IsolationForest model. Both models have been trained on historical data to learn the normal distribution of the data and identify outlier data points that deviate from this distribution. Model judgment based on One-Class SVM relies on the following formula: The weights w, slack variables ξ, and the offset ρ of the decision function have all been determined in the previous model training. If a target is flagged as an anomaly by a machine learning model, it is further labeled and analyzed to determine its nature.

8. The method for dynamic target detection and tracking of unmanned surface vessel remote sensing images according to claim 2, characterized in that... The specific process of feature engineering using the PCA, T-SNE, and UMAP technologies is as follows: S1. Using PCA, first calculate the covariance matrix of the data, then find the eigenvectors and eigenvalues ​​of the covariance matrix, and sort them according to the size of the eigenvalues; select the first k principal components for data dimensionality reduction. PCA(X)=X×eigenvectors(Cov(X)) Where X is the original data, and Cov(X) is the covariance matrix of X; S2. Using the T-SNE technique, based on the similarity between two points in the original space and the low-dimensional space, the Kullback-Leibler divergence between them is minimized to reduce the dimensionality of the data; Where, p ij and q ij These represent the similarity of two points in the original space and the lower-dimensional space, respectively. S3. Using UMAP technology, a topological representation of the original space is constructed, and then the best representation in the low-dimensional space is found; UMAP aims to minimize the change in distance and preserve the topological structure of the data; Where f is the mapping function, w ij It is the weight, which reflects the similarity between point i and point j.

Citation Information

Patent Citations

  • Self-adaptive multi-scale target detection method and system for unmanned ship environment

    CN113933828A

  • System and method for automatic tracking of marine objects

    KR1020180065411A