Passive Sonar Data Labeling Using AIS-Based Vessel Sound Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in passive sonar data analysis is the slow, inaccurate, and costly manual annotation of hydrophone recordings for vessel detection, classification, and identification, which is a bottleneck for training data-driven models like deep artificial neural networks, despite the availability of large datasets.
Innovation Solution
A computer-implemented method for automatically labeling passive sonar hydrophone data using publicly available vessel positioning data, such as AIS, to estimate sound metrics and apply thresholds for accurate labeling, enabling efficient and automated generation of labeled datasets for machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation of hydrophone recordings is performed by trained professionals, then labeling accuracy is improved, but productivity is worsened (slow and inefficient)
Solution Approach 1:
The system performs preliminary actions by automatically generating candidate labels using vessel position data and sound propagation models before human operators review them. This pre-processing step filters and prepares potential vessel detections, reducing the manual workload while maintaining accuracy through operator verification of automated suggestions.
2Measurement precision
If manual annotation by trained professionals is used, then labeling accuracy is improved, but loss of time is worsened (costly and inefficient)
Solution Approach 1:
The system introduces an intermediary automated labeling component that generates candidate labels using vessel position data and acoustic propagation models. This intermediary process handles the initial labeling task, allowing human operators to focus on reviewing and verifying results rather than performing manual annotation from scratch, thus reducing total annotation time while preserving accuracy.
3Quantity of substance
If large amounts of hydrophone recordings are collected, then quantity of training data is improved, but difficulty of detecting and measuring is worsened (harder to label)
Solution Approach 1:
The system segments the labeling task by automatically processing recordings in batches using vessel position data and sound propagation models. Each recording is divided into time segments where vessel presence is determined based on acoustic metrics, allowing large datasets to be processed systematically rather than as monolithic manual annotation projects.
Solution Approach 2:
The system enables self-service automated labeling by using publicly available vessel position data (AIS) and acoustic propagation models to generate labels without requiring human operators for each recording. This self-labeling capability allows the system to handle large volumes of data autonomously, reducing the complexity of manual labeling while maintaining data quality through automated consistency.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This method provides fast, accurate, and efficient labeling of hydrophone recordings, improving the training of machine learning models for vessel detection, classification, and identification, and enabling larger quantities of high-quality labeled data.
Implementation Method 1
determining, for a particular vessel identification information, an estimated received sound metric for each vessel throughout the detection time-period, based on the timestamped vessel data
Implementation Method 2
The estimated received sound metric may be a signal-to-interference-plus-noise ratio (SINR) signal estimated based on an estimated sound of the respective vessel and estimated noise and/or ambient sound
Data Source
Figure 1
Figure 2~3
Figure 4~6
AI summary
A computer-implemented method is provided of automatically labelling passive sonar hydrophone data for use as a supervised learning dataset, comprising: for a selected passive sonar hydrophone: obtaining actual received sound data throughout a detection time-period, obtaining a hydrophone location, and determining a maximum detection radius around the hydrophone location; retrieving timestamped vessel data associated with vessels within the maximum detection radius during the detection time-period, the timestamped vessel data comprising for each vessel: vessel location information and vessel identification information; determining, for a particular vessel identification information, an estimated received sound metric for each vessel throughout the detection time-period, based on the timestamped vessel data; labelling the actual received sound data, the labels including the particular vessel identification information and being associated with times during the detection time-period where the estimated received sound metric is above a threshold.