Underwater sound target radiation noise classification method

Through the combination of multimodal feature fusion and lightweight deep learning models, the data dependence and real-time problems of existing water acoustic target radiation noise classification methods are solved, and efficient and accurate water acoustic target radiation noise classification is achieved, which is suitable for marine monitoring and safety applications.

CN120164481APending Publication Date: 2025-06-17ZHONGCHUAN NO 9 DESIGN & RES INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510054031.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing classification methods for radiating noise of water acoustic targets have problems such as strong data dependence, insufficient feature extraction, and poor real-time performance, making it difficult to achieve high accuracy and low latency classification in complex marine environments.

Method used

The multimodal feature fusion method is adopted, combining spectrum, time domain and spatial domain features, and is trained through lightweight deep learning models (such as MobileNet), and transfer learning and hyperparameter optimization technologies are introduced to improve the generalization and real-time processing capabilities of the model.

Benefits of technology

It realizes efficient and accurate classification of water acoustic target radiation noise in environments with limited resources, with good real-time and low energy consumption characteristics, and is suitable for a variety of marine monitoring and safety application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164481A_ABST
    Figure CN120164481A_ABST
Patent Text Reader

Abstract

The invention relates to the related field of underwater acoustic signal processing, in particular to an underwater acoustic target radiation noise classification method, which constructs an efficient and accurate underwater acoustic target radiation noise classification model through the technologies of lightweight model selection, feature fusion, transfer learning, hyper-parameter optimization, model pruning, quantification and the like. The model not only can realize real-time processing in an environment with limited resources, but also has good classification performance, low energy consumption and high flexibility, is suitable for various application occasions, and particularly has remarkable advantages in an underwater acoustic target classification and real-time monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater acoustic signal processing, and specifically to a method for classifying underwater acoustic target radiated noise. Background Technique

[0002] The classification of underwater acoustic target radiated noise is one of the key technologies in the fields of ocean monitoring, ocean security, etc. Traditional methods for classifying underwater acoustic target radiated noise mainly rely on means such as spectral analysis and time-domain analysis, and combine prior knowledge for classification. However, due to the complex and changeable ocean environment, parameters such as the characteristic frequency band and intensity of the target radiated noise will be affected by various factors such as ocean currents, temperature, salinity, and marine organisms, resulting in the accuracy and robustness of traditional methods in practical applications being difficult to meet the requirements.

[0003] With the development of deep learning technology, noise classification methods based on deep learning have gradually become a research hotspot. Such methods can learn the deep features of target radiated noise on the basis of a large amount of data by constructing complex neural network models, thereby improving the accuracy and robustness of classification. However, the existing methods for classifying underwater acoustic target radiated noise based on deep learning still have the following problems: strong data dependence: usually a large amount of labeled data is required to train the model, and the cost of actually obtaining such data is high and time-consuming. Insufficient comprehensive feature extraction: Many methods only rely on spectral features and ignore the features in the time domain and spatial domain, resulting in limited classification performance. Poor real-time performance: The computational complexity of some models is high and it is difficult to be applied in real-time processing. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for classifying underwater acoustic target radiated noise to solve the problems raised in the above background technique.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for classifying underwater acoustic target radiated noise, including the following steps:

[0006] Step 1, Data preprocessing

[0007] Data acquisition: Collect target radiated noise data through underwater acoustic sensors under different sea areas and different environmental conditions; Data cleaning: Remove noise and interference signals in the collected data to ensure the cleanliness and effectiveness of the data; Data annotation: Annotate the collected data to generate training and test data sets;

[0008] Step 2, Feature extraction

[0009] Spectrum features: Perform short-time Fourier transform on the preprocessed data to generate a spectrogram and extract spectrum features; Time-domain features: Calculate the statistical features of the preprocessed data in the time domain; Spatial-domain features: Use a multi-sensor array to extract the spatial distribution features of the target radiated noise, specifically including direction of arrival and beamforming;

[0010] Step 3, Feature fusion

[0011] Multi-modal feature fusion: Perform multi-modal fusion on the spectrum features, time-domain features, and spatial-domain features to form a comprehensive feature vector; Feature dimensionality reduction: Use principal component analysis or t-SNE method to reduce the dimensionality of the comprehensive feature vector and reduce the computational complexity;

[0012] Step 4, Model construction

[0013] Lightweight deep learning model: Select a lightweight deep learning model with low computational complexity and few parameters, specifically select the MobileNet model; Model training: Use the fused feature vector to train the lightweight deep learning model, and improve the classification performance of the model through cross-validation and hyperparameter optimization; Model optimization: Introduce transfer learning technology and use a pre-trained model to improve the generalization ability and classification accuracy of the model;

[0014] Step 5, Real-time classification

[0015] Model deployment: Deploy the trained model to an embedded device or edge computing node to achieve low-latency real-time classification; Classification result output: During the real-time processing, output the classification result of the target radiated noise, supporting multiple output formats.

[0016] Preferably, the specific processing steps of the data preprocessing in step 1 are as follows:

[0017] Step 11, Acquisition device selection: Select underwater acoustic sensors, specifically including one or more of hydrophones, sonars, and underwater acoustic arrays, select a suitable sensor according to the characteristics of the target radiated noise, and adjust the number and position of the sensors according to actual requirements and resource conditions;

[0018] Step 12, Acquisition environment selection: Collect data in different sea areas, including coastal waters, open ocean, and deep sea, and collect data under different environmental conditions, including different sea currents, temperatures, salinities, and weather conditions;

[0019] Step 13, Acquisition parameter setting: Set a suitable sampling frequency according to the frequency range of the target radiated noise. Specifically, the sampling frequency should be greater than or equal to twice the highest frequency of the target radiated noise to meet the Nyquist sampling theorem; Select different time periods for data collection, including day, night, and tidal changes;

[0020] Step 14, Data Storage: Store the data collected in Step 13 in a standard file format and regularly back up the collected data to prevent data loss;

[0021] Step 15, Background Noise Removal: Use an adaptive filter to separate the background noise in the collected data; decompose the collected data into different scales through wavelet transform, and then remove the high-frequency noise;

[0022] Step 16, Interference Signal Removal: Use a band-pass filter or a band-stop filter in the frequency domain to remove the interference signals of specific frequencies in the collected data; use a high-pass filter or a low-pass filter in the time domain to remove the interference signals in specific time periods of the collected data; use adaptive noise cancellation technology to cancel the background noise and interference signals in the collected data;

[0023] Step 17, Data Annotation: Select experts with rich marine acoustics knowledge to use professional data annotation tools for data annotation, where the annotation content includes target type, environmental conditions, collection time, and collection location;

[0024] Step 18, Dataset Division: Training set: Used to train the deep learning model, accounting for 70% of the total data; Validation set: Used for cross-validation and hyperparameter optimization, accounting for 15% of the total data; Test set: Used to evaluate the final performance of the model, accounting for 15% of the total data.

[0025] Preferably, the specific steps of the feature extraction in Step 2 are as follows:

[0026] Step 21, Spectral Feature Extraction: Divide the preprocessed data in Step 1 into multiple short-time windows, with the length of each window set to 256 points and the overlap rate of 50%. To reduce the leakage effect at the window edges, apply a window function to each short-time window, and perform a discrete Fourier transform on each windowed short-time window to generate a spectrogram. The spectrogram is a two-dimensional matrix, where the horizontal axis represents time, the vertical axis represents frequency, and each element in the matrix represents the frequency energy at the corresponding time point, that is, splice the spectral results of multiple short-time windows into a spectrogram; extract key features from the generated spectrogram. The spectral features include but are not limited to the following: spectral mean, spectral variance, spectral peak, spectral bandwidth, and spectral energy; combine the above-extracted spectral features into a feature vector, where each dimension of the feature vector represents a spectral feature;

[0027] Step 22, Time-Domain Feature Extraction: Calculate the mean, variance, difference between the maximum and minimum values, kurtosis, and skewness of the preprocessed data in the time domain, and combine the above-extracted time-domain features into a feature vector, that is, each dimension of the feature vector represents a time-domain feature;

[0028] Step 23, Spatial-Domain Feature Extraction:

[0029] Direction of Arrival (DOA) estimation: Multiple sensor arrays are used, specifically linear arrays and uniform circular arrays. By calculating the time delay of the noise signals received by different sensors, the direction of arrival is estimated. Using the time delay information and the geometric layout of the sensor array, the direction of arrival is calculated.

[0030] Beamforming: The noise signals received by each sensor are obtained. According to the target direction and the sensor positions, the signals of each sensor are weighted, and the weighted signals are synthesized into a beam pattern.

[0031] Feature extraction: Direction energy distribution: The energy distribution characteristics of the target direction are extracted from the beam pattern; DOA change rate: The change rate of the direction of arrival over time is calculated to reflect the movement of the target; The above-mentioned spatially domain features extracted, namely the direction energy distribution and the DOA change rate, are combined into a feature vector, and each dimension of the feature vector represents a spatially domain feature.

[0032] Preferably, the specific steps of feature fusion in step 3 are as follows:

[0033] Step 31, Multimodal feature fusion

[0034] Feature vector concatenation: The three feature vectors in step 2, namely the spectral feature, the time-domain feature, and the spatially domain feature, are concatenated in sequence into a high-dimensional comprehensive feature vector;

[0035] Normalization processing: The concatenated feature vector is normalized to ensure the consistency of the numerical ranges between different features;

[0036] Feature selection: The most relevant features are selected from the comprehensive feature vector to further improve the classification performance of the model. The feature selection methods include feature selection based on information gain, feature selection based on variance, and feature selection based on correlation coefficient;

[0037] Step 32, Feature dimensionality reduction: The high-dimensional comprehensive feature vector in step 31 is converted into a low-dimensional feature vector to reduce the computational complexity and improve the real-time processing ability of the model. The specific feature dimensionality reduction methods used include principal component analysis and t-SNE.

[0038] Preferably, the specific steps of model construction in step 4 are as follows:

[0039] Step 41, MobileNet model structure

[0040] Input layer: Set the size of the input layer according to the dimension of the feature vector after dimensionality reduction in step 3; Depthwise separable convolutional layer: Use multiple depthwise separable convolutional layers to extract features; Pooling layer: Use the max pooling layer to reduce the size of the feature map; Fully connected layer: Use the fully connected layer for further processing and classification of features; Output layer: Set the number of neurons in the output layer to the number of classes, and use the softmax function to calculate the class probabilities;

[0041] Step 42, Model training: Randomly select 70% of the data from the fused feature vectors as the training set, 15% as the validation set, and 15% as the test set. Subsequently, use the cross-entropy loss function. Then divide the training set into k subsets. Each time during training, use k - 1 subsets as the training data, and the remaining 1 subset as the validation data. Repeat this k times, each time selecting a different subset as the validation set, record the performance metrics of each validation set, and take the average as the final performance evaluation result. Then use random search for hyperparameter optimization, evaluate the performance of each hyperparameter combination on the validation set, and select the hyperparameter combination with the best performance;

[0042] Step 43, Transfer learning: Select the MobileNet model pre-trained on ImageNet, load the weights of the MobileNet model pre-trained on ImageNet, freeze the first 10 depthwise separable convolutional layers of the pre-trained model, only fine-tune the subsequent fully connected layers, use the fused feature vectors for fine-tuning, and improve the performance through cross-validation and hyperparameter optimization; Evaluate the performance of the fine-tuned model on the validation set and the test set.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows: Through technologies such as the selection of lightweight models, feature fusion, transfer learning, hyperparameter optimization, model pruning, and quantization, the present invention constructs an efficient and accurate underwater acoustic target radiated noise classification model. This model can not only achieve real-time processing in an environment with limited resources, but also has good classification performance, low energy consumption, and high flexibility, and is applicable to a variety of application scenarios, especially having significant advantages in underwater acoustic target classification and real-time monitoring systems. Description of the Drawings

[0044] Figure 1 It is a schematic flowchart of the method of the present invention. Detailed Embodiments

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] Please refer to Figure 1 , the present invention provides a technical solution: an underwater acoustic target radiated noise classification method, including the following steps:

[0047] Step 1, data preprocessing

[0048] Data acquisition: Collect target radiated noise data through underwater acoustic sensors under different sea areas and different environmental conditions; data cleaning: Remove noise and interference signals in the collected data to ensure the cleanliness and effectiveness of the data; data annotation: Annotate the collected data to generate training and test data sets;

[0049] Step 2, feature extraction

[0050] Spectrum features: Perform short-time Fourier transform on the preprocessed data to generate a spectrogram and extract spectrum features; time-domain features: Calculate the statistical features of the preprocessed data in the time domain; spatial-domain features: Use a multi-sensor array to extract the spatial distribution features of the target radiated noise, specifically including direction of arrival and beamforming;

[0051] Step 3, feature fusion

[0052] Multi-modal feature fusion: Perform multi-modal fusion on spectrum features, time-domain features and spatial-domain features to form a comprehensive feature vector; feature dimensionality reduction: Use principal component analysis or t-SNE method to reduce the dimensionality of the comprehensive feature vector and reduce the computational complexity;

[0053] Step 4, model construction

[0054] Lightweight deep learning model: Select a lightweight deep learning model with low computational complexity and few parameters, specifically select the MobileNet model; model training: Use the fused feature vector to train the lightweight deep learning model, and improve the classification performance of the model through cross-validation and hyperparameter optimization; model optimization: Introduce transfer learning technology and use a pre-trained model to improve the generalization ability and classification accuracy of the model;

[0055] Step 5, real-time classification

[0056] Model deployment: Deploy the trained model to an embedded device or an edge computing node to achieve low-latency real-time classification; classification result output: During the real-time processing, output the classification result of the target radiated noise, supporting multiple output formats.

[0057] Further, the specific processing steps of data preprocessing in step 1 are as follows:

[0058] Step 11. Acquisition device selection: Select underwater acoustic sensors, specifically including one or more of hydrophones, sonars, and underwater acoustic arrays. Select appropriate sensors according to the characteristics of the target radiated noise, and adjust the number and location of the sensors according to actual requirements and resource conditions;

[0059] Step 12. Acquisition environment selection: Collect data in different sea areas, including coastal waters, open oceans, and deep seas. At the same time, collect data under different environmental conditions, including different ocean currents, temperatures, salinities, and weather conditions;

[0060] Step 13. Acquisition parameter setting: Set an appropriate sampling frequency according to the frequency range of the target radiated noise. Specifically, the sampling frequency should be greater than or equal to twice the highest frequency of the target radiated noise to meet the Nyquist sampling theorem; Select different time periods for data collection, including daytime, night, and tidal changes;

[0061] Step 14. Data storage: Store the data collected in Step 13 in a standard file format, and regularly back up the collected data to prevent data loss;

[0062] Step 15. Background noise removal: Use an adaptive filter to separate the background noise in the collected data; Decompose the collected data into different scales through wavelet transform, and then remove the high-frequency noise;

[0063] Step 16. Interference signal removal: Use a band-pass filter or a band-stop filter in the frequency domain to remove interference signals at specific frequencies in the collected data; Use a high-pass filter or a low-pass filter in the time domain to remove interference signals during specific time periods in the collected data; Use adaptive noise cancellation technology to cancel the background noise and interference signals in the collected data;

[0064] Step 17. Data annotation: Select experts with rich knowledge of underwater acoustics to use professional data annotation tools for data annotation. The annotation content includes target type, environmental conditions, collection time, and collection location;

[0065] Step 18. Dataset division: Training set: Used to train the deep learning model, accounting for 70% of the total data; Validation set: Used for cross-validation and hyperparameter optimization, accounting for 15% of the total data; Test set: Used to evaluate the final performance of the model, accounting for 15% of the total data.

[0066] Furthermore, the specific steps of feature extraction in Step 2 are as follows:

[0067] Step 21: Spectrum feature extraction: The preprocessed data in Step 1 is segmented into multiple short-time windows, with the length of each window set to 256 points and the overlap rate of 50%. To reduce the leakage effect at the window edges, each short-time window is windowed, and the discrete Fourier transform is performed on each windowed short-time window to generate a spectrogram. The spectrogram is a two-dimensional matrix, where the horizontal axis represents time, the vertical axis represents frequency, and each element in the matrix represents the frequency energy at the corresponding time point, that is, the spectrum results of multiple short-time windows are spliced into a spectrogram; Key features are extracted from the generated spectrogram. The spectrum features include but are not limited to the following: spectrum mean, spectrum variance, spectrum peak, spectrum bandwidth, and spectrum energy; The above-extracted spectrum features are combined into a feature vector, and each dimension of the feature vector represents a spectrum feature;

[0068] Step 22: Time-domain feature extraction: Calculate the mean, variance, difference between the maximum and minimum values, kurtosis, and skewness of the preprocessed data in the time domain. The above-extracted time-domain features are combined into a feature vector, that is, each dimension of the feature vector represents a time-domain feature;

[0069] Step 23: Spatial-domain feature extraction:

[0070] Direction of arrival (DOA) estimation: Use multiple sensor arrays, specifically linear arrays and uniform circular arrays. By calculating the time delay of the noise signals received by different sensors, estimate the direction of arrival, and calculate the direction of arrival using the time delay information and the geometric layout of the sensor array;

[0071] Beamforming: Obtain the noise signals received by each sensor. According to the target direction and sensor positions, perform weighted processing on the signals of each sensor, and synthesize the weighted signals into a beam pattern;

[0072] Feature extraction: Direction energy distribution: Extract the energy distribution feature of the target direction from the beam pattern; DOA change rate: Calculate the change rate of the direction of arrival over time to reflect the movement of the target; The above-extracted spatial-domain features, that is, the direction energy distribution and DOA change rate, are combined into a feature vector, and each dimension of the feature vector represents a spatial-domain feature.

[0073] Furthermore, the specific steps of feature fusion in Step 3 are as follows:

[0074] Step 31: Multimodal feature fusion

[0075] Feature vector splicing: Splice the three feature vectors in Step 2, that is, the spectrum features, time-domain features, and spatial-domain features, in sequence into a high-dimensional comprehensive feature vector;

[0076] Normalization processing: Perform normalization processing on the spliced feature vector to ensure that the numerical ranges between different features are consistent;

[0077] Feature selection: Select the most relevant features from the comprehensive feature vector to further improve the classification performance of the model. The feature selection methods include feature selection based on information gain, feature selection based on variance, and feature selection based on correlation coefficient.

[0078] Step 32, Feature dimensionality reduction: Convert the high-dimensional comprehensive feature vector in Step 31 into a low-dimensional feature vector to reduce the computational complexity and improve the real-time processing ability of the model. The specific feature dimensionality reduction methods used include principal component analysis and t-SNE.

[0079] Furthermore, the specific steps of model construction in Step 4 are as follows:

[0080] Step 41, MobileNet model structure

[0081] Input layer: Set the size of the input layer according to the dimension of the feature vector after dimensionality reduction in Step 3; Depthwise separable convolutional layer: Use multiple depthwise separable convolutional layers to extract features; Pooling layer: Use the max pooling layer to reduce the size of the feature map; Fully connected layer: Use the fully connected layer for further processing and classification of features; Output layer: Set the number of neurons in the output layer to the number of classes, and use the softmax function to calculate the class probabilities.

[0082] Step 42, Model training: Randomly select 70% of the data from the fused feature vector as the training set, 15% as the validation set, and 15% as the test set. Then use the cross-entropy loss function. Next, divide the training set into k subsets. Each time during training, use k - 1 subsets as the training data, and the remaining 1 subset as the validation data. Repeat this k times, each time selecting a different subset as the validation set, record the performance metrics of each validation set, and take the average as the final performance evaluation result. Then use random search for hyperparameter optimization, evaluate the performance of each hyperparameter combination on the validation set, and select the hyperparameter combination with the best performance.

[0083] Step 43, Transfer learning: Select the MobileNet model pre-trained on ImageNet, load the weights of the MobileNet model pre-trained on ImageNet, freeze the first 10 depthwise separable convolutional layers of the pre-trained model, only fine-tune the subsequent fully connected layer, and use the fused feature vector for fine-tuning to improve the performance through cross-validation and hyperparameter optimization; Evaluate the performance of the fine-tuned model on the validation set and the test set.

[0084] Efficient utilization of computing resources in the present invention: Selection of lightweight models: The present invention selects lightweight deep learning models with low computational complexity and few parameters (such as MobileNet). While maintaining high classification performance, these models significantly reduce the consumption of computing resources, enabling the model to achieve real-time processing in resource-constrained environments (such as embedded devices, mobile devices, etc.); Model pruning: Through model pruning techniques, unimportant weights are removed, further reducing the number of parameters and computational complexity of the model, and improving the running efficiency of the model; Model quantization: By converting floating-point numbers in the model to fixed-point numbers, the storage requirements and computational requirements of the model are reduced, enabling the model to run efficiently on resource-constrained devices;

[0085] High classification performance: Feature fusion: The present invention uses the fused feature vectors for model training. These feature vectors contain feature information from multiple different sources, improving the quality of the input data of the model, and thus enhancing the classification performance. Transfer learning: By introducing transfer learning techniques and fine-tuning using the weights of pre-trained models, the generalization ability and classification accuracy of the model are improved. The pre-trained models have been trained on large-scale datasets and have good generalization ability, which can provide powerful initial weights for the classification tasks of the present invention. Hyperparameter optimization: Through methods such as cross-validation and random search for hyperparameter optimization, the best combination of hyperparameters is found, further enhancing the classification performance of the model;

[0086] Fast training and deployment: Shorter training time: The lightweight model has low computational complexity and few parameters, so the training time is significantly shortened. This enables the model to be trained quickly in the case of large data volume or limited training resources. Easy deployment: Due to the lightweight characteristics of the model, the model of the present invention can be easily deployed to resource-constrained devices, such as embedded devices, drones, underwater robots, etc., without expensive computing resources.

[0087] Real-time performance and response ability: Real-time processing: The combination of lightweight models and optimization techniques enables the model of the present invention to achieve real-time processing, timely respond to the classification requirements of the radiated noise of underwater acoustic targets, and is suitable for scenarios of real-time monitoring and rapid decision-making. Low latency: Due to the low computational complexity, few parameters, and fast processing speed of the model, low latency can be achieved in real-time applications, improving the response speed and efficiency of the system.

[0088] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for classifying the radiated noise of underwater acoustic targets, characterized in that: The following steps are involved: Step 1: Data preprocessing Data collection: Use hydroacoustic sensors to collect target radiation noise data in different sea areas and environmental conditions; Data cleaning: remove noise and interference signals from the collected data to ensure the cleanliness and validity of the data; Data annotation: annotate the collected data and generate training and test data sets; Step 2: Feature extraction Spectral features: perform short-time Fourier transform on the preprocessed data to generate a spectrum diagram and extract spectral features; Time domain features: calculate the statistical features of the preprocessed data in the time domain; Spatial domain features: use a multi-sensor array to extract the spatial distribution characteristics of the target radiation noise, including the direction of arrival and beamforming; Step 3: Feature Fusion Multimodal feature fusion: spectral features, time domain features and spatial domain features are multimodally fused to form a comprehensive feature vector; feature dimension reduction: principal component analysis or t-SNE method is used to reduce the dimension of the comprehensive feature vector to reduce the computational complexity; Step 4: Model construction Lightweight deep learning model: Select a lightweight deep learning model with low computational complexity and a small number of parameters, specifically the MobileNet model; Model training: Use the fused feature vectors to train the lightweight deep learning model, and improve the classification performance of the model through cross-validation and hyperparameter optimization; Model optimization: Introduce transfer learning technology and use pre-trained models to improve the generalization ability and classification accuracy of the model; Step 5: Real-time classification Model deployment: Deploy the trained model to embedded devices or edge computing nodes to achieve low-latency real-time classification; Classification result output: During real-time processing, the classification results of the target radiated noise are output and multiple output formats are supported.

2. A method for classifying underwater acoustic target radiation noise according to claim 1, characterized in that: The specific processing steps of data preprocessing in step 1 are as follows: Step 11, collection equipment selection: select a hydroacoustic sensor, specifically including one or more of a hydrophone, sonar, and a hydroacoustic array. Select a suitable sensor according to the characteristics of the target radiation noise, and adjust the number and location of the sensors according to actual needs and resource conditions; Step 12: Collection environment selection: Data collection is conducted in different sea areas, including nearshore, offshore, and deep sea, and data is collected under different environmental conditions, including different ocean currents, temperatures, salinities, and weather conditions; Step 13, acquisition parameter setting: set a suitable sampling frequency according to the frequency range of the target radiation noise. The specific sampling frequency should be greater than or equal to twice the highest frequency of the target radiation noise to meet the Nyquist sampling theorem; select different time periods for data collection, including daytime, nighttime, and tidal changes; Step 14, data storage: store the data collected in step 13 in a standard file format, and regularly back up the collected data to prevent data loss; Step 15: Background noise removal: Use an adaptive filter to separate the background noise in the collected data; The collected data is decomposed into different scales through wavelet transform, and then high-frequency noise is removed; Step 16, interference signal removal: use a bandpass filter or a band-stop filter in the frequency domain to remove interference signals of specific frequencies of the collected data; use a high-pass filter or a low-pass filter in the time domain to remove interference signals of specific time periods of the collected data; use adaptive noise cancellation technology to cancel background noise and interference signals of the collected data; Step 17: Data annotation: experts with rich knowledge of ocean acoustics are selected to use professional data annotation tools to annotate the data, where the annotation content includes target type, environmental conditions, acquisition time and acquisition location; Step 18. Dataset division: Training set: used to train deep learning models, accounting for 70% of the total data; Validation set: used for cross-validation and hyperparameter optimization, accounting for 15% of the total data; Test set: used to evaluate the final performance of the model, accounting for 15% of the total data.

3. The method for classifying underwater acoustic target radiation noise according to claim 1, characterized in that: The specific steps of feature extraction in step 2 are as follows: Step 21, spectrum feature extraction: divide the data preprocessed in step 1 into multiple short-time windows, the length of each window is set to 256 points, and the overlap rate is 50%. In order to reduce the leakage effect at the window edge, each short-time window is windowed, and each windowed short-time window is subjected to discrete Fourier transform to generate a spectrum diagram. The spectrum diagram is a two-dimensional matrix, the horizontal axis represents time, the vertical axis represents frequency, and each element in the matrix represents the frequency energy at the corresponding time point, that is, the spectrum results of multiple short-time windows are spliced ​​into a spectrum diagram; extract key features from the generated spectrum diagram, the spectrum features include but are not limited to the following: spectrum mean, spectrum variance, spectrum peak, spectrum bandwidth and spectrum energy; combine the above-extracted spectrum features into a feature vector, and each dimension of the feature vector represents a spectrum feature; Step 22, time domain feature extraction: Calculate the mean, variance, difference between the maximum and minimum values, kurtosis and skewness of the preprocessed data in the time domain, and combine the above extracted time domain features into a feature vector, that is, each dimension of the feature vector represents a time domain feature; Step 23: Spatial domain feature extraction: Direction of arrival estimation: Use multiple sensor arrays, specifically linear arrays and uniform circular arrays, to estimate the direction of arrival by calculating the time delay of noise signals received by different sensors. Use the time delay information and the geometric layout of the sensor array to calculate the direction of arrival. Beamforming: Obtain the noise signal received by each sensor, perform weighted processing on the signal of each sensor according to the target direction and sensor position, and synthesize the weighted signals into a beam pattern; Feature extraction: Directional energy distribution: Extract the energy distribution characteristics of the target direction from the beam pattern; DOA change rate: Calculate the rate of change of the direction of arrival over time to reflect the movement of the target; combine the above-extracted spatial domain features, namely the directional energy distribution and the DOA change rate, into a feature vector, and each dimension of the feature vector represents a spatial domain feature.

4. The method for classifying underwater acoustic target radiation noise according to claim 1, characterized in that: The specific steps of feature fusion in step 3 are as follows: Step 31: Multimodal feature fusion Feature vector concatenation: concatenate the three feature vectors in step 2, namely, spectrum feature, time domain feature and space domain feature, into a high-dimensional comprehensive feature vector in sequence; Normalization: Normalize the concatenated feature vectors to ensure that the numerical ranges of different features are consistent; Feature selection: Select the most relevant features from the comprehensive feature vector to further improve the classification performance of the model. Feature selection methods include feature selection based on information gain, feature selection based on variance, and feature selection based on correlation coefficient. Step 32, feature dimensionality reduction: Convert the high-dimensional comprehensive feature vector in step 31 into a low-dimensional feature vector to reduce the computational complexity and improve the real-time processing capability of the model. The specific feature dimensionality reduction methods used include principal component analysis and t-SNE.

5. The method for classifying underwater acoustic target radiation noise according to claim 1, characterized in that: The specific steps of model construction in step 4 are as follows: Step 41: MobileNet model structure Input layer: set the size of the input layer according to the dimension of the feature vector after dimensionality reduction in step 3; Depthwise separable convolution layer: use multiple depthwise separable convolution layers to extract features; Pooling layer: use the maximum pooling layer to reduce the size of the feature map; Fully connected layer: use the fully connected layer for further processing and classification of features; Output layer: set the number of neurons in the output layer to the number of categories, and use the softmax function to calculate the category probability; Step 42, model training: randomly extract 70% of the data from the fused feature vector as the training set, 15% as the validation set, and 15% as the test set, then use the cross entropy loss function to divide the training set into k subsets, use k-1 subsets as training data each time for training, and the remaining 1 subset as validation data, repeat k times, each time select a different subset as the validation set, record the performance indicators of each validation set, take the average as the final performance evaluation result, and then use random search to optimize hyperparameters, evaluate the performance of each hyperparameter combination on the validation set, and select the hyperparameter combination with the best performance; Step 43, transfer learning: Select the MobileNet model pre-trained on ImageNet, load the weights of the MobileNet model pre-trained on ImageNet, freeze the first 10 depthwise separable convolutional layers of the pre-trained model, fine-tune only the subsequent fully connected layers, use the fused feature vector for fine-tuning, improve performance through cross-validation and hyperparameter optimization; evaluate the performance of the fine-tuned model on the validation set and test set.

Citation Information

Cited By

  • Pest detection and expelling method based on sound recognition

    CN121122291A

  • Container non-intrusive detection method and device based on single-point sound pressure time domain signal

    CN121188580A

  • Non-invasive container inspection method and equipment based on single-point sound pressure time-domain signals

    CN121188580B