Underwater target positioning method, system and equipment based on multi-task learning, and medium
The underwater target localization method based on multi-task learning utilizes a convolutional neural network to construct a multi-task network model, solving the problems of error amplification and high computational load in underwater sound source localization in long-distance scenarios, and achieving high-precision, low-latency, and low-energy-consumption localization results.
Patent Information
- Application Number
- CN202511359664.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-02-17
AI Technical Summary
Existing underwater sound source localization technologies suffer from geometrically amplified errors in long-distance scenarios, are highly dependent on platform maneuverability and large-aperture array deployment, have high computational load, and lack robustness. They are difficult to meet the requirements for real-time updates and low energy consumption, and their engineering implementation costs are high.
A multi-task learning approach is adopted, which constructs a multi-task network model through a convolutional neural network and trains it using array-received signal data to achieve single forward inference output of azimuth, horizontal distance and sound source depth. A residual backbone network and a differentiated weighted joint loss are designed to avoid error accumulation.
It achieves high-precision, low-latency, and low-power underwater target localization in long-distance scenarios, reducing computational complexity and hardware costs, and enhancing adaptability to noise and array errors.
Smart Images

Figure CN121541143A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of acoustic underwater positioning, and particularly relates to an underwater target positioning method, system, device and storage medium based on multi-task learning. BACKGROUND
[0002] Although there are many kinds of existing underwater acoustic source positioning technologies, these methods all have deficiencies in practical engineering applications. First, almost all geometric ranging schemes that rely on time difference of arrival or angle difference of arrival will have geometrically amplified errors as the propagation distance increases: in an underwater long-distance scenario, a small deviation in time difference measurement can be amplified into a position error of hundreds of meters or even kilometers. Second, some methods have rigid dependence on platform maneuvering, large-aperture array deployment or long-time observation. For example, target motion analysis methods require significant maneuvering to ensure convergence, and the deployment and attitude control of vertical arrays or horizontal large-aperture arrays in deep water are complex and costly, and once the array shape is distorted, the beamforming and multipath separation capabilities will be severely affected. Third, most three-dimensional search algorithms (such as matched field processing algorithms) have huge computational loads, and real-time applications are restricted; while multi-path feature matching, although relatively simplified in computational load, still requires global search or long sequence tracking when distance and depth are coupled and unobservable, resulting in a delay that cannot meet the second-level update requirement. Fourth, the robustness of traditional physical model methods and existing single-task or vertical array-multi-task deep learning methods under conditions of multipath, low signal-to-noise ratio and array misalignment is still insufficient, and often requires accurate environmental priors or additional correction steps; once the sound speed profile or seabed parameters are mismatched by more than a few percent, the positioning error will increase by an order of magnitude. Finally, the engineering landing and maintenance costs are always high, large-scale array hardware itself is expensive and difficult to recover, and if the parameter quantity of an end-to-end deep network is too large, it requires high-power GPU support, which is difficult to deploy on resource-limited platforms such as offshore nodes or unmanned underwater vehicles. SUMMARY
[0003] The present application provides an underwater target positioning method, system and storage medium based on multi-task learning, which can simultaneously output azimuth, horizontal distance and sound source depth through a single forward inference.
[0004] In a first aspect, the present application provides an underwater target positioning method based on multi-task learning, which comprises:
[0005] Generating array received signal data including a multi-path array signal model based on a sound field model;
[0006] Constructing a multi-task network model using a convolutional neural network;
[0007] Constructing a training data set based on the array received signal data, the training data set including sample data of different azimuths, distances and sound source depths;
[0008] training the multi-task network model using the training data set to obtain a trained multi-task network model;
[0009] using the trained multi-task network model to locate the underwater target; wherein the locating result includes the azimuth angle, distance and sound source depth of the underwater target.
[0010] In a second aspect, the present application provides a system for locating underwater targets based on multi-task learning, comprising:
[0011] a data generation module configured to generate array received signal data including a multi-path array signal model based on a sound field model;
[0012] a model construction module configured to construct a multi-task network model using a convolutional neural network;
[0013] a training data construction module configured to construct a training data set based on the array received signal data, the training data set including sample data of different azimuth angles, distances and sound source depths;
[0014] a training module configured to train the multi-task network model using the training data set to obtain a trained multi-task network model;
[0015] a locating module configured to use the trained multi-task network model to locate the underwater target; wherein the locating result includes the azimuth angle, distance and sound source depth of the underwater target.
[0016] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus;
[0017] the memory is configured to store a computer program;
[0018] the processor is configured to execute the program stored on the memory to implement the steps of the method for locating underwater targets based on multi-task learning according to any one of the first aspect.
[0019] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of the method for locating underwater targets based on multi-task learning according to any one of the first aspect.
[0020] Compared with the prior art, the above technical solution provided by the embodiments of the present application has the following advantages: the azimuth angle, horizontal distance and sound source depth can be output simultaneously through single forward inference, and the error accumulation of the traditional multi-stage pipeline is avoided through the design of the residual backbone network, the three-way parallel output head and the differential weighting joint loss, so as to meet the comprehensive requirements of the measurement platform for high precision, low delay and low energy consumption. For example, the error accumulation is eliminated: the same feature flow simultaneously regresses three parameters to avoid cascade amplification; the delay is significantly reduced: the inference times are reduced from'multi-section' to 'one section', and the actual measurement frame processing time can be reduced to tens of milliseconds; the residual structure and multi-task collaborative learning strengthen the adaptability to noise, array error and sound speed error. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, those skilled in the art can obtain other drawings from these drawings without any creative effort.
[0023] Figure 1 A flowchart of a multi-task learning-based underwater target positioning method provided by the embodiments of the present application;
[0024] Figure 2 A multi-task network model training method shown in some embodiments of the present application;
[0025] Figure 3 A specific model structure diagram of the multi-task network model provided by the embodiments of the present application;
[0026] Figure 4 A structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0027] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the scope of protection of the present application.
[0028] Underwater acoustic positioning of targets usually relies on the propagation characteristics of acoustic waves in water bodies to invert the position, distance, and depth of the target. Due to the objective limitations of the underwater channel, such as significant sound channel effects, complex multipath propagation, difficulty in obtaining environmental parameters, and high cost of array deployment in deep sea compared with shallow sea, it is difficult to balance positioning accuracy and robustness. The current public technical route can be summarized into the following categories:
[0029] The first category is the classical three-element subarray method based on the geometric ranging concept, which quickly solves the target position by measuring the time difference of arrival or the angle difference of arrival between three sensors. This method has simple structure and low implementation cost, and can obtain high accuracy in near-field or medium-distance scenarios, but the ranging error increases geometrically with the propagation distance; in long-range positioning, a small time difference measurement error can cause hundreds of meters or even kilometers of position deviation.
[0030] The second category, Target Motion Analysis (TMA), inverts the target motion trajectory through a series of bearing or distance observations during platform maneuvering. TMA has low sensitivity to prior information such as environmental sound speed profile, is suitable for single-platform combat scenarios, but requires significant platform maneuvering to converge; for slow-moving or stationary targets, the amount of observation information is insufficient, and it is prone to solution divergence or slow convergence.
[0031] The third category is Matched Field Processing (MFP) and its derivative methods. MFP uses an acoustic propagation model to calculate a “copy field”, matches the measured sound field with the theoretical sound field, and achieves high-precision positioning; frequency difference matched field (FDSL MFP) and normal mode matching further simplify the matching quantity in the feature domain and reduce the computational complexity. However, this type of method is highly dependent on environmental parameters such as sound speed, seabed topography, and bottom material. Once the model error exceeds 1-2%, the positioning error may be orders of magnitude larger; at the same time, the computational load brought by large-scale three-dimensional search restricts real-time application.
[0032] The fourth category is the matching method based on the characteristics of multipath arrival, including the joint matching of multipath time delay, arrival angle, and interference period. This idea uses the multipath information in the complex deep sea sound field to enhance the constraints, which can to some extent alleviate the performance degradation caused by environmental mismatch, and is easy to realize in real time. However, the multipath characteristics at a single time usually have distance-depth coupling, which requires the aid of array angle measurement, platform maneuvering, or long-time sequence tracking to achieve observability, increasing the overall complexity of the system.
[0033] The fifth type is single hydrophone multipath positioning technology. It uses the time delay between direct wave and seabed reflected wave for autocorrelation speed measurement, and combines with a simple motion model to calculate the target position. The hardware cost is extremely low and the deployment is flexible. However, this method relies on strong and stable seabed reflection, and in low signal-to-noise or weak reflection environment, the multipath peak is easily submerged by noise; at the same time, in order to overcome the distance-depth coupling, the target often needs to move along the radial direction and be tracked for a long time, which limits the application scenarios.
[0034] The sixth type is vertical array or horizontal array arrival angle multipath joint method. The vertical array has good elevation resolution and can simultaneously solve the range and depth; the horizontal array has high azimuth resolution and is easy to obtain large aperture on the towed platform. However, large aperture array needs complex deployment and attitude control device, and is extremely sensitive to array shape distortion and sea current disturbance; once the array shape calibration error accumulates, the beamforming gain and multipath resolution will decrease significantly.
[0035] Finally, in recent years, the deep learning end-to-end positioning method tries to directly learn the nonlinear mapping between the sound field and the source position through neural network, in order to reduce the dependence on accurate physical model. This kind of method has shown certain potential on controlled experiment or simulation data, but generally adopts "single task-serial processing" method, which has the problems of difficult training sample acquisition, insufficient generalization ability of network to real environment changes and lack of physical interpretability, and has not formed reliable application of engineering scale.
[0036] Figure 1 A flowchart of an underwater target positioning method based on multi-task learning provided by an embodiment of the present application. In some embodiments, the flowchart can include the following operations:
[0037] Step 101, generating array received signal data including a multipath array signal model based on a sound field model.
[0038] The sound field model is a mathematical model that describes the propagation of sound in a medium (such as water). The sound field model can simulate the propagation behavior of sound in the environment, including reflection, refraction, attenuation and scattering, etc.
[0039] The multipath array signal model is a model that describes the arrival of sound waves at the receiving array through multiple paths (such as direct path, reflected path), which results in multiple delayed and attenuated copies of the signal, i.e. multipath effect. For example, in the marine environment, sound waves emitted from the sound source may directly reach the receiving array, or may reach the receiving array after being reflected by the sea surface or the seabed, forming multiple signal components.
[0040] Array received signal data refers to the acoustic signal data received by an array composed of multiple sensors (such as hydrophones), which are usually recorded in the form of a time series and include spatial and temporal information of acoustic waves. For example, a linear array is composed of 8 hydrophones, and the sound pressure signals recorded by each hydrophone change with time, which are combined into a multi-dimensional data array.
[0041] In some embodiments, the array received signal data is expressed by the following formula (1):
[0042]
[0043] where x(t)k is the array received signal data at time t, k is the number of sound sources, p is the number of multipath reflections, P = 0 represents the direct wave, and P ≥ 1 represents the Pth reflection path, α k,p is the complex amplitude attenuation factor of the Pth reflection path of the kth sound source, is the incidence angle of the Pth reflection path of the kth sound source, s k (t) is the sound source signal, is the steering vector of the uniform linear array, and the mth component is shown in the following formula (2):
[0044] [α(θ)] m = exp(j2πf0τ m ) (2)
[0045] where τ m represents the phase difference of the mth array source relative to the reference array element in the uniform linear array.
[0046] In some embodiments, the propagation of acoustic waves in underwater environments can be simulated by a sound field model, considering the multipath effect, calculating the path of acoustic waves from the sound source to each sensor in the array, and generating the corresponding signal data. Based on the sound source parameters (such as location, frequency) and environmental parameters (such as water depth, sound speed profile), then using numerical methods (such as ray tracing or wave equation solving) to calculate the delay, amplitude and phase of each path, and finally synthesizing the array received signal.
[0047] Step 102, constructing a multi-task network model using a convolutional neural network.
[0048] Convolutional neural network is a deep learning model specially designed for processing data with grid structure (such as images, signals), which automatically extracts local features through convolutional layers, reduces dimensions through pooling layers, and performs classification or regression through fully connected layers.
[0049] A multi-task network model is a neural network architecture that can learn multiple related tasks simultaneously, sharing underlying feature representations to improve learning efficiency and generalization. For example, a multi-task network might predict azimuth, range, and depth simultaneously, sharing the front convolutional layers but having multiple output branches, each corresponding to a task.
[0050] In some embodiments, a convolutional neural network architecture can be designed with an input layer adapted to receive signal data (e.g., time series or spectrograms), intermediate layers including multiple convolutional and pooling layers for feature extraction, and an output layer designed as multiple heads (i.e., multiple output branches), each corresponding to a localization task (e.g., azimuth prediction, range prediction, depth prediction). Network parameters (e.g., kernel size, number of layers) need to be adjusted according to signal characteristics. For example, a CNN can be constructed with input as short-time Fourier transform spectrograms of array signals, using 3 convolutional layers (each followed by ReLU activation and pooling layers), then flattening the features and feeding into three fully connected output layers, outputting azimuth (continuous value), range (continuous value), and depth (continuous value) respectively, using shared convolutional layers to reduce the number of parameters.
[0051] In some embodiments, the multi-task network model constructed includes an input preprocessing module, a common feature extraction layer, a feature fusion module, and a multi-task output module.
[0052] The input preprocessing module is used to receive the input covariance matrix and normalize the real and imaginary parts of R (the covariance matrix described later) respectively and then concatenate them along the input channel dimension.
[0053] The input preprocessing module is a component of the neural network, specifically responsible for the preliminary processing and processing of the original input data, making it more suitable for subsequent network layers to extract features and learn.
[0054] The covariance matrix is a square matrix whose elements represent the covariances between the components of a random vector. In array signal processing, it is estimated from received data and used to characterize the spatial correlation of the signal. For example, the covariance matrix of an 8-element array is an 8x8 complex matrix.
[0055] Normalization is a data preprocessing technique that scales data by a certain proportion, making it fall within a specific interval (e.g., [0, 1] or [-1, 1]), to eliminate the influence of dimension and speed up model convergence. For example, min-max normalization linearly transforms the original data to the [0, 1] interval.
[0056] Input channel dimension refers to the dimension in the neural network input data that represents different feature sources. For example, in image processing, the channel dimension usually represents color channels (such as RGB three channels). For example, a grayscale image input can be considered as having 1 channel, while an RGB image input has 3 channels.
[0057] In some embodiments, the input preprocessing module first receives a complex covariance matrix R as input. Then, the matrix R is separated into two real matrices: one composed of the real parts of all elements, and the other composed of the imaginary parts of all elements. Then, normalization processing is performed on the two real matrices respectively (for example, using min-max normalization, scaling the values in each matrix to the range [0, 1]). Finally, the normalized real part matrix and imaginary part matrix are considered as two different input channels, and are concatenated along the channel dimension to form a new tensor structure with two channels as the input of the subsequent network. For example, the input is an 8x8 complex covariance matrix. The module first extracts an 8x8 real matrix and an 8x8 imaginary matrix. Assuming the maximum value of the real matrix is 10 and the minimum value is -5, each element is converted to [0, 1] through the normalization formula. Similar operations are performed on the imaginary matrix. Finally, the two normalized 8x8 matrices are stacked to form a 2x8x8 three-dimensional tensor (2 channels, each channel 8x8), which is passed to the next layer.
[0058] The common feature extraction layer includes a first residual block, a second residual block and a third residual block connected in sequence, for feature extraction based on the output of the input preprocessing module.
[0059] The common feature extraction layer is part of a neural network, composed of a series of layers (such as convolutional layers, pooling layers), whose purpose is to extract abstract feature representations from input data, and these features are shared by all subsequent tasks. For example, a CNN for multi-task learning, the first few convolutional layers are common to all tasks.
[0060] Residual block is a basic building block in deep neural networks, which introduces "shortcut connection" or "jump connection" to directly pass the input to the output and add it to the result after several layers of transformation, to solve the problem of gradient vanishing and degradation in deep networks, and facilitate the training of deeper networks. For example, a basic residual block may include two convolutional layers and a ReLU activation function, the input is directly added to the final output by skipping these two layers.
[0061] Feature extraction refers to the process of transforming input data through various layers of a neural network (e.g., convolutional layers) to progressively extract increasingly higher-level feature representations useful for the task. For example, in image processing, lower-level convolutional layers might extract edge features, while deeper convolutional layers might extract part or whole object features.
[0062] Sequential connection refers to a way of connecting network layers or modules in a sequential manner, where the output of a previous layer is directly used as the input of the next layer. For example, the output of the first residual block is directly used as the input of the second residual block, and the output of the second residual block is directly used as the input of the third residual block.
[0063] The multi-channel tensor output by the pre-processing module is first sent to the first residual block. Inside each residual block, a series of operations (e.g., convolution, batch normalization, activation function) are performed, and the output is added to the original input of the block (through a shortcut connection) to form the final output of the residual block. The output of the first residual block is then sequentially sent to the second residual block for similar processing, and the output of the second residual block is sent to the third residual block. Through these three sequentially connected residual blocks, the input data is gradually transformed into more abstract and higher-level feature representations (often referred to as feature maps), which include shared information useful for solving all tasks (orientation, distance, depth estimation). For example, the input is a 2x8x8 tensor. The first residual block includes two convolutional layers using 3x3 convolutional kernels, outputting a 32-channel feature map. The output of this block is the result of the input after convolutional transformation plus the original input (adjusted in dimension by a 1x1 convolution to match). Its output (32x8x8) enters the second residual block, which outputs a 64-channel feature map (64x8x8). Finally, the third residual block outputs a 128-channel feature map (128x8x8). This 128x8x8 feature map is the output of the common feature extraction layer.
[0064] The feature fusion module is used to flatten the feature map output by the common feature extraction layer into a one-dimensional vector and sequentially pass it through a three-layer fully connected network for dimension mapping.
[0065] The feature fusion module is a component in a neural network that integrates and transforms features extracted from previous layers (which may be multi-dimensional) to prepare for the final task-specific output. The feature fusion module can serve to fuse global information and perform dimension conversion. For example, in a CNN, the feature fusion module includes operations to flatten the feature map and pass it through a fully connected layer.
[0066] A feature map is an output of a convolutional neural network after convolutional operations, which can be a multi-dimensional array (usually three-dimensional), representing the spatial distribution and intensity of certain features in the input data. For example, a feature map with size 128x8x8 indicates that the network has extracted 128 different features, each with response values on an 8x8 spatial grid.
[0067] A one-dimensional vector is a data structure consisting of a single dimension of numerical values. For example, a list [v1, v2..., v512] containing 512 numerical values is a one-dimensional vector.
[0068] Flattening is an operation that rearranges multi-dimensional data (such as a two-dimensional matrix or a three-dimensional tensor) into a one-dimensional vector, which can be adapted to the input requirements of a fully connected layer. For example, flattening a three-dimensional tensor of size 2x3x4 results in a one-dimensional vector containing 24 elements.
[0069] A fully connected network is a type of neural network layer where each neuron in the layer is connected to all neurons in the previous layer, used to learn non-linear combinations of features and perform dimension transformation. For example, a fully connected layer can have 512 neurons, each receiving the weighted sum of all outputs from the previous layer.
[0070] Dimension mapping refers to the transformation of the dimension (size) of input data through a neural network layer (such as a fully connected layer) to another specified dimension. For example, a fully connected layer can map a 1024-dimensional input vector to a 512-dimensional output vector.
[0071] The feature fusion module first receives the multi-dimensional feature map (such as a three-dimensional tensor) output by the common feature extraction layer. Then, using the flattening operation, all elements of the multi-dimensional feature map are arranged in a long one-dimensional vector in order. Then, this one-dimensional vector is sent to the first fully connected layer, which maps it to a new, usually lower, dimension. The output of the first fully connected layer is then sent to the second fully connected layer for further dimension mapping and feature transformation. The output of the second fully connected layer is finally sent to the third fully connected layer to complete the final dimension mapping, outputting a fixed-dimensional vector that integrates all extracted feature information, preparing for the final multi-task output. For example, the common feature extraction layer outputs a 128x8x8 feature map. The feature fusion module first flattens it into a one-dimensional vector containing 8192 elements (128x8x8 = 8192). Then, the vector passes through the first fully connected layer, with an output dimension of 1024. Next, it passes through the second fully connected layer, with an output dimension of 512. Finally, it passes through the third fully connected layer, with an output dimension of 256. This 256-dimensional vector is the final output of the feature fusion module.
[0072] A multi-task output module to output the predicted azimuth, range and source depth through three parallel regression tasks based on the output of the feature fusion module.
[0073] The multi-task output module is the part of the neural network responsible for producing the final prediction results. It usually consists of multiple parallel sub-networks or output layers, each producing output for a specific task.
[0074] Parallel regression tasks refer to multiple regression problems being solved simultaneously under a multi-task learning framework. The multiple regression tasks share the front part of the network but have independent computation paths (in parallel) in the output layer. The output of each path is a continuous numerical value. For example, in a localization task, azimuth estimation, range estimation and depth estimation are three parallel regression tasks.
[0075] Azimuth is the horizontal angle of the target relative to a reference direction (e.g. true north), usually measured in degrees. For example, an azimuth of 90 degrees means the target is in the true east direction.
[0076] Range is the horizontal straight-line distance between the target and the receiving array, usually measured in meters. For example, a range of 1500 meters means the target is 1500 meters away from the array.
[0077] Source depth is the vertical depth of the target underwater, measured from sea level, usually in meters. For example, a source depth of 50 meters means the target is 50 meters below sea level.
[0078] The multi-task output module receives a fixed-dimensional feature vector output by the feature fusion module. This feature vector is then fed into three independent fully connected layers (or small sub-networks) simultaneously. Each fully connected layer is responsible for a regression task: the first fully connected layer maps its input to a single numerical value as the predicted azimuth; the second fully connected layer also outputs a single numerical value as the predicted range; the third fully connected layer outputs a single numerical value as the predicted source depth. These three output layers work in parallel without interference, collectively constituting the final output of the network. For example, the feature fusion module outputs a 256-dimensional vector. The multi-task output module has three independent fully connected layers, each with 256 input nodes and 1 output node. The 256-dimensional vector is simultaneously input to these three layers. The first output layer calculates and outputs a value, e.g. 45.2, representing a predicted azimuth of 45.2 degrees. The second output layer outputs a value, e.g. 1850.7, representing a predicted range of 1850.7 meters. The third output layer outputs a value, e.g. 98.3, representing a predicted depth of 98.3 meters.
[0079] See Figure 3 , Figure 3is a schematic diagram of a specific model structure of the multi-task network model provided in the embodiments of the present application. It includes an input layer Input, a residual block ResBlock, a flattening layer Flatten, a fully connected layer FC, and an output layer.
[0080] The input layer is a double-channel complex spectrogram (2ch).
[0081] The residual block ResBlock1 / ResBlock2 / ResBlock3 each contains 3 layers of 3x3 convolution, and the number of channels decreases step by step C1-C2-C3 (for example, 512-256-128). Each main branch in the residual block contains 3 3x3 convolutions, and the bypass is a 1x1 convolution.
[0082] The flattening layer flattens the feature map into a one-dimensional vector
[0083] The fully connected layer FC includes 3 layers of fully connected layers (FC1 / FC2 / FC3).
[0084] The output layer is 3 parallel regression tasks (azimuth angle θ, distance R, and depth Zs).
[0085] Step 103, constructing a training data set based on the array received signal data.
[0086] The training data set includes sample data of different azimuths, distances, and sound source depths.
[0087] The training data set is a data set for training a machine learning model, composed of input samples and corresponding labels, and the purpose is to let the model learn the mapping from input to output. For example, in the positioning task, the training data set includes multiple groups of array received signal data (input) and corresponding azimuth angle, distance, and depth
[0088] Sample data is a single data instance in the training data set, including input features (such as signal data) and target values (such as positioning parameters). For example, a sample can be a 1-second signal segment received by a hydrophone array, and the label is an azimuth angle of 30 degrees, a distance of 1500 meters, and a depth of 50 meters.
[0089] In some embodiments, sample can be systematically extracted from the generated array receiving signal data, ensuring coverage of a wide range of sound source position parameters, i.e. different azimuth angles (such as 0 to 360 degrees), different distances (such as 100 meters to 5000 meters) and different sound source depths (such as 10 meters to 500 meters). Each sample corresponds to a unique sound source position, the input is the array signal data (possibly pre-processed, such as normalization or feature extraction), and the label is the true azimuth angle, distance and depth value. The dataset is usually divided into training set, validation set and test set. For example, 10,000 samples are generated using the sound field model, each sample corresponds to a randomly generated sound source position (azimuth angle, distance, depth), the array signal data is stored in matrix form (number of sensors x number of time points), the label is stored in vector form, and then saved as a dataset file to obtain the training dataset.
[0090] In some embodiments, constructing the training dataset based on the array receiving signal data can include the following operations:
[0091] S10, based on the array receiving signal data, obtaining an array manifold matrix.
[0092] The array manifold matrix is a mathematical matrix, the column vectors of which represent the response of the array to a unit amplitude plane wave from a specific direction (azimuth angle and elevation angle), which describes the spatial response characteristics of the array and is a core concept in array signal processing.
[0093] In some embodiments, the spatial characteristics of the array receiving signal data can be extracted to construct a matrix representing the response of the array to sound waves from different directions.
[0094] S11, adding Gaussian white noise to the array manifold matrix and performing signal discretization;
[0095] Gaussian white noise is a random noise with Gaussian (normal) distribution of statistical characteristics, and the power spectral density is uniformly distributed in the entire frequency domain (i.e. the characteristic of "white").
[0096] Signal discretization refers to quantizing continuous parameters (such as azimuth angle, signal-to-noise ratio) to generate a finite number of discrete sample points for numerical calculation and dataset construction.
[0097] In order to simulate the noise existing in the actual environment and enhance the diversity of the dataset, a matrix with the same dimension as the array manifold matrix and Gaussian white noise elements is first generated, and then weighted by a certain signal-to-noise ratio (SNR) and added to the array manifold matrix. Then, the continuous parameters (such as azimuth angle, signal-to-noise ratio) on which the signal data after adding noise depends are discretized to generate a series of discrete working condition points.
[0098] In some embodiments, the adding of the Gaussian white noise to the array manifold matrix is represented by the following formula (3):
[0099] x(t) = A eff s(t) + n(t) (3)
[0100] wherein A eff = [A0A1…A P ] ∈ C M×K(P+1) is an array manifold matrix, s(t) = [s1(t), …, s K (t)] T is each sound source signal, and n(t) is a Gaussian white noise.
[0101] In some embodiments, the signal discretization can be performed by the following formula (4):
[0102] X = A eff S + N (4)
[0103] wherein X is a discretized received data matrix, A eff is an array manifold matrix, S is a sound source signal matrix, and N is a Gaussian white noise, X, N ∈ C M×L , S ∈ C K×L , and L is the number of snapshots.
[0104] Based on the discretized signal data, the covariance matrix is obtained by the following formula:
[0105] R = X(t)X H (t)
[0106] wherein R is a covariance matrix, X H is a conjugate transpose, and t is a time.
[0107] S12, the covariance matrix is obtained based on the discretized signal data.
[0108] The discretized signal data refers to a set of data obtained after adding noise and parameter discretization processing, and each data sample corresponds to a specific discretized parameter combination (such as a specific azimuth angle and a specific signal-to-noise ratio) under the noise-added array manifold vector or matrix.
[0109] The covariance matrix is a square matrix, and its elements represent the covariance between the components of a random vector. In array signal processing, it can be estimated from multiple snapshot data to represent the spatial correlation of the signal. For example, for an 8-element array, its covariance matrix is an 8x8 complex matrix, and the diagonal elements are the powers of the signals received by each element, and the non-diagonal elements are the correlations between different element signals.
[0110] For each discrete parameter combination, the generated noisy array response (usually treated as one snapshot) is used to estimate the theoretical covariance matrix under that condition by computing its outer product (or the product with the conjugate transpose).
[0111] S13, inputting the covariance matrix as the sample input data and generating sample labels by a hybrid labeling strategy to construct the training dataset.
[0112] The hybrid labeling strategy is a method of generating sample labels, which combines multiple sources or multiple types of labeling information, possibly including hard labels (such as one-hot encoding), soft labels (such as probability distribution), or labeling fusion from different sensors to provide richer and more robust supervision information. For example, in a localization task, the label of a sample not only includes the true azimuth value (hard label), but also may include the angle ambiguity information (soft label, such as an angle probability distribution) when the multipath effect exists at this position according to the sound field model. The execution process of the hybrid labeling strategy can be referred to the description of the generation of sample labels in Figure 2 .
[0113] The sample label is the true value or target output corresponding to the sample input data, used to guide model training in supervised learning. For example, for a localization task, the sample label can be a three-dimensional vector including azimuth, distance, and depth.
[0114] The training dataset is a collection of data used to train machine learning models, consisting of input samples and corresponding labels. For example, the final constructed training dataset may include tens of thousands of samples, each sample's input is a covariance matrix, and the label is a vector including azimuth, distance, and depth, as well as label information.
[0115] Step 104, training the multi-task network model using the training dataset to obtain a trained multi-task network model.
[0116] Training refers to the process of adjusting neural network model parameters using training data sets through optimization algorithms to minimize prediction errors, so that the model can learn the rules from the input data. For example, the training process usually involves forward propagation to calculate the output, calculate the loss function, and update the weights through back propagation.
[0117] The trained multi-task network model is the model obtained after the training process, whose parameters (such as weights and biases) have been optimized and can accurately predict new input data while outputting the results of multiple tasks. For example, a trained multi-task CNN model can output the estimated values of azimuth, distance, and depth simultaneously after inputting an array signal.
[0118] In some embodiments, the input samples (array received signal data) in the training dataset can be batched into the multi-task network model, the output values of each task are calculated by forward propagation, then compared with the true labels, the total loss is calculated using a loss function (such as mean square error for regression tasks) (possibly weighted sum of the losses of multiple tasks), then the gradients are calculated by the back propagation algorithm, and the model parameters are updated using an optimizer (such as Adam or SGD). The training process is iterated multiple times (epochs) until the loss converges or a stopping condition is reached. The validation set is used to monitor the generalization performance and prevent overfitting. For example, 8000 samples in the training dataset are used for training, the batch size is 32, and the training is performed for 100 epochs. The validation set loss is calculated after each epoch, and the final model that performs best on the validation set is selected as the trained multi-task network model.
[0119] Step 105, using the trained multi-task network model to locate the underwater target; wherein the positioning result includes the azimuth, distance and sound source depth of the underwater target.
[0120] The underwater target refers to an underwater object or sound source that needs to be located, such as a submarine, an underwater drone or a marine organism, which emits sound waves or is detected by sound waves. For example, a submarine sailing underwater, its engine produces noise, which becomes a positioning target.
[0121] Positioning is the process of determining the position of a target in three-dimensional space, for example, estimating the azimuth, distance and depth of the target. For example, by analyzing the acoustic signal, it is determined that the target is in the northeast direction (azimuth 45 degrees), 2000 meters away, and 100 meters deep.
[0122] In some embodiments, the array received signal data (after the same preprocessing as the training data) collected in real time or in testing can be input into the trained multi-task network model, and the model automatically calculates and outputs the predicted values of the azimuth, distance and sound source depth through forward propagation. These outputs are the positioning results, which can be used for display or further processing. For example, in actual application, a hydrophone array is deployed in the ocean to continuously monitor acoustic signals; when a potential target signal is detected, a segment of signal data is extracted and input into the trained multi-task CNN model; the model outputs an azimuth of 55 degrees, a distance of 1800 meters, and a depth of 120 meters, thereby completing the three-dimensional positioning of the underwater target.
[0123] Figure 2 is a multi-task network model training method shown in some embodiments of the present application. In some embodiments, the embodiment can include the following operations:
[0124] Step 201, input the sample data in the training dataset into the multi-task network model to be trained to obtain the predicted value.
[0125] The multi-task network model to be trained refers to a neural network model that has not completed parameter optimization and needs to be trained and learned. The initial values of the parameters (such as weights and biases) of the multi-task network model to be trained are usually randomly set. For example, a CNN model with just initialized weights has a structure including an input preprocessing module, a common feature extraction layer, a feature fusion module, and a multi-task output module.
[0126] The predicted value refers to an output result calculated by the model after inputting sample input data into the multi-task network model to be trained, which is a current estimation of the target value by the model.
[0127] The input data part (for example, the preprocessed covariance matrix tensor) of the sample is taken out from the training data set in batches or one by one, and is input into the current state of the multi-task network model. The model calculates the predicted values of the azimuth, the distance, and the depth of the sound source through forward propagation calculation and processing of each network layer (such as a convolutional layer and a fully connected layer) according to the current internal parameters (weights and biases) of the model. For example, 32 samples are taken out from the training data set, and the input data of each sample is a 2x8x8 tensor (representing the real part and the imaginary part of the covariance matrix). The batch of data is input into the model to be trained, and the model performs forward calculation and outputs a 32x3 matrix, wherein each row includes three predicted values (azimuth, distance, and depth) of a sample.
[0128] In some embodiments, the training stage can adopt double-layer batch processing: first, read and cache in disk batches with a data block size of no more than 100 samples, and then send into the network forward and backward propagation in GPU small batches with a size of no more than 32 samples allowed by the video memory; immediately release the cache after completing a disk batch to reduce the peak memory occupation. The peak video memory and system memory are reduced to less than 10% of the original full load, which can prevent memory explosion caused by one-time loading and is compatible with low video memory devices.
[0129] In some embodiments, knowledge distillation can also be used to replace direct quantization training, for example, a large teacher model is first trained, and then distilled into a small student model, which maintains the accuracy while significantly reducing the parameters. After distillation, 8-bit quantization can be performed to further compress the MCU and other extremely low-power hardware.
[0130] In some embodiments, model pruning combined with low-rank decomposition can also be used to replace direct quantization training, for example, channel / structure pruning is performed on the convolution kernel, or the weights are decomposed into low-rank matrices.
[0131] In step 202, based on the predicted values and sample labels corresponding to the sample input data, an azimuth index, a distance index, and a depth index are calculated to evaluate the performance of the model.
[0132] In some embodiments, the sample label can be obtained by the following way:
[0133] S20, generating an initial label corresponding to the array receiving signal data based on the sound field model.
[0134] The initial label refers to the target parameter value generated directly by a theoretical model or simulation calculation without correction, which serves as the preliminary version or basis of the sample label. For example, in a simulation environment, if the sound source is set to be located at an azimuth angle of 30 degrees, a distance of 2000 meters, and a depth of 100 meters, then the initial label corresponding to the array receiving signal data generated for this position is [30, 2000, 100].
[0135] In some embodiments, the sound source position parameters used when generating the array receiving signal data can be recorded at the same time as the data is simulated using the sound field model. These parameters, i.e., the azimuth angle, the distance, and the sound source depth, are directly extracted as the initial label corresponding to this set of array receiving signal data.
[0136] S21, obtaining a real label corresponding to the measured array data.
[0137] The measured array data refers to the sound signal data directly measured and recorded by the actually deployed hydrophone array in a real ocean environment or experimental pool. The measured array data includes all physical effects and noise in the actual environment, and is opposite to the simulation-generated array receiving signal data.
[0138] The real label refers to the reference value of the real position parameter of the measured target obtained directly by high-precision measurement means, which has a much higher precision than the expected precision of the model to be trained, and is usually used as the ground truth for evaluating the performance of the model.
[0139] In some embodiments, a high-precision external measurement system can be used simultaneously to measure and record the real position of the target sound source when collecting the measured array data in actual sea experiments or pool experiments. For the azimuth angle and distance, optical measurement equipment such as GPS system, ultra-short baseline (USBL), or total station can be used to obtain the horizontal position of the sound source and convert it into the azimuth angle and distance relative to the receiving array. For the sound source depth, a pressure depth sensor can be used for direct measurement. The high-precision position parameters measured in this way are used as the real label and correspond to the measured array data collected at the same time.
[0140] S22, adjusting the initial label based on the real label to obtain the sample label.
[0141] Adjustment refers to the process of modifying or correcting the original data (here, initial labels) according to the reference standard (here, true labels), aiming to make the original data closer to the true situation or meet certain requirements.
[0142] In some embodiments, a certain amount of paired data can be collected first, i.e., for the same or very similar target positions, both the sound field model generated (array received signal data, initial labels) and the actually measured (actually measured array data, true labels) are available. By comparing and analyzing the initial labels and true labels in these paired data, systematic errors or deviation rules are found (for example, there is a fixed proportional error in the distance value of the initial label, or a fixed deviation in the azimuth angle). Then, according to the analyzed error rule, an adjustment model or correction function is established (for example, a simple linear scaling or offset). Finally, using this adjustment model, all the initial labels generated by the sound field model are batch corrected to obtain effective sample labels that are closer to the true situation and can be used for final training.
[0143] Azimuth indicator is an evaluation standard for quantifying the performance of model azimuth angle prediction, usually a scalar value, the smaller the value, the better the performance. For example, mean absolute error (MAE), which is the average of the absolute values of the differences between the predicted azimuth angles and the true azimuth angles of all samples.
[0144] Distance indicator is an evaluation standard for quantifying the performance of model distance prediction, usually a scalar value, the smaller the value, the better the performance. For example, root mean square error (RMSE), which is the square root of the average of the squares of the differences between the predicted distances and the true distances of all samples.
[0145] Depth indicator is an evaluation standard for quantifying the performance of model sound source depth prediction, usually a scalar value, the smaller the value, the better the performance. For example, mean absolute error (MAE), which is the average of the absolute values of the differences between the predicted depths and the true depths of all samples.
[0146] Model performance refers to the degree of performance of a machine learning model on a specific task, which is measured by evaluation indicators.
[0147] In some embodiments, after obtaining the predicted values of a batch or an entire epoch, these predicted values can be compared with the corresponding true labels in the data set. For the azimuth angle task, the azimuth indicator (such as MAE) is used to calculate the azimuth angle predicted values and true values of all samples, and a numerical value is obtained as the azimuth indicator. Similarly, the distance indicator and the depth indicator are calculated using the calculation methods of the distance indicator and the depth indicator, respectively. The three indicators are calculated independently and are used to evaluate the performance of the model on the three different tasks.
[0148] For example, the mean absolute error is used to evaluate the error between the model predicted azimuth and the signal direction value and the true value, and its functional expression is shown in the following formula (6):
[0149]
[0150] where N is the number of samples, y i and respectively represent the label of the i-th sample and the predicted value of the model.
[0151] In some embodiments, a differentiated weighted joint loss function can also be used to evaluate the model prediction performance. The differentiated weighted joint loss function is shown in the following formula (7).
[0152]
[0153] where L is the total loss function, is the azimuth indicator loss function term, θ (0) respectively represent the corresponding model predicted value and sample label, is the distance indicator loss function term, R (0) respectively represent the corresponding model predicted value and sample label, is the depth indicator loss function term, Z (0) respectively represent the corresponding model predicted value and sample label, w θ , w R , w Z are the corresponding weight coefficients, and the weights satisfy w θ ≥ 0.5, w R = w Z , w θ + w R + w Z = 1, for example, 0.6:0.2:0.2.
[0154] Step 203, if any of the azimuth indicator, distance indicator and depth indicator exceeds the preset threshold value, adjusting the model parameters of the multi-task network model based on the evaluation result.
[0155] The preset threshold is a numerical limit set in advance, which is used to judge whether the model performance is acceptable or whether it needs to be further adjusted. If the evaluation indicator exceeds (is greater than) this threshold, it is considered that the performance is not up to standard.
[0156] The evaluation result refers to the model performance quantitative conclusion obtained by calculating the azimuth indicator, distance indicator and depth indicator.
[0157] Model parameters is a broad term and here refers specifically to settings that need to be optimized or changed during the training process, including network structure and / or hyperparameters.
[0158] Network structure refers to the organization of the neural network, including the type of layers (such as convolutional layers, fully connected layers), the number of layers, the number of neurons in each layer, the connection method, etc. For example, increasing the residual block of the public feature extraction layer from 3 to 4, or adding a fully connected layer before each regression task in the multi-task output module.
[0159] In some embodiments, three performance indicators can be calculated at each evaluation stage (such as at the end of each epoch) and compared with their respective preset thresholds. If the value of any one indicator (azimuth, range, or depth) exceeds the threshold set for it, it indicates that the performance of the model on that corresponding task has not yet met the requirements. At this time, the parameters of the model need to be adjusted according to the problem reflected by the evaluation results (which task is poor in performance). Adjusting the network structure may mean modifying the number, type, or size of the layers (for example, adding more layers in the output module for the task with poor performance). Adjusting the hyperparameters may mean changing the learning rate, optimizer settings, or regularization parameters, etc. After adjustment, the training process will usually continue or restart in order to improve performance in subsequent iterations.
[0160] Compared with the existing three-subarray, TMA, MFP, positioning using single-task convolutional network, and multi-task learning scheme based on vertical array, the underwater target three-dimensional positioning network of the multi-task horizontal line array of the application realizes the synchronous output of three-dimensional position parameters through the parallel setting of azimuth, range, and sound source depth regression heads after sharing the residual backbone in a single forward operation, significantly shortens the inference delay and reduces repeated feature calculation; at the same time, multi-task joint training excavates the physical correlation between each target parameter, and can still control the mean square error of the azimuth angle to be within 1.2°, the horizontal distance error to be within 50m, and the depth error to be within 5m under low signal-to-noise ratio conditions, the precision and robustness are better than the prior art; in addition, parameter sharing reduces the overall model size by about 50% compared with three independent networks, reduces the deployment algorithm and storage demand, has higher real-time performance, generalization, and embedded application feasibility, thereby having good application prospect in the field of underwater passive sonar target positioning.
[0161] Based on the same inventive concept, the embodiments of the application also provide an underwater target positioning system based on multi-task learning, the system comprising:
[0162] A data generation module for generating array received signal data including a multi-path array signal model based on a sound field model;
[0163] A model construction module for constructing a multi-task network model using a convolutional neural network;
[0164] a training data construction module, configured to construct a training data set based on the array receiving signal data, the training data set comprising sample data of different azimuths, distances and sound source depths;
[0165] a training module, configured to train the multi-task network model using the training data set to obtain a trained multi-task network model;
[0166] a positioning module, configured to position an underwater target using the trained multi-task network model; wherein the positioning result comprises an azimuth, a distance and a sound source depth of the underwater target.
[0167] As shown in Figure 3 Embodiments of the present application provide an electronic device, comprising a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112 and the memory 113 complete mutual communication through the communication bus 114,
[0168] the memory 113 is configured to store a computer program;
[0169] In an embodiment of the present application, the processor 111 is configured to execute the program stored in the memory 113 to implement the method for underwater target positioning based on multi-task learning provided by any one of the preceding method embodiments, comprising:
[0170] generating array receiving signal data comprising a multi-path array signal model based on a sound field model;
[0171] constructing a multi-task network model using a convolutional neural network;
[0172] constructing a training data set based on the array receiving signal data, the training data set comprising sample data of different azimuths, distances and sound source depths;
[0173] training the multi-task network model using the training data set to obtain a trained multi-task network model;
[0174] positioning an underwater target using the trained multi-task network model; wherein the positioning result comprises an azimuth, a distance and a sound source depth of the underwater target.
[0175] Embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method for underwater target positioning based on multi-task learning provided by any one of the preceding method embodiments.
[0176] It has to be noted that, in the present document, relational terms are intended only to convey a possible relationship between elements or
[0177] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, which modifications and changes are to be understood as intended to be encompassed by the general scope of the application. Accordingly, the application is not to be limited to the above described or illustrated embodiments, but is intended to encompass all embodiments consistent with the principles of the application.
Claims
1. An underwater target localization method based on multi-task learning, characterized in that, The method includes: Array received signal data, including a multipath array signal model, is generated based on the acoustic field model; Constructing a multi-task network model using convolutional neural networks; A training dataset is constructed based on the signal data received by the array, and the training dataset includes sample data with different orientations, distances, and sound source depths. The multi-task network model is trained using the training dataset to obtain the trained multi-task network model. The trained multi-task network model is used to locate underwater targets; wherein the location results include the azimuth, distance and sound source depth of the underwater target.
2. The method according to claim 1, characterized in that, The array receives signal data in the following manner: Where x(t)k is the array received signal data at time t, k is the number of sound sources, p is the number of multipath reflections, P = 0 represents the direct wave, P ≥ 1 represents the Pth reflection path, and α k,p Let be the complex amplitude attenuation factor of the p-th reflection path of the k-th sound source. Let s be the incident angle of the p-th reflection path of the k-th sound source. k (t) represents the sound source signal. Let m be the guiding vector of a uniform linear array, and its m-th component be: [a(i)] m =exp(j2πf0τ m ) Where, τ m This represents the phase difference between the m-th array source and the reference array element in a uniform linear array.
3. The method according to claim 2, characterized in that, The construction of the training dataset based on the received signal data from the array includes: Based on the received signal data from the array, the array manifold matrix is obtained; Gaussian white noise is added to the array manifold matrix and the signal is discretized; The covariance matrix is obtained based on the data after signal discretization. The training dataset is constructed by using the covariance matrix as the sample input data and generating sample labels through a hybrid labeling strategy.
4. The method according to claim 3, characterized in that, The addition of Gaussian white noise to the array manifold matrix is represented as follows: x(t)=A eff s(t)+n(t) Among them, A eff =[A0A1…A P ]∈C M×K(P+1) Let s(t) = [s1(t), ..., s2(t)] be an array manifold matrix. K (t)] t Let n(t) be the signal from each sound source, and n(t) be Gaussian white noise. The signal is discretized using the following formula: X=A eff S+N Where X is the discretized received data matrix, and A eff Let S be the array manifold matrix, S be the sound source signal matrix, N be Gaussian white noise, and X, N ∈ C. M×L S∈C K×l L represents the number of snapshots; Based on the discretized signal data, the covariance matrix is obtained using the following formula: R=X(t)X H (t) Where R is the orthometric matrix, X H For the conjugate transpose, t is time.
5. The method according to claim 4, characterized in that, The multi-task network model includes an input preprocessing module, a common feature extraction layer, a feature fusion module, and a multi-task output module. The input preprocessing module is used to receive the input covariance matrix, and then normalize the real and imaginary parts of R separately before concatenating them according to the input channel dimension. The common feature extraction layer comprises a first residual block, a second residual block, and a third residual block connected in sequence, and is used to extract features based on the output of the input preprocessing module. The feature fusion module is used to flatten the feature map output by the common feature extraction layer into a one-dimensional vector and then perform dimension mapping through a three-layer fully connected network. The multi-task output module is used to output the predicted azimuth, distance, and sound source depth based on the output of the feature fusion module through three parallel regression tasks.
6. The method according to claim 3, characterized in that, The step of training the multi-task network model using the training dataset includes: Input the sample input data from the training dataset into the multi-task network model to be trained to obtain the predicted value; Based on the predicted values and the sample labels corresponding to the sample input data, the orientation index, distance index, and depth index are calculated to evaluate the model performance. If any of the orientation, distance, and depth indicators exceeds a preset threshold, the model parameters of the multi-task network model are adjusted based on the evaluation results; wherein, the model parameters include network structure and / or hyperparameters.
7. The method according to claim 6, characterized in that, The sample labels were obtained in the following way: Initial tags corresponding to the signal data received by the array are generated based on the sound field model; Obtain the actual labels corresponding to the measured array data; The initial label is adjusted based on the real label to obtain the sample label.
8. An underwater target localization system based on multi-task learning, characterized in that, The system includes: The data generation module is used to generate array received signal data, including a multipath array signal model, based on the sound field model; The model building module is used to build multi-task network models using convolutional neural networks; The training data construction module is used to construct a training dataset based on the signal data received by the array. The training dataset includes sample data with different orientations, distances, and sound source depths. The training module is used to train the multi-task network model using the training dataset to obtain the trained multi-task network model. The localization module is used to locate underwater targets using the trained multi-task network model; wherein the localization result includes the azimuth, distance and sound source depth of the underwater target.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the steps of the underwater target localization method based on multi-task learning as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater target localization method based on multi-task learning as described in any one of claims 1-7.