A preprocessing method for underwater direction of arrival estimation
By improving the BERT model to pre-treat underwater array signals, improve the signal-to-noise ratio of the signal and estimate the sound speed, the problem of the existing DOA estimation method degradation in underwater environments is solved, and higher estimation accuracy is achieved.
Patent Information
- Application Number
- CN202210615595.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The existing DOA estimation method has a degraded estimation performance in underwater ambient sound velocity uncertainty and low signal-to-noise ratio environments, making it difficult to effectively overcome the impact of noise and sound velocity on estimation.
The improved BERT deep learning model is used as the denoising autoencoder to preprocess the array signal received data, improve the signal-to-noise ratio of the signal and estimate the sound speed, thereby reducing the impact of noise and sound speed on DOA estimation.
By improving the preprocessing of the BERT model, the signal-to-noise ratio of the array signal received data is improved, the impact of noise and sound speed on DOA estimation is reduced, and the estimation accuracy of the existing DOA estimation methods in underwater environments is significantly improved.
Smart Images

Figure CN115186697B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of target positioning, and in particular to an underwater direction-of-arrival estimation preprocessing method. Background Art
[0002] Estimation of the direction of arrival (DOA) of spatial signals is one of the important research directions in the field of array signal processing. The spatial signals are received by the sensor array, and useful signal features and information contained in the signals are extracted to estimate the azimuth of the target. Common DOA estimation methods include: MUSIC algorithm, ESPRIT algorithm, propagation operator method, sparse array DOA estimation algorithm, etc. These algorithms have good results in a good channel environment.
[0003] Underwater DOA estimation is a common task in marine exploration, underwater sonar systems, etc. In the marine environment, there are various marine noises, such as natural environmental noise, marine biological noise, etc. The signal-to-noise ratio of the underwater acoustic channel will be greatly reduced. Therefore, the estimation accuracy of the above algorithm will be reduced when applied underwater, and overcoming the influence of the signal-to-noise ratio has become an urgent problem to be solved. Secondly, when the above DOA estimation algorithm is applied underwater, the underwater sound speed is often calculated as a known constant. In the actual underwater environment, the sound speed is affected by factors such as temperature and salinity, and its value is constantly changing. Calculating the sound speed as a constant will lead to large estimation errors. Therefore, overcoming the influence of the sound speed is also a problem that needs to be solved. Summary of the invention
[0004] The purpose of the present invention is to solve the problem that the estimation performance of the existing DOA estimation method is degraded in the environment of uncertain underwater sound speed and low signal-to-noise ratio, and to provide an underwater direction of arrival estimation preprocessing method. The method improves the BERT deep learning model, uses the improved BERT as a denoising autoencoder to preprocess the array signal receiving data, improves the signal-to-noise ratio of the output signal, and estimates the sound speed at the same time. The denoised received data obtained after preprocessing and the estimated sound speed are used as the input of the existing DOA estimation method, which can reduce the influence of noise and sound speed on DOA estimation and improve the estimation performance of the existing DOA estimation method in the environment of uncertain underwater sound speed and low signal-to-noise ratio.
[0005] The purpose of the present invention can be achieved by adopting the following technical solutions:
[0006] A method for preprocessing underwater direction of arrival estimation, the method comprising the following steps:
[0007] S1. Establishing an array signal model of a one-dimensional uniform linear array;
[0008] S2. Build an improved BERT model;
[0009] S3. Train and improve the BERT model;
[0010] S4. Reconstruct the received data X of the array signal to obtain the noisy matrix Z, and use the improved BERT model to preprocess the noisy matrix Z to obtain the denoising estimation matrix With estimated speed of sound Then estimate the denoising matrix Reconstruct the array signal to get the denoised received data
[0011] Furthermore, the process of step S1 is as follows:
[0012] The uniform linear array model scenario is shown in the attached figure. Figure 1 As shown in the figure, it is assumed that there are M array elements distributed on the uniform linear array, the array element spacing is d, and there are K far-field narrow-band acoustic wave signals as target signals. The i-th target signal is recorded as s i (t), i = 1, 2, ..., K, incident on the uniform linear array at the same time, and satisfying K < M, the center frequency of the signal is f, the speed of sound is c, and the wave arrival angle of the i-th target signal is θ i ,
[0013] The signal x1(t) received by the first array element is:
[0014]
[0015] Among them, s i (t) represents the i-th target signal, n1(t) represents the noise on the first array element channel;
[0016] The time delay τ of the i-th target signal between the first array element and the m-th array element i for:
[0017]
[0018] The signal x received by the mth array element at the same time m (t) is:
[0019]
[0020] n m (t) represents the noise on the mth array element channel, then the signal X(t) received by M array elements at the same time is represented by the vector:
[0021] X(t)=AS(t)+N(t) (Formula 4)
[0022] The received signal is X(t)=[x1(t),x2(t),…,x M (t)]T , superscript T represents the transpose operation, and the flow matrix A is:
[0023]
[0024] S(t)=[s1(t),s2(t),…,s K (t)] T is the target signal, N(t)=[n1(t),n2(t),…,n M (t)] T for noise;
[0025] After L snapshots, the array signal receiving data is X = [x1, x2, ..., x m …,x M ] T , the received noise is N = [n1,n2,…,n m ,…n M ] T , where the received data of the mth array element is x m =[x m (1),x m (t),…,x m (L)] T , the received noise is n m =[n m (1),n m (t),…,n m (L)] T .
[0026] Furthermore, the process of step S2 is as follows:
[0027] Improve the structure of the BERT model during the training phase as shown in the attached Figure 2 As shown in the attached figure, the structure in the prediction stage is Figure 3 As shown, the dotted part is the main improved part; the improved BERT model includes an input layer, an embedding layer, an encoder and an output layer connected in sequence, wherein,
[0028] The input layer is a noisy matrix Z = [z1, z2, ..., z m ,...,z M ], the dimension of the noise matrix Z is 2L×M, where the vector z m is the mth column vector, and the noisy matrix Z is reconstructed from the array signal receiving data X:
[0029]
[0030] Wherein, Re{X} and Im{X} are the real part and imaginary part of the array signal receiving data X respectively;
[0031] Considering that the power of the received signal will decrease due to attenuation in practice, the data will be standardized before being input into the embedding layer. The noisy matrix Z is normalized to obtain a normalized matrix E = [e1,e m ,…,e M ], where the vector e m is the mth column vector, and the normalized matrix E is expressed as:
[0032]
[0033] in, is the median of the elements in the noisy matrix Z, is half the length of the interval where the elements in the noisy matrix Z are located, v max is the maximum value of the elements in the noisy matrix Z, v min is the minimum value of the elements in the noisy matrix Z, and the matrix G represents a matrix whose elements are all 1; the normalized matrix E will be input into the embedding layer;
[0034] The embedding layer is composed of the input encoding E in , fragment code E seg and position code E pos Add together;
[0035] The input of the embedding layer is the normalized matrix E, and the sound speed code z C , signal-to-noise ratio coding z SNR , classification code z CLS Construct an input code E of dimension 2L×(M+3) in , enter code E in It is expressed as:
[0036] E in =[z C ,z SNR ,z CLS ,E] (Formula 8)
[0037] The speed of sound is encoded as z C , signal-to-noise ratio coding z SNR , classification code z CLS , are column vectors of length 2L, randomly initialized, and updated as parameters during training;
[0038] Fragmentation Code E seg The dimension of the input layer is 2L×(M+3). When the input noise matrix Z=[z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from the same L snapshots, the slice encoding E seg The column vectors of E are all column vectors with all elements equal to 0. A =[0,...,0] T ; When the input noise matrix Z of the input layer is [z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from data obtained from different L snapshots, the slice encoding E seg Before Column vectors are all column vectors E whose elements are all 0 A =[0,...,0] T ,back Column vectors are all column vectors E whose elements are all 1 B =[1,...,1] T ;
[0039] Positional Encoding pos The dimension is 2L×(M+3), and its i-th row and j-th column element is expressed as:
[0040]
[0041] Among them, floor(·) is a floor function; the role of position encoding is to mark the relative position relationship of the input sequence represented by each array element. Because the encoder of the BERT model does not have a sequential relationship in processing, it is necessary to add position encoding to the sequence;
[0042] The output E of the embedding layer out The dimension of the embedding layer is 2L×(M+3), and the output E out It is expressed as:
[0043] E out =E in +E seg +E pos (Formula 10)
[0044] The output of the embedding layer E out inputting into the encoder;
[0045] The encoder includes N encoder layers connected in sequence; each encoder layer includes four sublayers connected in sequence, namely a multi-head attention layer, a layer normalization layer, a fully connected feedforward network layer and a layer normalization layer, and the input and output dimensions of each sublayer are 2L×(M+3); the front and rear ends of the multi-head attention layer and the fully connected feedforward network layer both use residual connections;
[0046] The fully connected feedforward network consists of two linear transformations connected in sequence, with a ReLU(·) activation between the two linear transformations; the fully connected feedforward network is applied to the input E of this layer in the same way LN =[e LN,1 ,e LN,m ,...,e LN,M+3 ], e LN,m Enter E for this layer LN The mth column vector of the fully connected feedforward network, the mth column output e FFN,m It is expressed as:
[0047] e FFN,m =W2ReLU(W1e LN,m +b1)+b2 (Formula 11)
[0048] Among them, W1 is the weight of the first linear transformation, the dimension is 8L×2L, b1 is the bias term of the first linear transformation, the dimension is 8L×1, W2 is the weight of the second linear transformation, the dimension is 2L×8L, b1 is the bias term of the second linear transformation, the dimension is 2L×1;
[0049] The output of the encoder is where the vector z′ C The vector z′ is the sound speed estimation code. SNR For the signal-to-noise ratio estimation code, the vector z′ CLS For classification estimation coding, both are column vectors of length 2L, and the matrix is the normalized denoising estimation matrix, with a dimension of 2L×M, and the vector is the normalized denoising estimation matrix The mth column vector of ;
[0050] During the training phase, the output layer includes two parallel linear regressors, a linear classifier and an affine layer; the sound speed estimation is encoded Coding with SNR estimation Input into two linear regressors respectively to get the estimated sound speed Estimated signal-to-noise ratio Encoding categorical estimates Input to a linear classifier to predict the relationship between two sets of sequences; normalize the denoised estimation matrix Input to the affine layer to get the denoising estimation matrix Denoising Estimation Matrix It is expressed as:
[0051]
[0052] Among them, the denoising estimation matrix is The matrix size is 2L×M, and the vector For the matrix The mth column vector of ;
[0053] In the prediction stage, the output layer includes a linear regressor and an affine layer in parallel; the sound speed estimate is encoded Input to the linear regressor to get the estimated sound speed Normalize the denoised estimate matrix Input to the affine layer to get the denoising estimation matrix
[0054] Furthermore, the process of step S3 is as follows:
[0055] The data sets are divided into simulation data sets and measured data sets. The simulation data sets are based on different signal source numbers K, sound speed c, wave arrival direction angles θ1, θ2, …, θ K , signal-to-noise ratio snr, generate multiple sets of array signal receiving data X with white noise and its corresponding denoised receiving data X′, noisy matrix Z, denoising matrix Z′ to train the improved BERT model, where the array signal denoised receiving data X′ with dimension M×L is expressed as:
[0056] X′=AS (Formula 13) Where S=[S(1),S(t),…,S(L)] is the L snapshots of the target signal; denoising matrix Z′=[z′1,z′2,…,z′ m ,...,z′ M ] is the reconstruction matrix of the array signal denoising received data X′, with a dimension of 2L×M, and the vector z′ m is the mth column vector of the denoising matrix Z′, and the denoising matrix Z′ is expressed as:
[0057]
[0058] Among them, Re{X′} and Im{X′} are the real part and imaginary part of the array signal denoising received data X′ respectively; the simulation data set is used for pre-training the improved BERT model, and the measured data set is used for fine-tuning the improved BERT model; a large number of data sets are used to train the BERT model in the pre-training stage, and then only a small amount of data sets of the current task are needed to train the improved BERT model in the fine-tuning stage, which can achieve excellent results in the current task without using a large amount of data to train an improved BERT model suitable for the current task, thereby improving the applicability of the model; the training set and the test set are divided into 9:1;
[0059] The present invention uses four training objectives to supervise the training of the improved BERT model;
[0060] The first training goal is to predict the denoised signal sequence. The improved BERT model randomly masks 15% of the column vectors of the noisy matrix Z, and then predicts the column vectors of the denoised matrix Z′ corresponding to these masked column vectors, which is the usual cloze task. If the mth column vector z of the noisy matrix m is masked, then there is an 80% probability that the vector will be replaced by the masking vector z mask , mask vector z mask is the model parameter; there is a 10% probability of being replaced by a random vector z′ rand , which is a column vector of a random denoising matrix Z′ in the data set; there is a 10% probability of no replacement;
[0061] In the random masking process, the unmasked input is Z = [z1,z2,...,z m ,...,z M ], assuming we choose the column vector z m As the masking goal, our masking process is:
[0062] 80% probability: Change the column vector z m Replaced by the mask vector z mask , for example, [z1,z2,...,z m ,...,z M ]→[z1,z2,...,z mask ,...,z M ];
[0063] 10% probability: use a random vector z′ rand Replace the column vector z m , for example, [z1,z2,...,z m ,...,z M ]→[z1,z2,...,z′ rand ,...,z M ];
[0064] 10% probability: keep vector z m unchanged, for example, [z1,z2,...,z m ,...,z M ]→[z1,z2,...,z m ,...,z M ];
[0065] Mask the mth column vector z of the noisy matrix m After that, the mth column vector z′ of the corresponding denoising matrix will be predicted m ; The prediction value of improved BERT is The optimization goal is to minimize the mth column vector z′ of the denoising matrix mWith the predicted value The squared error between:
[0066]
[0067] The second training objective is post-order array prediction; in natural language processing tasks, this training is used to capture the relationship between two sentences; in the present invention, in order to train a model that can capture the relationship between the received data of two sets of arrays, we train a binary post-order array prediction task; specifically, when selecting the previous noisy matrix Z for each sample Column vector and When the column vector is 50%, the probability of the noise matrix Z is The column vector is the actual posterior Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from the same L snapshots of data, with the sample label marked as 1; the posterior of the noisy matrix Z is 50% The column vector is a random segment of the array signal in the data set. The noisy matrix Z of the received data X rand After Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from data obtained from L different snapshots, and the sample labels are marked as -1; finally, the classification estimation encoding output by the BERT model encoder is improved The input is fed into a binary classifier for subsequent array prediction;
[0068] The following two examples can illustrate the post-order array prediction task;
[0069] There are two noisy matrices Z1 = [z 1,1 ,z 1,2 ,...,z 1,m ,...,z 1,M ],Z2=[z 2,1 ,z 2,2 ,...,z 2,m ,...,z 2,M ],z i,j is the j-th column vector of the i-th noisy matrix;
[0070] Input layer input: [z 1,1 ,z 1,2 ,...,z 1,m ,...,z 1,M ]; Tags: 1.
[0071] Input layer input: Tags: -1.
[0072] The label of sample Z is y∈{-1,1}, and the weight of the binary classifier is w CLS , the bias term is b CLS , the predicted value is sign(·) is the sign function, · is the vector dot product, and its optimization goal is to minimize the label y of sample Z and the predicted value of the binary classifier. The squared error between:
[0073]
[0074] Training goal three is to predict the speed of sound; sample Z is obtained in an environment with a speed of sound of c, and the speed of sound estimation code output by the improved BERT encoder is Input to a linear regressor with weight w C , the bias term is b C , the estimated speed of sound is The optimization goal is to minimize the sound speed c and the estimated sound speed The squared error between:
[0075]
[0076] Training goal 4: SNR prediction; the SNR of sample Z is snr, and the SNR estimate output by the improved BERT encoder is encoded Input to the second linear regressor, the weight of the second linear regressor is w SNR , the bias term is b SNR , the estimated signal-to-noise ratio is The optimization goal is to minimize the signal-to-noise ratio SNR and the estimated signal-to-noise ratio The squared error between:
[0077]
[0078] Supervised training is performed based on four training objectives until the model converges on the test set to obtain a trained improved BERT model.
[0079] Furthermore, the process of step S4 is as follows:
[0080] The noisy array signal receiving data X is reconstructed to obtain the noisy matrix Z, and then the noisy matrix Z is input into the improved BERT model to obtain the output denoising estimation matrix and estimate the speed of sound Note the denoising estimation matrix The matrix composed of the first L row vectors is The matrix composed of L row vectors is Array signal denoising received data The denoising matrix is estimated Reconstructed, where the array signal denoises the received data The real part of for:
[0081]
[0082] Imaginary part for:
[0083]
[0084] Finally, the array signal denoising received data is obtained and estimate the speed of sound The above flow chart can be obtained from the attached Figure 4 express.
[0085] Compared with the prior art, the present invention has the following advantages and effects:
[0086] 1. The present invention utilizes the advantages of the improved BERT model in capturing the timing and spatial sequence characteristics of the received signal in a uniform linear array, extracts the signal component in the received data, reduces the noise component, and improves the signal-to-noise ratio of the array signal received data after preprocessing. It has good noise reduction performance in a low signal-to-noise ratio environment and can improve the estimation accuracy of the existing underwater DOA estimation algorithm.
[0087] 2. The present invention uses an improved BERT model to estimate the sound speed. In an environment where the underwater sound speed is uncertain, compared with the DOA estimation algorithm that does not perform sound speed estimation, the present invention can improve the estimation accuracy of the existing underwater DOA estimation algorithm and reduce the impact of the underwater variable sound speed environment on the existing underwater DOA estimation algorithm.
[0088] 3. Compared with the traditional DOA estimation algorithm, the computational complexity and complexity of the estimation preprocessing of the present invention only increase the computational complexity and complexity of the improved BERT model, thereby ensuring the feasibility of the estimation preprocessing. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0090] Figure 1 is a schematic diagram of a uniform line model scene of the method of the present invention;
[0091] Figure 2 It is a schematic diagram of the structure of the improved BERT model used in the present invention during the training process;
[0092] Figure 3It is a schematic diagram of the structure of the improved BERT model used in the present invention in the prediction process;
[0093] Figure 4 It is a flow chart of the underwater direction of arrival estimation preprocessing method disclosed in the present invention. DETAILED DESCRIPTION
[0094] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0095] Example 1
[0096] As attached Figure 4 This embodiment discloses a method for preprocessing underwater direction of arrival estimation, and the preprocessing method comprises the following steps:
[0097] S1. Establish an array signal model of a one-dimensional uniform linear array.
[0098] The uniform linear array model scenario is shown in the attached figure. Figure 1 As shown in the figure, it is assumed that there are M array elements distributed on the uniform linear array, the array element spacing is d, and there are K far-field narrow-band acoustic wave signals as target signals. The i-th target signal is recorded as s i (t), i = 1, 2, ..., K, incident on the uniform linear array at the same time, and satisfying K < M, the center frequency of the signal is f, the speed of sound is c, and the wave arrival angle of the i-th target signal is θ i ,
[0099] The signal x1(t) received by the first array element is:
[0100]
[0101] Among them, s i (t) represents the i-th target signal, n1(t) represents the noise on the first array element channel;
[0102] The time delay τ of the i-th target signal between the first array element and the m-th array element i for:
[0103]
[0104] The signal x received by the mth array element at the same time m (t) is:
[0105]
[0106] n m (t) represents the noise on the mth array element channel, then the signal X(t) received by M array elements at the same time is represented by the vector:
[0107] X(t)=AS(t)+N(t) (Formula 4)
[0108] The received signal is X(t)=[x1(t),x2(t),…,x M (t)] T , superscript T represents the transpose operation, and the flow matrix A is:
[0109]
[0110] S(t)=[s1(t),s2(t),…,s K (t)] T is the target signal, N(t)=[n1(t),n2(t),…,n M (t)] T for noise;
[0111] After L snapshots, the array signal receiving data is X = [x1, x2, ..., x m …,x M ] T , the received noise is N = [n1,n2,…,n m ,…n M ] T , where the received data of the mth array element is x m =[x m (1),x m (t),…,x m (L)] T , the received noise is n m =[n m (1),n m (t),…,n m (L)] T .
[0112] S2. Build an improved BERT model.
[0113] Improve the structure of the BERT model during the training phase as shown in the attached Figure 2 As shown in the attached figure, the structure in the prediction stage is Figure 3 As shown, the dotted part is the main improved part; the improved BERT model includes an input layer, an embedding layer, an encoder and an output layer connected in sequence, wherein,
[0114] The input layer is a noisy matrix Z = [z1, z2, ..., z m ,...,z M ], the dimension of the noise matrix Z is 2L×M, where the vector z m is the mth column vector, and the noisy matrix Z is reconstructed from the array signal receiving data X:
[0115]
[0116] Where Re{X} and Im{X} are the real and imaginary parts of the array signal receiving data X respectively; X = [x1, x2, …, x m …,x M ] T The dimension is M×L, Re{X} T and Im{X} T The dimensions of are all L×M, so The dimension is 2L×M.
[0117] Normalize the noisy matrix Z to obtain a normalized matrix E=[e1,e m ,…,e M ], where the vector e m is the mth column vector, and the normalized matrix E is expressed as:
[0118]
[0119] in, is the median of the elements in the noisy matrix Z, is half the length of the interval where the elements in the noisy matrix Z are located, v max is the maximum value of the elements in the noisy matrix Z, v min is the minimum value of the elements in the noisy matrix Z, and the matrix G represents a matrix whose elements are all 1; the normalized matrix E will be input into the embedding layer;
[0120] The embedding layer is composed of the input encoding E in , fragment code E seg and position code E pos Add together;
[0121] The input of the embedding layer is the normalized matrix E, and the sound speed code z C , signal-to-noise ratio coding z SNR , classification code z CLS Construct an input code E of dimension 2L×(M+3) in , enter code E in It is expressed as:
[0122] E in =[z C ,zSNR ,z CLS ,E] (Formula 8)
[0123] The speed of sound is encoded as z C , signal-to-noise ratio coding z SNR , classification code z CLS , are column vectors of length 2L, randomly initialized, and updated as parameters during training;
[0124] Fragmentation Code E seg The dimension of the input layer is 2L×(M+3). When the input noise matrix Z=[z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from the same L snapshots, the slice encoding E seg The column vectors of E are all column vectors with all elements equal to 0. A =[0,...,0] T ; When the input noise matrix Z of the input layer is [z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from data obtained from different L snapshots, the slice encoding E seg Before Column vectors are all column vectors E whose elements are all 0 A =[0,...,0] T ,back Column vectors are all column vectors E whose elements are all 1 B =[1,...,1] T ;
[0125] Positional Encoding pos The dimension is 2L×(M+3), and its i-th row and j-th column element is expressed as:
[0126]
[0127] Where floor(·) is the floor rounding function;
[0128] The output E of the embedding layer out The dimension of the embedding layer is 2L×(M+3), and the output E out It is expressed as:
[0129] E out =E in +E seg +E pos (Formula 10)
[0130] The output of the embedding layer E out inputting into the encoder;
[0131] The encoder includes N encoder layers connected in sequence; each encoder layer includes four sublayers connected in sequence, namely a multi-head attention layer, a layer normalization layer, a fully connected feedforward network layer and a layer normalization layer, and the input and output dimensions of each sublayer are 2L×(M+3); the front and rear ends of the multi-head attention layer and the fully connected feedforward network layer both use residual connections;
[0132] The fully connected feedforward network consists of two linear transformations connected in sequence, with a ReLU(·) activation between the two linear transformations; the fully connected feedforward network is applied to the input E of this layer in the same way LN =[e LN,1 ,e LN,m ,...,e LN,M+3 ], e LN,m Enter E for this layer LN The mth column vector of the fully connected feedforward network, the mth column output e FFN,m It is expressed as:
[0133] e FFN,m =W2ReLU(W1e LN,m +b1)+b2 (Formula 11)
[0134] Among them, W1 is the weight of the first linear transformation, the dimension is 8L×2L, b1 is the bias of the first linear transformation, the dimension is 8L×1, W2 is the weight of the second linear transformation, the dimension is 2L×8L, b1 is the bias of the second linear transformation, the dimension is 2L×1;
[0135] The output of the encoder is where the vector z′ C The vector z′ is the sound speed estimation code. SNR For the signal-to-noise ratio estimation code, the vector z′ CLS For classification estimation coding, both are column vectors of length 2L, and the matrix is the normalized denoising estimation matrix, with a dimension of 2L×M, and the vector is the normalized denoising estimation matrix The mth column vector of ;
[0136] During the training phase, the output layer includes two parallel linear regressors, a linear classifier and an affine layer; the sound speed estimation is encoded Coding with SNR estimation Input into two linear regressors respectively to get the estimated sound speed Estimated signal-to-noise ratio Encoding categorical estimates Input to a linear classifier to predict the relationship between two sets of sequences; normalize the denoised estimation matrix Input to the affine layer to get the denoising estimation matrix Denoising Estimation Matrix It is expressed as:
[0137]
[0138] Among them, the denoising estimation matrix is The matrix size is 2L×M, and the vector For the matrix The mth column vector of ;
[0139] In the prediction stage, the output layer includes a linear regressor and an affine layer in parallel; the sound speed estimate is encoded Input to the linear regressor to get the estimated sound speed Normalize the denoised estimate matrix Input to the affine layer to get the denoising estimation matrix
[0140] S3. Train and improve the BERT model.
[0141] The data sets are divided into simulation data sets and measured data sets. The simulation data sets are based on different signal source numbers K, sound speed c, wave arrival direction angles θ1, θ2, …, θ K , signal-to-noise ratio snr, generate multiple sets of array signal receiving data X with white noise and its corresponding denoised receiving data X′, noisy matrix Z, denoising matrix Z′ to train the improved BERT model, where the array signal denoised receiving data X′ with dimension M×L is expressed as:
[0142] X′=AS (Formula 13) Where S=[S(1),S(t),…,S(L)] is the L snapshots of the target signal; denoising matrix Z′=[z′1,z′2,…,z′ m ,...,z′ M ] is the reconstruction matrix of the array signal denoising received data X′, with a dimension of 2L×M, and the vector z′ m is the mth column vector of the denoising matrix Z′, and the denoising matrix Z′ is expressed as:
[0143]
[0144] Among them, Re{X′} and Im{X′} are the real and imaginary parts of the array signal denoising received data X′ respectively; the simulation data set is used for pre-training the improved BERT model, and the measured data set is used for fine-tuning the improved BERT model; the training set and the test set are divided into 9:1;
[0145] The present invention uses four training objectives to supervise the training of the improved BERT model;
[0146] The first training goal is to predict the denoised signal sequence. The improved BERT model randomly masks 15% of the column vectors of the noisy matrix Z, and then predicts the column vectors of the denoised matrix Z′ corresponding to these masked column vectors, which is the usual cloze task. If the mth column vector z of the noisy matrix m is masked, then there is an 80% probability that the vector will be replaced by the masking vector z mask , mask vector z mask is the model parameter; there is a 10% probability of being replaced by a random vector z′ rand , which is a column vector of a random denoising matrix Z′ in the data set; there is a 10% probability of no replacement;
[0147] Mask the mth column vector z of the noisy matrix m After that, the mth column vector z′ of the corresponding denoising matrix will be predicted m ; The prediction value of improved BERT is The optimization goal is to minimize the mth column vector z′ of the denoising matrix m With the predicted value The squared error between:
[0148]
[0149] The second training objective is post-order array prediction; in natural language processing tasks, this training is used to capture the relationship between two sentences; in the present invention, in order to train a model that can capture the relationship between the received data of two sets of arrays, we train a binary post-order array prediction task; specifically, when selecting the previous noisy matrix Z for each sample Column vector and When the column vector is 50%, the probability of the noise matrix Z is The column vector is the actual posterior Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from the same L snapshots of data, with the sample label marked as 1; the posterior of the noisy matrix Z is 50% The column vector is a random segment of the array signal in the data set. The noisy matrix Z of the received data X rand After Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from data obtained from L different snapshots, and the sample labels are marked as -1; finally, the classification estimation encoding output by the BERT model encoder is improved The input is fed into a binary classifier for subsequent array prediction;
[0150] The label of sample Z is y∈{-1,1}, and the weight of the binary classifier is w CLS , the bias term is b CLS , the predicted value is sign(·) is the sign function, · is the vector dot product, and its optimization goal is to minimize the label y of sample Z and the predicted value of the binary classifier. The squared error between:
[0151]
[0152] Training goal three is to predict the speed of sound; sample Z is obtained in an environment with a speed of sound of c, and the speed of sound estimation code output by the improved BERT encoder is Input to a linear regressor with weight w C , the bias term is b C , the estimated speed of sound is The optimization goal is to minimize the sound speed c and the estimated sound speed The squared error between:
[0153]
[0154] Training goal 4: SNR prediction; the SNR of sample Z is snr, and the SNR estimate output by the improved BERT encoder is encoded Input to the second linear regressor, the weight of the second linear regressor is w SNR , the bias term is b SNR , the estimated signal-to-noise ratio is The optimization goal is to minimize the signal-to-noise ratio SNR and the estimated signal-to-noise ratio The squared error between:
[0155]
[0156] Supervised training is performed based on four training objectives until the model converges on the test set to obtain a trained improved BERT model.
[0157] S4. Reconstruct the received data X of the array signal to obtain the noisy matrix Z, and use the improved BERT model to preprocess the noisy matrix Z to obtain the denoising estimation matrix With estimated speed of sound Then estimate the denoising matrix Reconstruct the array signal to get the denoised received data
[0158] The noisy array signal receiving data X is reconstructed to obtain the noisy matrix Z, and then the noisy matrix Z is input into the improved BERT model to obtain the output denoising estimation matrix and estimate the speed of sound Note the denoising estimation matrix The matrix composed of the first L row vectors is The matrix composed of L row vectors is Array signal denoising received data The denoising matrix is estimated Reconstructed, where the array signal denoises the received data The real part of for:
[0159]
[0160] Imaginary part for:
[0161]
[0162] The above algorithm flow chart can be obtained from the attached Figure 4 express.
[0163] De-noise the array signal to receive data With estimated speed of sound As the input of existing DOA estimation methods (such as ESPRIT algorithm), the direction of arrival angle can be estimated.
[0164] Example 2
[0165] This embodiment discloses an underwater direction of arrival estimation preprocessing method, and the specific working steps are as follows:
[0166] T1. Establish the array signal model of one-dimensional uniform linear array;
[0167] The number of elements M of the uniform linear array is set to 10, the element spacing is d = 0.05m, the center frequency of the signal is f = 15kHz, the number of signals K = 1, the number of snapshots is L = 384, the noise N is white noise, the array signal receives data X, and its matrix size is M × L = 10 × 384.
[0168] T2. Build an improved BERT model;
[0169] The improved BERT model includes an input layer, an embedding layer, an encoder, and an output layer connected in sequence. The input layer is a noisy matrix Z with a matrix size of 2L×M=768×10; the embedding layer consists of an input encoder E in , fragment code E seg and position code E posThe encoder includes 6 encoder layers connected in sequence; each encoder layer includes four sub-layers connected in sequence, namely a multi-head attention layer, a layer normalization layer, a fully connected feedforward network layer and a layer normalization layer, and the input and output dimensions of each sub-layer are 2L×(M+3)=768×13; the front and back ends of the multi-head attention layer and the fully connected feedforward network layer use residual connections; the fully connected feedforward network includes two linear transformations connected in sequence, and there is a ReLU(·) activation between the two linear transformations; in the training stage, the output layer includes two parallel linear regressors, a linear classifier and an affine layer; in the prediction stage, the output layer includes a parallel linear regressor and an affine layer.
[0170] T3. Train and improve the BERT model;
[0171] T31. Construct a simulation data set. Generate multiple sets of array signal receiving data X with white noise and its corresponding denoised receiving data X', noisy matrix Z, and denoised matrix Z' at intervals of 1° for the wave direction of arrival angle of 0-60°, 1dB for the signal-to-noise ratio of -5-5dB, and 1m / s for the sound speed of 1500-1520m / s to form a simulation data set.
[0172] T32, the training set and the test set are divided into 9:1. Supervised training is performed based on four training objectives until the model converges on the test set. The BERT model is improved by pre-training with a simulated data set, and then fine-tuned with a measured data set to improve the BERT model.
[0173] T4. In this embodiment, the signal wave arrival direction angle θ=30°, under the environment of signal-to-noise ratio snr=0dB and sound speed c=1511m / s, array signal reception data X is generated, and the array signal reception data X is reconstructed using formula (6) to obtain a noisy matrix Z. The noisy matrix Z is preprocessed using the improved BERT model to obtain a denoising estimation matrix With estimated speed of sound Then estimate the denoising matrix Using formula (19) and formula (20) to reconstruct the array signal denoising received data Array signal receiving data X and array signal denoising receiving data Add Hamming windows respectively, and after Fourier transform, calculate the signal-to-noise ratio of the two in the frequency domain. The signal-to-noise ratio of the array signal receiving data X is 0.23dB, and the array signal denoising receiving data The signal-to-noise ratio of the data is 5.26dB. The signal-to-noise ratio of the data is improved after preprocessing, indicating that the preprocessing method has noise reduction performance.
[0174] Without preprocessing, the sound speed estimation is not performed, and the actual sound speed c = 1511m / s is not known. The default sound speed is usually 1500m / s. The array signal receiving data X and the default sound speed are used as the input of the TLS-ESPRIT algorithm, and the estimated wave direction angle is 29.63°; the array signal denoising receiving data With estimated speed of sound As the input of the TLS-ESPRIT algorithm, the estimated direction of arrival angle is 30.13°, which shows that the preprocessing method reduces the impact of the underwater variable sound speed environment on the existing DOA estimation algorithm and improves the estimation accuracy of the existing underwater DOA estimation algorithm.
[0175] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A method for preprocessing underwater direction of arrival estimation, characterized in that: The pretreatment method comprises the following steps: S1. Establishing an array signal model of a one-dimensional uniform linear array; S2. Construct an improved BERT model; the improved BERT model includes an input layer, an embedding layer, an encoder, and an output layer connected in sequence, wherein: The input layer is a noisy matrix Z = [z1, z2, ..., z m ,...,z M ], the dimension of the noise matrix Z is 2L×M, where the vector z m is the mth column vector, and the noisy matrix Z is reconstructed from the array signal receiving data X: Wherein, Re{X} and Im{X} are the real part and imaginary part of the array signal receiving data X respectively; Normalize the noisy matrix Z to obtain a normalized matrix E=[e1,e m ,…,e M ], where the vector e m is the mth column vector, and the normalized matrix E is expressed as: in, is the median of the elements in the noisy matrix Z, is half the length of the interval where the elements in the noisy matrix Z are located, v max is the maximum value of the elements in the noisy matrix Z, v min is the minimum value of the elements in the noisy matrix Z, and the matrix G represents a matrix whose elements are all 1; the normalized matrix E will be input into the embedding layer; The embedding layer is composed of the input encoding E in , fragment code E seg and position code E pos Add together; The input of the embedding layer is the normalized matrix E, and the sound velocity code z C , signal-to-noise ratio coding z SNR , classification code z CLS Construct an input code E of dimension 2L×(M+3) in , enter code E in It is expressed as: E in = [z C , z SNR , z CLS , E] (Formula 8) The speed of sound is encoded as z C , signal-to-noise ratio coding z SNR , classification code z CLS , are column vectors of length 2L, randomly initialized, and updated as parameters during training; Fragmentation Code E seg The dimension of the input layer is 2L×(M+3). When the input noise matrix Z=[z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from the same L snapshots, the slice encoding E seg The column vectors of E are all column vectors with all elements equal to 0. A =[0,...,0] T ; When the input noise matrix Z of the input layer is [z1,z2,...,z m ,...,z M ] Column vector and When the column vectors are derived from data obtained from different L snapshots, the slice encoding E seg Before Column vectors are all column vectors E whose elements are all 0 A =[0,...,0] T ,back Column vectors are all column vectors E whose elements are all 1 B =[1,...,1] T ; Positional Encoding pos The dimension is 2L×(M+3), and its i-th row and j-th column element is expressed as: Among them, floor() is a rounding function; S3. Train and improve the BERT model; S4. Reconstruct the received data X of the array signal to obtain the noisy matrix Z, and use the improved BERT model to preprocess the noisy matrix Z to obtain the denoising estimation matrix With estimated speed of sound Then estimate the denoising matrix Reconstruct the array signal to get the denoised received data 2. The underwater direction of arrival estimation preprocessing method according to claim 1, characterized in that: The process of step S1 is as follows: Assume that there are M array elements distributed on the uniform linear array, the array element spacing is d, and there are K far-field narrow-band acoustic wave signals as target signals. The i-th target signal is denoted as s i (t), i = 1, 2, ..., K, incident on the uniform linear array at the same time, and satisfy K < M, the center frequency of the signal is f, the speed of sound is c, and the wave arrival angle of the i-th target signal is θ i , The signal x1(t) received by the first array element is: Among them, s i (t) represents the i-th target signal, n1(t) represents the noise on the first array element channel; The time delay τ of the i-th target signal between the first array element and the m-th array element i for: The signal x received by the mth array element at the same time m (t) is: n m (t) represents the noise on the mth array element channel, then the signal X(t) received by M array elements at the same time is represented by the vector: X(t)=AS(t)+N(t) (Formula 4) The received signal is X(t)=[x1(t),x2(t),…,x M (t)] T , the superscript T indicates the transposition operation, and the flow matrix A is: S(t)=[s1(t),s2(t),…,s K (t)] T is the target signal, N(t)=[n1(t),n2(t),…,n M (t)] T for noise; After L snapshots, the array signal receiving data is X = [x1, x2, ..., x m …,x M ] T , the received noise is N = [n1,n2,…,n m ,…n M ] T , where the received data of the mth array element is x m =[x m (1),x m (t),…,x m (L)] T , the received noise is n m =[n m (1),n m (t),…,n m (L)] T .
3. The underwater direction of arrival estimation preprocessing method according to claim 2, characterized in that: The output E of the embedding layer out The dimension of the embedding layer is 2L×(M+3), and the output E out It is expressed as: E out = E in +E seg +E pos (Formula 10) The output of the embedding layer E out inputting into the encoder; The encoder includes N encoder layers connected in sequence; each encoder layer includes four sublayers connected in sequence, namely a multi-head attention layer, a layer normalization layer, a fully connected feedforward network layer and a layer normalization layer, and the input and output dimensions of each sublayer are 2L×(M+3); the front and rear ends of the multi-head attention layer and the fully connected feedforward network layer both use residual connections; The fully connected feedforward network consists of two linear transformations connected in sequence, with a ReLU() activation between the two linear transformations; the fully connected feedforward network is applied to the input E of this layer in the same way LN =[e LN,1 ,e LN,m ,...,e LN,M+3 ], e LN,m Enter E for this layer LN The mth column vector of the fully connected feedforward network, the mth column output e FFN,m It is expressed as: e FFN,m = W2ReLU(W1e LN,m +b1)+b2 (Official 11) Among them, W1 is the weight of the first linear transformation, the dimension is 8L×2L, b1 is the bias term of the first linear transformation, the dimension is 8L×1, W2 is the weight of the second linear transformation, the dimension is 2L×8L, b1 is the bias term of the second linear transformation, the dimension is 2L×1; The output of the encoder is where the vector z′ C The vector z′ is the sound speed estimation code. SNR For the signal-to-noise ratio estimation code, the vector z′ CLS For classification estimation coding, both are column vectors of length 2L, and the matrix is the normalized denoising estimation matrix, with a dimension of 2L×M, and the vector is the normalized denoising estimation matrix The m-th column vector of .
4. The underwater direction of arrival estimation preprocessing method according to claim 3, characterized in that: During the training phase, the output layer includes two parallel linear regressors, a linear classifier and an affine layer; the sound speed estimation is encoded Coding with SNR estimation Input into two linear regressors respectively to get the estimated sound speed Estimated signal-to-noise ratio Encoding categorical estimates Input to a linear classifier to predict the relationship between two sets of sequences; normalize the denoised estimation matrix Input to the affine layer to get the denoising estimation matrix Denoising Estimation Matrix It is expressed as: Among them, the denoising estimation matrix is The matrix size is 2L×M, and the vector For the matrix The m-th column vector of .
5. The underwater direction of arrival estimation preprocessing method according to claim 3, characterized in that: In the prediction stage, the output layer includes a linear regressor and an affine layer in parallel; the sound speed estimate is encoded Input to the linear regressor to get the estimated sound speed Normalize the denoised estimate matrix Input to the affine layer to get the denoising estimation matrix 6. The underwater direction of arrival estimation preprocessing method according to claim 3, characterized in that: The process of step S3 is as follows: The data set is divided into a simulation data set and a measured data set; the simulation data set is based on the number of signal sources K, the speed of sound c, the wave arrival direction angle θ1, θ2,…, θ K , signal-to-noise ratio snr, generate multiple sets of array signal receiving data X with white noise and its corresponding denoised receiving data X′, noisy matrix Z, denoising matrix Z′ to train the improved BERT model, where the array signal denoised receiving data X′ with dimension M×L is expressed as: X′=AS (Formula 13) Where S = [S(1), S(t), ..., S(L)] is the L snapshots of the target signal; the denoising matrix Z′ = [z′1, z′2, ..., z′ m ,...,z′ M ] is the reconstruction matrix of the array signal denoising received data X′, with a dimension of 2L×M, and the vector z′ m is the mth column vector of the denoising matrix Z′, and the denoising matrix Z′ is expressed as: Among them, Re{X′} and Im{X′} are the real and imaginary parts of the array signal denoising received data X′ respectively; the simulation data set is used for pre-training the improved BERT model, and the measured data set is used for fine-tuning the improved BERT model; the training set and the test set are divided into 9:1; Supervised training based on four training objectives to improve the BERT model; The first training goal is to predict the denoised signal sequence. The improved BERT model randomly masks 15% of the column vectors of the noisy matrix Z, and then predicts the column vectors of the denoised matrix Z′ corresponding to these masked column vectors, which is the usual cloze task. If the mth column vector z of the noisy matrix m is masked, then there is an 80% probability that the vector will be replaced by the masking vector z mask , mask vector z mask is the model parameter; there is a 10% probability of being replaced by a random vector z′ rand , which is a column vector of a random denoising matrix Z′ in the data set; there is a 10% probability of no replacement; Mask the mth column vector z of the noisy matrix m After that, the mth column vector z′ of the corresponding denoising matrix will be predicted m ; The prediction value of improved BERT is The optimization goal is to minimize the mth column vector z′ of the denoising matrix m With the predicted value The squared error between: The second training objective is post-order array prediction; in natural language processing tasks, this training is used to capture the relationship between two sentences; in order to train a model that can capture the relationship between the received data of the two arrays, a binary post-order array prediction task is trained; specifically, when selecting the previous noisy matrix Z for each sample Column vector and When the column vector is 50%, the probability of the noise matrix Z is The column vector is the actual posterior Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from the same L snapshots of data, with the sample label marked as 1; the posterior of the noisy matrix Z is 50% The column vector is a random segment of the array signal in the data set. The noisy matrix Z of the received data X rand After Column vector, that is, the front of the noise matrix Z of the sample Column vector and The column vectors are derived from data obtained from L different snapshots, and the sample labels are marked as -1; finally, the classification estimation encoding output by the BERT model encoder is improved The input is fed into a binary classifier for subsequent array prediction; The label of sample Z is y∈{-1,1}, and the weight of the binary classifier is w CLS , the bias term is b CLS , the predicted value is sign(·) is the sign function, · is the vector dot product, and its optimization goal is to minimize the label y of sample Z and the predicted value of the binary classifier. The squared error between: Training goal three is to predict the speed of sound; sample Z is obtained in an environment with a speed of sound of c, and the speed of sound estimation code output by the improved BERT encoder is Input to a linear regressor with weight w C , the bias term is b C , the estimated speed of sound is The optimization goal is to minimize the sound speed c and the estimated sound speed The squared error between: Training goal 4: SNR prediction; the SNR of sample Z is snr, and the SNR estimate output by the improved BERT encoder is encoded Input to the second linear regressor, the weight of the second linear regressor is w SNR , the bias term is b SNR , the estimated signal-to-noise ratio is The optimization goal is to minimize the signal-to-noise ratio SNR and the estimated signal-to-noise ratio The squared error between: Supervised training is performed based on four training objectives until the model converges on the test set to obtain a trained improved BERT model.
7. The underwater direction of arrival estimation preprocessing method according to claim 6, characterized in that: The process of step S4 is as follows: The noisy array signal receiving data X is reconstructed to obtain the noisy matrix Z, and then the noisy matrix Z is input into the improved BERT model to obtain the output denoising estimation matrix and estimate the speed of sound , denote the denoising estimation matrix The matrix composed of the first L row vectors is The matrix composed of L row vectors is , array signal denoising received data The denoising matrix is estimated Reconstructed, where the array signal denoises the received data The real part of for: Imaginary part for: Finally, the array signal denoising received data is obtained and estimate the speed of sound
Citation Information
Patent Citations
Underwater target direction-of-arrival estimation method based on convolutional neural network
CN114397621A
Deep learning beam domain channel estimation method based on approximate message passing algorithm
WO2020253690A1
Cited By
Direction of arrival estimation method fusing improved genetic algorithm and MUSIC
CN121805939A