A brain electrical response deep learning classification and identification method based on underwater acoustic signal stimulation

By combining the advanced cognition and rapid perception capabilities of the human brain, and utilizing EEG signals and deep learning methods, the DenseNet-121 improved network is used to classify and identify underwater acoustic target signals. This solves the problem of difficulty in identification under low signal-to-noise ratio conditions in traditional methods, and achieves higher accuracy and real-time processing capabilities.

CN115470821BActive Publication Date: 2026-02-17NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211132459.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-17
Publication Date
2026-02-17
Estimated Expiration
2042-09-17

AI Technical Summary

Technical Problem

Existing underwater acoustic target recognition methods perform poorly under low signal-to-noise ratio conditions, making it difficult to distinguish between underwater background noise and underwater acoustic target signals. Traditional machine learning algorithms cannot meet the needs of complex scenarios.

Method used

Combining the advanced cognition and rapid perception capabilities of the human brain, and using electroencephalogram (EEG) signals as input, this study employs deep learning methods to classify and identify underwater acoustic target signals, and utilizes an improved DenseNet-121 network for feature extraction and classification.

Benefits of technology

It improves the accuracy and real-time processing capability of underwater acoustic target identification, reduces processing time and cost, and is suitable for complex underwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470821B_ABST
    Figure CN115470821B_ABST
Patent Text Reader

Abstract

The application provides a brain electrical response deep learning classification and identification method based on underwater acoustic signal stimulation, which is based on brain electrical signals and combines deep learning to classify and identify underwater acoustic target signals. The method fully utilizes the differences in the subjective sensitivity of people to underwater noise and underwater acoustic signals, the rapid auditory perception of the human brain to different types of underwater acoustic signals and the differences in the conscious information extraction, takes the brain electrical signals as the main input signals, and classifies and identifies the underwater acoustic signals after corresponding processing of the brain electrical signals. The application plays the "filtering characteristics" of the human brain which is more advanced and more targeted than computers, and the characteristics of more accurate classification and identification, and can improve the poor target recognition effect in a low signal-to-noise ratio environment in the traditional underwater acoustic target classification method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a classification and recognition method for underwater acoustic target signals based on electroencephalogram (EEG) signals and deep learning methods, which can be applied to complex underwater environments. Background Technology

[0002] Underwater acoustic target identification is an important component of modern maritime surveillance systems and a key technology in underwater detection and acoustic countermeasures. Accurate identification of underwater targets is of great significance for the utilization of marine resources and space, national defense security, and economic development, and has now become one of the hot research topics in the field of underwater acoustics.

[0003] Early underwater acoustic target recognition relied primarily on experienced sonar operators, but the accuracy of their judgments was significantly influenced by their experience, physical condition, and psychological factors. Traditional machine learning algorithms, such as support vector machines, are still widely used, but their accuracy remains limited by the small size of underwater acoustic signal samples and high background noise levels. In recent years, with the rapid development of machine learning methods, breakthroughs have been achieved in underwater acoustic target recognition using automated machine identification techniques. Deep learning algorithms, primarily based on various convolutional network structures, can identify underwater acoustic targets relatively accurately, and the transfer learning approach between the source and target domains has, to some extent, overcome the problem of small underwater acoustic signal samples. However, although the efficiency and speed of machine-based target recognition methods have significantly improved compared to the past, relying solely on computer algorithms for target recognition still struggles to meet the demands of complex scenarios, especially the complexity of the underwater environment and the adversarial nature of targets, which greatly affect the accuracy of underwater acoustic target recognition.

[0004] The brain is the most complex organ in the human body. Compared to computers, the human brain possesses a unique ability to process complex situations, capable of generating neural responses to key and sensitive information within milliseconds. Its advantages can be divided into two aspects: first, advanced cognitive abilities, meaning the human brain has exceptional cognitive capabilities for unstructured complex information, including emotional processing, semantic understanding, and temporal correlation; second, rapid perceptual abilities, the human brain can always automatically and quickly extract statistical results or corresponding patterns from perceptual information, and this process is usually unconscious and does not require task relevance. When the human brain perceives external stimuli, engages in thought processes, and generates subjective consciousness, a series of biochemical reactions and electrical activities accompany the operation of its nervous system, producing measurable electrochemical signals, namely electroencephalograms (EEGs). Currently, classification and recognition tasks that primarily use EEG signals are mostly focused on visual responses, with almost no applications in underwater target detection and recognition. Summary of the Invention

[0005] To address the shortcomings of traditional underwater acoustic signal detection methods, which primarily rely on signal feature extraction and machine learning, such as poor performance under low signal-to-noise ratio conditions and difficulty in distinguishing underwater background noise from underwater acoustic target signals, this invention proposes a method for classifying and recognizing underwater acoustic target signals based on electroencephalogram (EEG) signals and deep learning. This method leverages the varying degrees of subjective sensitivity of humans to underwater noise and underwater acoustic signals, as well as the differences in the human brain's rapid auditory perception and conscious information extraction of different types of underwater acoustic signals. Using EEG signals as the primary input signal, the method processes the EEG signals before classifying and recognizing the underwater acoustic signals.

[0006] The technical solution of this invention is as follows:

[0007] The aforementioned deep learning classification and recognition method based on electroencephalogram (EEG) responses to underwater acoustic signal stimulation includes the following steps:

[0008] Step 1: Collect audio signals of known underwater acoustic signals and label each type of underwater acoustic signal;

[0009] Step 2: Using the processed underwater acoustic signal audio from Step 1, the subject wears an EEG signal acquisition device and in-ear headphones to conduct an EEG experiment and collect the subject's initial EEG signal after hearing the underwater acoustic signal audio.

[0010] Step 3: Perform data preprocessing on the initial EEG signals collected in Step 2 to obtain pure EEG signals and corresponding underwater acoustic signal labels, forming a training dataset;

[0011] Step 4: Use the training dataset obtained in Step 3 to train the neural network and obtain the optimal network parameter model;

[0012] Step 5: For the underwater acoustic signal to be identified and classified, the same subject puts on the EEG signal acquisition device and in-ear headphones, and performs the EEG experiment again to obtain the EEG response data of the underwater acoustic target to be identified and classified.

[0013] Step 6: Following the data preprocessing process in Step 3, preprocess the EEG response data obtained in Step 5 to obtain clean EEG signals;

[0014] Step 7: Input the clean EEG signal obtained in Step 6 into the optimal network parameter model trained in Step 4 to obtain the final classification and recognition result.

[0015] Furthermore, in step 1, the sound pressure level of the underwater acoustic signal is adjusted so that the subject is not affected by excessive volume when hearing it.

[0016] Furthermore, during the EEG experiment in step 2, the audio of each type of underwater acoustic signal was played the same number of times.

[0017] Furthermore, the initial EEG signal X(n) obtained in step 2 is:

[0018] X(n)=S(n)+N EOG (n)+N EMG (n)+N x (n)

[0019] The dimension of X(n) is the underwater acoustic signal category × the number of times each signal is played × the number of electrodes used × the dimension of the signal x(n) collected by each electrode, where x(n) is a column vector and is a one-dimensional nonlinear time series signal; S(n) represents the EEG signal containing useful information, N EOG (n) represents electrooculography noise, N EMG (n) represents electromyographic noise, N x (n) represents other noise, including electrocardiogram noise and power frequency noise.

[0020] Furthermore, the process of data preprocessing for the initial EEG signals in step 3 is as follows:

[0021] First, the EEG signal is high-pass filtered at 0.5–1 Hz to remove low-frequency noise N. x (n) and some low-frequency electrooculography noise N EOG (n), thus obtaining signal X1(n);

[0022] The obtained signal X1(n) is then subjected to fourth-order wavelet packet decomposition. The high-frequency components with wavelet packet tree indices 27–30, corresponding to frequencies of 49–64 Hz, are set to zero, and the signal is reconstructed to obtain N after filtering out high-frequency noise and electromyographic noise. EMG The purer EEG signal X2(n) of (n);

[0023] Finally, singular spectrum analysis was performed on the EEG signal X2(n) to obtain the pure EEG signal X3(n).

[0024] Furthermore, the neural network used in step 4 is an improved DenseNet-121 network with dilated convolutions.

[0025] Furthermore, the improved DenseNet-121 network with dilated convolutions is a one-dimensional DenseNet-121 network. The DenseNet-121 network contains four Dense Blocks, each containing 6, 12, 24, and 16 Dense Layers composed of 1×1 and 1×3 convolutional layers, respectively, to extract features by calculating convolutions. A Transition Layer is connected between every two Dense Blocks to halve the dimension of the output features. Finally, a global average pooling layer and a fully connected layer are connected, and the final result is obtained by classification using the Softmax function.

[0026] Furthermore, in the DenseNet-121 network, the 1×3 convolutional layers in the Dense Layer are dilated convolutional layers with dilation rates varying with the number of layers: the 1×3 convolutional layers in each Dense Layer are successively dilated with dilation rates of 2... i-1 The dilation rate of the dilated convolutional layers i = 1, 2, ..., 5, up to the sixth Dense Layer, is reduced from 2... 0 Run the loop.

[0027] Furthermore, the underwater acoustic targets to be identified and classified in step 5 belong to one or more of the underwater acoustic signals already classified in step 1.

[0028] Beneficial effects

[0029] This invention proposes a method for classifying and recognizing underwater acoustic targets based on electroencephalogram (EEG) signals. It leverages the human brain's superior and more targeted filtering characteristics compared to computers, along with its more accurate classification capabilities, to improve upon the poor target recognition performance in low signal-to-noise ratio environments of traditional underwater acoustic target recognition methods. The method requires either a complex but higher-sampling-rate 32- or even 64-lead electrode cap, or a portable 16-lead EEG for convenient outdoor experiments. The method is also very simple to operate, suitable for healthy adults without hearing impairments, and theoretically, those with relevant work experience will achieve even greater accuracy. Furthermore, since the sampling rate of EEG signals is much lower than that of underwater acoustic signals, the processing time and cost for the same duration of EEG signals are significantly reduced compared to underwater acoustic signals. Comparison with traditional underwater target recognition methods verifies that this invention improves classification accuracy to a certain extent. Moreover, due to the high temporal resolution and small data size of EEG signals, it can demonstrate real-time processing capabilities in future applications, showing broad application prospects in underwater acoustic target recognition.

[0030] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0031] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0032] Figure 1 The EMOTIV 16-channel wireless portable EEG device used in this example method;

[0033] Figure 2 : A schematic diagram of the equipment connection and experiment for the EEG signal acquisition step in this example method;

[0034] Figure 3 : A schematic diagram of the improved DenseNet-121 network structure with dilated convolutions used in this method. Detailed Implementation

[0035] This invention achieves underwater acoustic target signal classification and recognition through three stages. The first stage is the model training dataset acquisition stage, where pre-classified underwater acoustic signals are used to design relevant EEG experiments for subjects, and the corresponding EEG signals and underwater acoustic signal classification labels are collected as the dataset for training the optimal network parameter model. The second stage is the network model training stage, which mainly uses the model training dataset acquired in the first stage to train the parameters of a deep neural network, obtaining a network model with optimal parameters. The third stage is the underwater target classification and recognition stage, where the underwater acoustic target signal to be classified is used as input and fed into the optimal network parameter model for classification and recognition, obtaining the final classification and recognition result and classification and recognition accuracy. The implementation steps of each stage are as follows:

[0036] Phase 1: Model training dataset collection phase

[0037] Step 1: Collect audio signals of known underwater acoustic signals, label each type of underwater acoustic signal, and adjust parameters such as sound pressure level of each signal to ensure that the subject is not affected by excessive volume, which could affect the quality of the collected EEG data.

[0038] Step 2: Conduct the EEG experiment and collect EEG signals. Prepare the processed underwater acoustic signal audio from Step 1. Have the subject wear the EEG signal acquisition device and in-ear headphones. After confirming that all electrodes on the electrode cap are connected, begin the experiment. During the experiment, ensure that each type of underwater acoustic signal audio is played the same number of times, and ensure that the subject is relaxed and gets necessary rest. Finally, obtain the initial EEG signal array X(n), where...

[0039] X(n)=S(n)+N EOG (n)+NEMG (n)+N x (n) (1)

[0040] The array X(n) has dimensions equal to the number of underwater acoustic signal categories × the number of times each signal is played × the number of electrodes used × the dimension of the signal x(n) collected by each electrode. x(n) is a column vector, representing a one-dimensional nonlinear time-series signal. S(n) represents the EEG signal containing useful information, and N... EOG (n) represents electrooculography noise, N EMG (n) represents electromyographic noise, N x (n) represents other noise with smaller amplitude, such as electrocardiogram noise, power frequency noise, etc.

[0041] Step 3: Preprocess the EEG signals X(n) acquired in Step 2 to obtain clean EEG signals and corresponding underwater acoustic signal labels, forming a training dataset. In practice, the EEG signals are first subjected to a high-pass filter of 0.5–1 Hz to remove low-frequency noise N. x (n) and some low-frequency electrooculography noise N EOG (n), to obtain signal X1(n):

[0042] X1(n)=HPF(X(n)) (2)

[0043] HPF stands for high-pass filter.

[0044] Then, using Matlab's wavelet packet transform function, a fourth-order wavelet packet decomposition is performed on the obtained signal. The high-frequency components with wavelet packet tree indices 27–30, corresponding to frequencies of 49–64 Hz, are set to zero, and the reconstructed signal is obtained after filtering out high-frequency noise and electromyographic noise, N. EMG The purer EEG signal X2(n) of (n):

[0045] X2(n)=WPTRec (4) (WT(X1(n))) (3)

[0046] Among them, WPTRec (4) This represents the fourth-order wavelet packet transform.

[0047] Finally, singular spectrum analysis is performed on the results obtained in the previous step to further remove noise and smooth the signal. Singular spectrum analysis is a principal component analysis method based on nonlinear time series. The idea is to construct a trajectory matrix using phase space reconstruction, group the eigenvectors by performing singular value decomposition on the trajectory matrix, and finally group the time series using diagonal averaging. The steps of singular spectrum analysis are as follows:

[0048] Step 1: Set the window length to w, and construct the trajectory matrix X using a one-dimensional time series x(n), n = 1, 2, 3, ..., N, where the length of each row of the matrix is ​​L = N - w + 1. The trajectory matrix is ​​as follows:

[0049]

[0050] Step 2: Calculate the covariance matrix XX T Then, perform singular value decomposition on it to obtain w eigenvalues ​​λ. j Satisfying λ1≥λ2≥…≥λ w ≥0 and the corresponding eigenvector U j At this time, X satisfies:

[0051]

[0052]

[0053] in,

[0054] Step 3: According to equation (5), divide the subscript set {1,2,3,...,w} into m disjoint subsets {I1,...,I...} m}, then X = X I1 +X I2 +...+X Ik ...+X Im Let I = {i1,...,i} p}, such that the composite matrix X corresponding to I Ik =X i1 +X i2 +...+X ip Each composite matrix X Ik Each can be viewed as a distinct component, representing a specific motion feature of the original sequence. The contribution rates of the different components are:

[0055]

[0056] Generally speaking, components with higher contribution rates contain more useful information, while components with lower contribution rates contain more noise.

[0057] Step 4: In the composite matrix (X) Ik ) ij For a matrix X, calculate the average of all diagonal elements along the diagonal i+j=n+1 of the matrix k=1,2,3,...,m. Ij Transformed into a new one-dimensional time series y j (n), n=1,2,3,...,N,y j The construction form of (n) is as follows:

[0058]

[0059] The original time series signal can be reconstructed into y(n) = (y1, y2, ..., y m ).

[0060] In this method, the column vector x2(n) of the EEG signal array X2(n), i.e., the relatively pure EEG signal of each trial, is decomposed into several different signal components using SSA. Components with higher contribution values ​​are retained, while those with lower contribution values ​​are removed. Then, time series reconstruction is performed to further remove noise and smooth the signal. Finally, a pure EEG signal X3(n) is obtained. X3(n) and its corresponding four-class labels are used as the dataset for training the optimal network parameter model.

[0061] X3(n)=SSA(X2(n)) (9)

[0062] The second stage, the network model training stage:

[0063] Step 4: Train the neural network using pure EEG signal X3(n) to obtain the optimal network parameter model. This method preferably uses an improved DenseNet-121 network with dilated convolutions. DenseNet is a densely connected neural network where the input of each layer comes from the output of all previous layers. The feature information calculated by each layer of the network is utilized more effectively, which significantly improves the accuracy of classification and recognition.

[0064] Since EEG signals and their features are one-dimensional signals, a one-dimensional DenseNet-121 network is used. The DenseNet-121 network consists of four Dense Blocks, each containing 6, 12, 24, and 16 Dense Layers, primarily composed of 1×1 and 1×3 convolutional layers, respectively. Feature extraction is performed through convolution, and the increase in feature dimension after each Dense Layer is determined by the growth rate, which is the main structure of the convolutional neural network. A Transition Layer connects every two Dense Blocks, halving the output feature dimension. Finally, a global average pooling layer and a fully connected layer are applied, and the final result is obtained through classification using the Softmax function.

[0065] To reduce the number of training parameters and improve the network's learning rate, the 1×3 ordinary convolutional layers of the Dense Layer are transformed into dilated convolutional layers with a dilation rate varying with the number of layers. This increases the receptive field, allowing each convolutional layer to output a wider range of information, thereby reducing the number of training parameters. Within each Dense Block, the 1×3 ordinary convolutional layers of each Dense Layer are successively transformed into dilated convolutional layers with a dilation rate of 2... i-1 The dilation rate of the dilated convolutional layers (i = 1, 2, ..., 5) continues until the sixth Dense Layer, where the dilation rate returns to 2. 0 The loop continues: the first 1×3 ordinary convolutional layer of the Dense Layer is changed to a dilation rate of 2. 0 The first layer is a dilated convolutional layer with the same receptive field as the original layer, while the second 1×3 ordinary convolutional layer with a dilation rate of 2 becomes a dilated layer with a dilation rate of 2. 1 The dilated convolutional layers expand the receptive field from 1×3 to 1×13, increasing with each subsequent layer until the fifth DenseLayer, where the receptive field expands to 1×48. The sixth Dense Layer is identical to the first, and so on for subsequent DenseLayers. A schematic diagram of the improved DenseNet-121 network with dilated convolutions is attached. Figure 3 As shown.

[0066] The network model training dataset obtained in the first stage is trained multiple times by adjusting the network structure hyperparameters according to the ratio of training set: validation set: test set = 20%: 30%: 50% to obtain the optimal network parameter model. The accuracy of network recognition and the running time required by the network are then calculated.

[0067] Phase 3: Underwater acoustic target recognition phase

[0068] Step 5: The same subject performs another EEG experiment to obtain EEG response data of the underwater acoustic target to be identified and classified. Here, the underwater acoustic target to be identified and classified belongs to one or more of the pre-classified underwater acoustic signals in the first-stage network model training dataset.

[0069] To assess the recognition accuracy, the audio signals of the underwater acoustic targets to be classified were processed according to step 1, and then the experiment was conducted according to step 2. Data was collected, and classification labels for the underwater acoustic targets to be classified were prepared, resulting in the EEG response data of the underwater acoustic target signals to be classified. It is important to note that conducting consecutive experiments on the same subject can lead to fatigue, increasing the probability of distraction, meditation, or other brain activities, thus affecting the quality of the experimental data and causing significant deviations in subsequent classification results. Therefore,

[0070] The experiment in step 5 should not be performed at the same time as the experiment in step 2. A longer rest period should be allowed between the two experiments to ensure that the subjects can concentrate on the experiment.

[0071] Step 6: Preprocess the EEG response data obtained in Step 5 according to the data preprocessing method in Step 3 to obtain a clean EEG signal. In order to judge the recognition accuracy in the future, the label of the underwater acoustic target signal to be classified is also known here.

[0072] Step 7: Input the clean EEG signal obtained in Step 6 into the optimal network parameter model of the DenseNet-121 improved network with dilated convolution obtained in Step 4 to obtain the final classification and recognition result, and compare it with the label of the underwater acoustic target signal to be classified to calculate the classification and recognition accuracy.

[0073] The embodiments of the present invention are described in detail below. These embodiments are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0074] Step 1: The underwater acoustic audio used in this example is the four types of lake test datasets collected by our research group in September 2018 on Danjiangkou Lake. These datasets represent the radiated noise of four different types of ships. The data was collected around the clock using an 8-element linear array with a sampling frequency of 48kHz.

[0075] The underwater audio played in this experiment was recorded from one channel. Five audio samples under the same conditions were taken for each type of lake test data. Each signal was played twice, with the playback order randomized. Therefore, the audio for each type of lake test data was played a total of 10 times, and each underwater audio sample was truncated to 10 seconds in length. A total of 40 trials were conducted across the four types of lake test audio. No instances of excessive volume or overstimulation of the subjects were observed during the listening tests; therefore, only the volume of each signal needed to be adjusted consistently without further adjustments.

[0076] Step 2: The experimental equipment used to collect the network model training dataset in this example is a wireless portable EEG device from EMOTIV, as shown in the attached image. Figure 1 As shown, the device has two sampling rates: 128Hz and 256Hz. This experiment uses the 128Hz sampling rate to collect data. The device has a total of 16 electrodes, and 14 electrodes were used in the experiment, namely AF3, AF4, F7, F8, F3, F4, FC5, FC6, T7, T8, P7, P8, O1, and O2.

[0077] This experiment was conducted at 3 PM to ensure participants were in good mental condition after their lunch break. A quiet laboratory setting was chosen, and room brightness was reduced to minimize light stimulation for the participants. During the experiment, cotton plugs were first inserted into each electrode, then moistened with saline solution. Following the international 10-20 standard, the EEG device was fitted, and the electrodes were correctly aligned. The EEG device was then connected to the signal acquisition equipment, and in-ear headphones were worn until the connection was stable before the experiment began. A schematic diagram of the entire EEG signal acquisition process is attached. Figure 2 As shown in the figure. In the experiment, the participants were asked to pay attention to the differences between the underwater acoustic signals without having to perform any other operations. The participants were required to maintain a comfortable body posture, minimize large limb movements, concentrate when they heard the underwater acoustic signals, and minimize blinking to avoid affecting the quality of the collected signals.

[0078] The final training dataset for the network model collected 40 EEG signals across four categories, each 20 seconds long. The time during which the subjects heard the underwater acoustic signals was the segment containing valid information, from the 5th second to the 15th second of each signal. To avoid the impact of sudden audio playback and termination on data quality, only 8 seconds (7-14 seconds) of each signal were collected, totaling 1024 sampling points. Furthermore, due to poor contact with the subject's scalp at electrodes T7 and T8, the collected data was abnormal and therefore discarded. The final dataset size was: 4 categories of underwater acoustic targets × 10 signals per category × 12 electrode channels × 1024 points per channel.

[0079] Step 3: Perform data preprocessing on each EEG signal collected in Step 2 to obtain a clean EEG signal after noise filtering.

[0080] The collected network model training data contained a lot of low-frequency noise from the device itself and eye movements, so a 0.5Hz high-pass filter was used to filter out this noise.

[0081] First, wavelet packet decomposition and reconstruction of the EEG signal were performed using a Matlab program. Among many wavelet basis functions, the db4 wavelet has certain similarities to the EEG signal waveform, so the db4 wavelet was used to perform fourth-order wavelet packet decomposition on the signal. The coefficient vectors with indices 27-30 (corresponding frequencies of 49-64Hz) on the obtained wavelet tree were reset to 0 before wavelet packet reconstruction was performed to obtain the EEG signal after filtering out high-frequency noise. This method retains the effective information to the maximum extent while reducing the data components that need to be processed, thereby improving the processing efficiency of this method.

[0082] The wavelet packet-reconstructed EEG signal is then decomposed into 15 components using SSA. The top 6 components, ranked by contribution rate from highest to lowest, are selected as the components for the reconstructed signal and participate in the reconstruction of the one-dimensional time-domain signal, while the remaining 9 components are discarded. Finally, a smooth and clean EEG signal is obtained.

[0083] Step 4: This example uses the appendix Figure 3 The DenseNet-121 network model with dilated convolutions shown is trained using the network model training dataset processed in step 3 as input. The network is set to receive 1×224 data points at a time, with a hop size of 64 points between inputs, a batch size of 4, a learning rate of 0.001, a dropout rate of 0.3, a growth rate of 12, and a training period of 100 epochs. In this example, when training the network parameter model, the training dataset is divided into a 20% training set, a 30% validation set, and a 50% test set. On a single GeForce RTX 2080 GPU, the optimal network parameter model achieves an accuracy of over 99% on the training set and an average accuracy of around 96% on the test set. The average training time for the entire process is approximately 1260 seconds.

[0084] Step 5: In this example, the subjects for the EEG response data acquisition experiment to be classified and identified are the same as those for the network model training data acquisition experiment in Step 2. The experimental equipment is the same, and the experimental time is also 3 PM, but not on the same day. The experimental location is a quiet, dimly lit workshop. Before the experiment, 40 additional underwater acoustic audio recordings of four types of ship radiated noise under the same working conditions and of the same category were processed according to Step 1. During the experiment, underwater acoustic audio recordings were played randomly, with each audio recording played twice. The audio recordings for each category label were played a total of 20 times. The EEG experiment was conducted according to Step 2, and EEG response data were collected, resulting in 80 EEG response data for the underwater acoustic targets to be classified across four categories. To maintain consistency with the training set data, the data from channels T7 and T8 were discarded, making the array size 4 categories of underwater acoustic targets × 20 signals per category of underwater acoustic targets × 12 electrode channels × 1024 points extracted from each channel. The data were labeled with the corresponding underwater acoustic target classification labels, and finally, the EEG response data of the underwater acoustic target signals to be classified and identified were obtained.

[0085] Step 6: Preprocess the data obtained in Step 5 according to the method in Step 3. First, perform high-pass filtering, then wavelet packet decomposition and reconstruction, and SSA decomposition and reconstruction to remove noise and non-useful information components, and obtain smooth and pure EEG signals and their labels, which will serve as the EEG response dataset to be classified in this method.

[0086] Step 7: Prepare the optimal network parameter model trained in Step 4, input the EEG signal to be classified obtained in Step 6 into the network, obtain the final classification and recognition result, and calculate the classification and recognition accuracy.

[0087] In this example, the EEG signals to be classified are fed into a DenseNet-121 network model with dilated convolutions, achieving a classification accuracy of 71%, demonstrating that this method has a certain level of classification accuracy. Furthermore, due to the low sampling rate and small data volume of EEG signals, the time cost required for network training is relatively low. During the network model training phase of this method, under the computation of a single GeForce RTX 2080 GPU, the average network training time required according to the hyperparameter settings in step 4 is only about 1290 seconds, far shorter than the network training time using raw underwater acoustic data as input, proving that this method has a relatively fast network model training speed for underwater acoustic target recognition. Since network model training is the most time-consuming stage in the entire data processing process, it can be said that this method has a relatively fast speed for underwater acoustic target recognition, thus demonstrating a certain degree of rapid processing capability and real-time processing potential.

[0088] Furthermore, in this example, if the dataset collected in step 5 and processed in step 6 is used as the training dataset for the network model, with other parameters remaining unchanged, the model training results are still quite good. The accuracy of the optimal network parameter model training set can still reach over 99%. Due to the increased training data volume, the accuracy of the test set decreases slightly on average, but it can still reach around 86%. The average training time for the entire process is approximately 2660 seconds. However, if the dataset collected in step 2 and processed in step 3 is used as the EEG response dataset to be classified, the accuracy obtained using the obtained optimal network parameter model can reach over 90% on average. This demonstrates the advantages of DenseNet as a deep neural network algorithm in handling data classification and recognition problems; the more training data the model has, the better the network performance and the more accurate the final classification result. On the other hand, it illustrates that the acquisition of EEG signals requires attention to many other influencing factors. Changes in the experimental location and the number of trials conducted by the subjects can cause shifts in the subjects' attention, which can affect the final data quality.

[0089] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A deep learning classification and recognition method for EEG response based on underwater acoustic signal stimulation, characterized in that: Includes the following steps: Step 1: Collect audio signals of known underwater acoustic signals and label each type of underwater acoustic signal; Step 2: Using the processed underwater acoustic signal audio from Step 1, the subject wears an EEG signal acquisition device and in-ear headphones to conduct an EEG experiment, collecting the subject's initial EEG signal X(n) after hearing the underwater acoustic signal audio: The dimension of X(n) is the underwater acoustic signal category × the number of times each signal is played × the number of electrodes used × the dimension of the signal x(n) collected by each electrode, where x(n) is a column vector and is a one-dimensional nonlinear time series signal; S(n) represents the EEG signal containing useful information, N EOG (n) represents electrooculography noise, N EMG (n) represents electromyographic noise, N x (n) represents other noise, including electrocardiogram noise and power frequency noise; Step 3: Perform data preprocessing on the initial EEG signals acquired in Step 2 to obtain clean EEG signals and corresponding underwater acoustic signal labels, forming a training dataset; the data preprocessing process is as follows: First, the EEG signal is high-pass filtered at 0.5~1Hz to remove low-frequency noise N. x (n) and some low-frequency electrooculography noise N EOG (n), thus obtaining signal X1(n); The obtained signal X1(n) is then subjected to fourth-order wavelet packet decomposition. The high-frequency components with wavelet packet tree indices 27-30 and corresponding frequencies of 49-64Hz are set to zero, and the signal is reconstructed to obtain N after filtering out high-frequency noise and electromyographic noise. EMG The purer EEG signal X2(n) of (n); Finally, singular spectrum analysis was performed on the EEG signal X2(n) to obtain the pure EEG signal X3(n); Step 4: Train the neural network using the training dataset obtained in Step 3 to obtain the optimal network parameter model; the neural network is an improved DenseNet-121 network with dilated convolutions; The improved DenseNet-121 network with dilated convolutions is a one-dimensional DenseNet-121 network. The DenseNet-121 network contains four Dense Blocks, each containing 6, 12, 24, and 16 Dense Layers composed of 1×1 and 1×3 convolutional layers, respectively. Feature extraction is performed by calculating convolutions. A Transition Layer is connected between every two Dense Blocks to halve the dimension of the output features. Finally, a global average pooling layer and a fully connected layer are connected, and the final result is obtained by classification using the Softmax function. In the DenseNet-121 network, the 1×3 convolutional layers in the Dense Layer are dilated convolutional layers with dilation rates varying with the number of layers: the 1×3 convolutional layers in each Dense Layer are successively dilated with dilation rates of 2... i-1 The dilation rate of the dilated convolutional layers i=1,2,...,5, up to the sixth Dense Layer, is reduced from 2... 0 Perform a loop; Step 5: For the underwater acoustic signal to be identified and classified, the same subject puts on the EEG signal acquisition device and in-ear headphones, and performs the EEG experiment again to obtain the EEG response data of the underwater acoustic target to be identified and classified. Step 6: Following the data preprocessing process in Step 3, preprocess the EEG response data obtained in Step 5 to obtain clean EEG signals; Step 7: Input the clean EEG signal obtained in Step 6 into the optimal network parameter model trained in Step 4 to obtain the final classification and recognition result.

2. The deep learning classification and recognition method for EEG response based on underwater acoustic signal stimulation according to claim 1, characterized in that: In step 1, the sound pressure level of the underwater acoustic signal is adjusted so that the subject is not affected by the excessive volume when hearing it.

3. The deep learning classification and recognition method for EEG response based on underwater acoustic signal stimulation according to claim 1, characterized in that: During the EEG experiment in step 2, the audio of each type of underwater acoustic signal was played the same number of times.

4. The deep learning classification and recognition method for EEG response based on underwater acoustic signal stimulation according to claim 1, characterized in that: The underwater acoustic targets to be identified and classified in step 5 belong to one or more of the underwater acoustic signals already classified in step 1.

Citation Information

Patent Citations

  • Music tone imagination distinguishing method based on electroencephalogram signals

    CN112799505A

  • Target detection method and system based on auditory brain-computer interface

    CN114781461A