A method for vocalization recognition of marine mammals based on a one-dimensional convolutional neural network

By designing a marine mammal vocal recognition method based on one-dimensional convolutional neural network, the acoustic signal is directly recognized, which solves the problems of low recognition rate and poor generalization ability in the prior art, achieves higher recognition efficiency and accuracy, and avoids the influence of image scale transformation.

CN115762534BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211378501.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-06-27
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

The prior art relies on converting acoustic signals into images in vocal recognition in marine mammals, resulting in low recognition rate and poor generalization ability, and the impact of image scale transformation on recognition efficiency cannot be effectively avoided.

Method used

A marine mammal vocal recognition method based on one-dimensional convolutional neural network was designed to directly identify the acoustic signals, avoiding the step of converting the acoustic signals into images. The one-dimensional convolutional neural network model 1D-AlexNet is used, which includes five convolutional layers, three pooling layers and three fully connected layers. A spatial pyramid pooling layer is added before the fully connected layer to unify the dimensions of the input data.

Benefits of technology

It achieves higher recognition efficiency and better recognition performance, with the test set recognition accuracy reaching about 96%, and avoids the impact of image scale transformation on recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762534B_ABST
    Figure CN115762534B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying vocalizations of marine mammals based on a one-dimensional convolutional neural network. First, preprocessing is performed on the vocalization samples of marine mammals. Secondly, a one-dimensional convolutional neural network that can be used for identifying vocalizations of marine mammals is designed to directly identify acoustic signals, avoiding the interference caused by the step of converting acoustic signals into images to the recognition results. Compared with the original two-dimensional convolutional neural network, the one-dimensional convolutional neural network designed in the present invention has higher recognition efficiency and better recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a method for identifying vocalizations of marine mammals. Background Art

[0002] Traditional methods for identifying and classifying vocalizations of marine mammals rely on manual feature extraction and heavily depend on prior knowledge, resulting in low recognition rates and poor generalization ability of the methods. Due to the great success of convolutional neural networks in optical images, they are often applied to underwater signal processing. Dorian proposed a method of converting the vocalization signals of humpback whales into spectrograms through fast Fourier transform and identifying them through a convolutional neural network, achieving excellent results under different noise backgrounds; Ferguson also adopted the method of converting acoustic signals into images and inputting the data into a convolutional neural network for monitoring, and used the method of data augmentation to improve the signal-to-noise ratio of the images; Smirnov detected vocalization signals by calculating the Mel-frequency cepstral coefficients of the vocalization signals of North Atlantic right whales and converting them into images and inputting them into a convolutional neural network; Madhusudhana combined convolutional neural networks with long short-term memory to improve the automatic recognition performance of the network. However, most current methods for identifying underwater vocalization targets through convolutional neural networks are to convert the acoustic signals collected underwater into images by certain means and use convolutional neural networks for identification, and different image conversion means have a great impact on the recognition performance of the network, and at the same time, the influence of image scale transformation on the network recognition efficiency cannot be avoided. Summary of the Invention

[0003] In order to overcome the deficiencies of the prior art, the present invention provides a method for identifying vocalizations of marine mammals based on a one-dimensional convolutional neural network. First, the vocalization samples of marine mammals are preprocessed, and then a one-dimensional convolutional neural network for identifying vocalizations of marine mammals is designed to directly identify acoustic signals, avoiding the interference of the step of converting acoustic signals into images on the recognition results. Compared with the original two-dimensional convolutional neural network, the one-dimensional convolutional neural network designed by the present invention has higher recognition efficiency and better recognition performance.

[0004] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0005] Step 1: Use a database of marine mammal sounds, mark and intercept the vocalization parts in the audio of marine mammal vocalizations in the database to obtain marine mammal vocalization samples as a marine mammal vocalization data set; then randomly divide the marine mammal vocalization data set into a training set and a test set;

[0006] Step 2: Construct a one-dimensional convolutional neural network 1D-AlexNet model for identifying vocalizations of marine mammals;

[0007] The one-dimensional convolutional neural network 1D-AlexNet model includes five convolutional layers, three pooling layers, one spatial pyramid pooling layer, and three fully-connected layers; after the input layer is the first convolutional layer with a kernel size of 1×11; the kernel size of the second convolutional layer is 1×5, and the kernel sizes of the remaining three convolutional layers are all 1×3;

[0008] After the first convolutional layer is the first max-pooling layer with a window size of 1×3 and a stride of 2, followed by the second convolutional layer, and then a second max-pooling layer with a window size of 1×3 and a stride of 2 is connected behind the second convolutional layer. Next are three consecutive convolutional layers, and after the fifth convolutional layer, a third max-pooling layer with a window size of 1×3 and a stride of 2 is connected; the three max-pooling layers are all used to reduce the dimension of the feature matrix;

[0009] After the third max-pooling layer is connected to the spatial pyramid pooling layer, followed by three consecutive fully-connected layers; the Dropout function is used to randomly deactivate 20% of the neurons before the first fully-connected layer and the second fully-connected layer;

[0010] The classification function of the one-dimensional convolutional neural network 1D-AlexNet model all uses the Softmax function, the loss function uses the cross-entropy loss function, and the backpropagation optimization algorithm uses the Adam optimization algorithm;

[0011] Step 3: Input the training set into the one-dimensional convolutional neural network 1D-AlexNet model for marine mammal vocalization recognition for training, and finally realize the vocalization recognition of marine mammals after the training is completed.

[0012] Preferably, the marine mammal sound database is the Mobysound marine mammal sound database.

[0013] Preferably, when the spatial pyramid pooling layer is applied, the number of levels of the spatial pyramid pooling layer and the sizes of the sampling windows at each level need to be set in advance, and then the window size and the moving step of the pooling window inside the spatial pyramid pooling layer are adaptively adjusted according to the sample size input to the model and the size of the spatial pyramid pooling sampling window;

[0014] Suppose a data sample with a size of 1×L and a depth of C is input into the spatial pyramid pooling layer, the size of the sampling window at a certain level in the spatial pyramid pooling layer is 1×N, the window size for the pooling operation at this level is 1×F, and the moving step of the pooling window is S. There are two calculation methods for the pooling window size and the moving step of the pooling window, which are determined by the determination value K, and the calculation method of the K value is:

[0015]

[0016] When the judgment value K satisfies the calculation methods of the pooling window size and its moving step are as follows:

[0017]

[0018]

[0019] When the judgment value K satisfies it is necessary to pad the input data sample. The calculation method of the pooling window size remains unchanged, and the pooling window moving step is:

[0020]

[0021] Among them, when padding the input data sample, both sides of the one-dimensional data sample are padded, and the calculation method of the number of zeros P filled during padding is:

[0022]

[0023] The verification method for the input and output sizes is:

[0024]

[0025] In the formula, W out is the size of the input feature matrix, W in is the data size after zero-padding, F is the convolution kernel size, S is the convolution kernel moving step, and are floor and ceiling respectively.

[0026] The beneficial effects of the present invention are as follows:

[0027] 1. The convergence speed of the 1D-AlexNet model of the present invention is close to that of AlexNet, but it has a higher recognition accuracy on the test set, about 96%. The 1D-AlexNet has a significant decrease in the number of parameters and the training duration of one epoch compared to AlexNet, and at the same time, the recognition accuracy on the test set has also been greatly improved.

[0028] 2. When using the 1D-AlexNet model of the present invention to implement the task of identifying the vocalizations of marine mammals, there is no need to convert the one-dimensional acoustic signal into a two-dimensional image. Therefore, there is no need to consider problems such as scaling, stretching, and distortion of the two-dimensional image, which greatly improves the usage efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a block diagram of the one-dimensional convolutional neural network structure of the present invention.

[0030] Figure 2Schematic diagram of the vocalization dataset of marine mammals, (a) Berardius arnuxii, (b) Berardius bairdii, (c) Indopacetus pacificus.

[0031] Figure 3 Curve showing the change of the recognition accuracy of the test set in the embodiment of the present invention with the number of iterations. Specific implementation manners

[0032] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0033] First, the present invention preprocesses the vocalization samples of marine mammals. Secondly, a one-dimensional convolutional neural network for vocalization recognition of marine mammals is designed to directly recognize the acoustic signals, avoiding the interference of the step of converting the acoustic signals into images on the recognition results.

[0034] 1) In terms of the macroscopic structure, the one-dimensional convolutional neural network is similar to the two-dimensional convolutional neural network, including an input layer, a hidden layer and an output layer. The biggest difference from the two-dimensional convolutional neural network lies in the input layer and the hidden layer.

[0035] The input layer of the two-dimensional convolutional neural network is generally a grayscale image with two channels or an RGB image with three channels, while the input layer of the one-dimensional convolutional neural network is a time-domain signal or a frequency-domain signal.

[0036] In the hidden layer, the convolutional kernel of the one-dimensional convolutional neural network is a one-dimensional strip-shaped convolutional kernel, which is similar to the traditional filter during the convolution operation and slides continuously in a specific direction and step size to traverse each sampling point. In terms of the pooling operation principle, after feature extraction in the convolutional layer, downsampling is performed to reduce the amount of calculation, the number of parameters, etc. The difference lies in the selection of the pooling window. The one-dimensional convolutional neural network selects a one-dimensional window that matches the input signal.

[0037] The fully connected layer in the one-dimensional convolutional neural network has the same function as the two-dimensional convolutional neural network, and is a layer structure used for classification, which non-linearly combines the extracted features to obtain the output. Before the fully connected layer, the one-dimensional convolutional neural network only needs to be unfolded in the depth direction. Since supervised learning is adopted, the output layer outputs the classification labels.

[0038] 2) For the convenience of comparison, the one-dimensional convolutional neural network in the present invention adopts a structure similar to the two-dimensional convolutional neural network AlexNet, including five convolutional layers, three pooling layers and three fully connected layers. The first layer is a convolutional kernel with a size of 1×11 to capture more macroscopic features, the second layer uses a convolutional kernel with a size of 1×5, and then convolutional kernels with a size of 1×3 are used. After the first, second and fifth convolutional layers, a max-pooling layer with a window size of 1×3 and a stride of 2 is followed to reduce the dimension of the feature matrix, and finally multiple fully connected layers are used to improve the non-linear classification performance.

[0039] The classification function of the network all uses the Softmax function, the loss function uses the cross-entropy loss function, and the backpropagation optimization algorithm selects the Adam optimization algorithm. After building the network with the above structure, the Dropout function is used to randomly deactivate 20% of the neurons before the first fully connected layer and the second fully connected layer to prevent overfitting.

[0040] 3) The problems faced by one-dimensional convolutional neural networks and two-dimensional convolutional neural networks are very different. Especially, there are huge differences in the spatial complexity of the model, that is, the number of parameters of the network, which directly affects the memory occupancy of the computer. As the network structure deepens, the model becomes more and more complex, the spatial complexity increases, the number of model parameters increases, especially the number of parameters in the fully connected layer will increase exponentially, and the problem of memory occupancy becomes more and more prominent. Therefore, when designing a one-dimensional convolutional neural network, the problem of spatial complexity cannot be ignored while deepening the network model structure.

[0041] In addition, due to the existence of the fully connected layer in the convolutional neural network, flattening processing is required. Therefore, when using a one-dimensional convolutional neural network model for the vocalization recognition task of marine mammals, the lengths of marine mammal vocalization samples of different lengths generally need to be unified into a certain fixed length and then input into the network model for training, which will greatly waste computing resources or lose the vocalization characteristics of marine mammals.

[0042] To solve the above problems, the present invention adopts the spatial pyramid pooling method and adds a spatial pyramid pooling layer before the fully connected layer. Taking a three-level spatial pyramid as an example, the input-output direction is from bottom to top. The feature matrix of any size with a depth of C obtained by passing the input sample through the convolutional layer is input into the spatial pyramid pooling layer for processing. Entering the spatial pyramid pooling layer, the rightmost pane divides each channel of the input data into 16 parts, so the input data is divided into 16×C parts; similarly, the middle pane divides the input data into 4×C parts, and the left pane divides the input data into 1×C parts. After division, pooling operations are performed on each part respectively, and the input data is transformed into a 21×C-dimensional matrix, which will be stretched into a feature vector with a length of 21C before entering the fully connected layer. In this way, it is ensured that regardless of the dimension of the initial input data, through the spatial pyramid pooling layer, the data is unified to the same dimension before entering the fully connected layer, thus solving the problem that the data samples input into the convolutional neural network need to be of the same size. At the same time, since the dimension of the data input into the fully connected layer is small, the number of parameters of the network is greatly reduced.

[0043] When applying the spatial pyramid pooling layer in the model, it is necessary to preset the number of levels of the spatial pyramid pooling layer and the sizes of the sampling windows at each level. Then, according to the size of the sample input to the model and the size of the sampling window at a certain level of the spatial pyramid pooling, the window size and the moving step of the pooling window for the pooling operation inside that level in the spatial pyramid pooling layer are adaptively adjusted.

[0044] Suppose a data sample with a size of 1×L and a depth of C is input to the spatial pyramid pooling layer. The size of the sampling window at a certain level in the spatial pyramid pooling layer is 1×N, the window size for the pooling operation at this level is 1×F, and the moving step of the pooling window is S. There are two calculation methods for the pooling window size and the moving step of the pooling window, which are determined by the judgment value K. The calculation method of the K value is:

[0045]

[0046] When the judgment value K satisfies the calculation methods for the pooling window size and its moving step are:

[0047]

[0048]

[0049] When the judgment value K satisfies it is necessary to pad the input data sample. The calculation method of the pooling window size remains unchanged, and the moving step of the pooling window is:

[0050]

[0051] Among them, when padding the input data sample, both sides of the one-dimensional data sample are padded. The calculation method of the number of zeros P filled during padding is:

[0052]

[0053] The verification method for the input and output sizes is:

[0054]

[0055] In the formula, W out is the size of the input feature matrix, W in is the size of the data after zero-padding, F is the size of the convolutional kernel, S is the moving step of the convolutional kernel, and are floor and ceiling respectively.

[0056] The final block diagram of the one-dimensional convolutional neural network structure adopted is as shown in Figure 1 shown. Specific embodiments:

[0058] In this embodiment, a one-dimensional convolutional neural network is used to learn and identify / classify the database samples, and compared with the two-dimensional convolutional neural network AlexNet to verify the performance.

[0059] The sample set selected in this embodiment comes from the Mobysound marine mammal sound database. Three types of marine mammals including Berardius arnuxii, Berardius bairdii, and Mesoplodon densirostris are included in the sample set. By analyzing each vocalization sample and marking and intercepting the vocalization part in the marine mammal vocalization audio based on the metadata, marine mammal vocalization samples in various forms including single echolocation signals and continuous echolocation signals are obtained as the marine mammal vocalization dataset for the one-dimensional convolutional neural network. The number of vocalization samples of various marine mammals is shown in Table 1.

[0060] Table 1 Number of vocalization samples of various marine mammals

[0061]

[0062] At the same time, in this embodiment, the marine mammal vocalization audio is transformed into a time-frequency map with a size of 680×540×3 through short-time Fourier transform, thereby constituting a marine mammal vocalization image dataset applicable to the two-dimensional convolutional neural network. The short-time Fourier transform adopts various parameters such as a unified number of Fourier transform points, overlap rate, window size, and Colorbar range, as shown in Table 2.

[0063] Table 2 List of parameters

[0064]

[0065] Since the lengths of the marine mammal vocalization samples are different, before transforming them into time-frequency maps, the vocalization durations of the marine mammals are unified to a certain fixed duration. If the vocalization duration of a marine mammal is less than this fixed duration, zeros are padded before and after this section of the vocalization audio; if the vocalization duration of a marine mammal exceeds this fixed duration, this audio is segmented into multiple segments with this fixed duration as the unit. Figure 2 The figure shows a schematic diagram of the marine mammal vocalization dataset.

[0066] For the AlexNet two-dimensional convolutional neural network model, the pre-processed time-frequency maps of the vocalizations of three types of marine mammals, namely Berardius arnuxii, Berardius bairdii, and Mesoplodon densirostris, are used as the dataset of the model for training and testing. The size of the time-frequency maps is 680×540×3. For the one-dimensional convolutional neural network model, the vocalization audios of the above three types of marine mammals are used as the dataset of the model for training and testing.

[0067] When verifying the network recognition performance, the cross-validation method is adopted. In the vocalization datasets of three types of marine mammals, 10% of the vocalization samples of each type of marine mammal are randomly selected as the test set, and the remaining samples are used as the training set and input into the network model for training. During training, the recognition rates of the network for the training set data and the test set data are recorded in each generation. After the neural network is trained for 20 epochs, the average value of the recognition rates of the last 5 generations when the training tends to be stable for the test set is saved, and it is used as the test set recognition accuracy of this network in this experiment. After independently repeating the above experimental process 5 times, the average value of the 5 results is taken as the final test set accuracy of this network. At the same time, to ensure the reliability of the verification results, when extracting the test set in each verification experiment, the situation where some samples are selected as the test set more than once is excluded. The verification results are shown in Table 3. The curve of the test set recognition accuracy of the model changing with the number of training generations is as Figure 3 shown.

[0068] Table 3 Comparison of Network Recognition Performance

[0069] .

Claims

1. A method for identifying vocalizations of marine mammals based on a one-dimensional convolutional neural network, characterized in that It includes the following steps: Step 1: Use the marine mammal sound database to mark and intercept the vocal parts in the vocal audio of marine mammals in the database, obtain the vocal samples of marine mammals as the marine mammal vocal dataset; then randomly divide the marine mammal vocal dataset into a training set and a test set; Step 2: Construct a one-dimensional convolutional neural network 1D-AlexNet model for marine mammal vocal recognition; The one-dimensional convolutional neural network 1D-AlexNet model includes five convolutional layers, three pooling layers, one spatial pyramid pooling layer, and three fully connected layers; after the input layer is the first convolutional layer with a convolution kernel size of 1×11; the convolution kernel size of the second convolutional layer is 1×5, and the convolution kernel sizes of the remaining three convolutional layers are all 1×3; After the first convolutional layer is the first max pooling layer with a window size of 1×3 and a stride of 2. Next is the second convolutional layer, and behind the second convolutional layer is connected another second max pooling layer with a window size of 1×3 and a stride of 2. Next are three consecutive convolutional layers, and after the fifth convolutional layer is connected a third max pooling layer with a window size of 1×3 and a stride of 2; the three max pooling layers are all used to reduce the dimension of the feature matrix; After the third max pooling layer is connected the spatial pyramid pooling layer, and then are three consecutive fully connected layers; before the first fully connected layer and the second fully connected layer, the Dropout function is used to randomly deactivate 20% of the neurons; When the spatial pyramid pooling layer is applied, the number of levels of the spatial pyramid pooling layer and the sizes of the sampling windows at each level need to be preset in advance, and then the window size and the moving step of the pooling window inside the spatial pyramid pooling layer are adaptively adjusted according to the sample size input to the model and the size of the spatial pyramid pooling sampling window; Input a data sample with a size of and a depth of into the spatial pyramid pooling layer. The size of the sampling window at a certain level in the spatial pyramid pooling layer is , and the window size for the pooling operation at this level is . The moving step of the pooling window is . The calculation methods of the pooling window size and the moving step of the pooling window include: By a determination value to make a determination The calculation method of the value is as follows: When the determination value meets the pooling window size and its moving step calculation method are as follows: When the determination value meets it is necessary to fill the input data sample, and the calculation method of the pooling window size remains unchanged. The moving step size of the pooling window is: Among them, when filling the input data sample, both sides of the one-dimensional data sample are filled, and the number of zeros filled is calculated as follows: The verification method for the input and output size is: Wherein, is the size of the input feature matrix, is the data size after zero-padding, is the size of the convolution kernel, is the convolution kernel moving step size, and are floor and ceiling respectively; The classification function of the one-dimensional convolutional neural network 1D-AlexNet model all uses the Softmax function, the loss function uses the cross-entropy loss function, and the backpropagation optimization algorithm uses the Adam optimization algorithm; Step 3: Input the training set into the one-dimensional convolutional neural network 1D-AlexNet model for marine mammal vocal recognition for training, and finally realize the vocal recognition of marine mammals after the training is completed.

2. The method for identifying vocalizations of marine mammals based on a one-dimensional convolutional neural network according to claim 1, wherein The marine mammal sound database is the Mobysound marine mammal sound database.

Citation Information

Patent Citations

  • Animal oestrus determination method and device

    CN111723785A

  • Marine mammal sounding real-time identification method based on convolutional neural network

    CN113870870A