Whale sound signal recognition method and system based on multi-scale time-frequency feature extraction

By employing adaptive multi-scale linear frequency modulated wavelet transform and an improved convolutional neural network framework, the problem of nonlinear feature processing in cetacean acoustic signal recognition is solved, improving recognition accuracy and robustness, and making it suitable for marine resource exploration and protection.

CN115547347BActive Publication Date: 2025-11-07XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210803510.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-11-07
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

Existing methods for recognizing whale vocal signals struggle to effectively handle the multi-component and nonlinear characteristics of whale signals, resulting in reduced time-frequency resolution and insufficient recognition accuracy. Deep learning networks also lead to the elimination of detailed features, further reducing recognition accuracy.

Method used

An adaptive multi-scale linear frequency modulated wavelet transform is used to extract multi-scale time-frequency parameter features of cetacean acoustic signals, and recognition is performed in an improved convolutional neural network framework. This is achieved by adding a context information extraction module and a bottom-up multi-scale channel attention feature fusion module to the top-down path of the feature pyramid network.

Benefits of technology

It improves the recognition accuracy and robustness of cetacean acoustic signals, and shows significant recognition advantages under different distances, signal-to-noise ratios and Doppler frequency offsets, while reducing information loss during feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115547347B_ABST
    Figure CN115547347B_ABST
Patent Text Reader

Abstract

The application provides a whale sound signal recognition method based on multi-scale time-frequency feature extraction, comprising the following steps: S1, acquiring a sound signal emitted by a whale in the sea; S2, extracting multi-scale time-frequency parameter features of the sound signal by using an adaptive multi-scale linear frequency modulation wavelet transform; and S3, inputting the time-frequency parameter features into an improved convolutional neural network framework for recognition, wherein the improved convolutional neural network framework specifically comprises: adding a context information extraction module to the highest layer of a top-down path of a feature pyramid network, and adding a multi-scale channel attention feature fusion module to a bottom-up path of the feature pyramid network. The application combines the time-frequency parameter features extracted by the adaptive multi-scale linear frequency modulation wavelet transform method and the time-frequency convolutional neural network framework designed based on the time-frequency parameter features, fully extracts and utilizes multi-scale nonlinear features in the whale sound signal, and improves the recognition accuracy and robustness of the whale sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of underwater acoustic signal recognition, and particularly relates to a method and system for recognizing whale acoustic signals based on multi-scale time-frequency feature extraction. BACKGROUND

[0002] Whales are widely distributed and numerous in the ocean, and their importance has attracted more and more interest from researchers. Whales spend most of their time in water, and use sound for communication, echolocation and other social activities. Whale signal detection, recognition and classification are prerequisites for studying the behavior of whales, and are of great significance for the development and protection of marine resources, communication bionics and the protection of sea areas.

[0003] The classical method of studying whales is to use visual methods to detect them, but most species are easy to hear but not easy to see. With the advancement of technology, more and more people now realize that passive acoustic monitoring (PAM) is a good technology for measuring and studying whales. Although PAM can obtain a large amount of sound data, it is often difficult to analyze manually. The development of automatic detection and classification technologies makes it faster and more accurate to analyze whale signals, and can eliminate the human errors often seen in the process of manual detection and classification.

[0004] Whale signals are usually divided into two categories: click signals and whistle signals. Many scholars have studied time-frequency transform methods for recognizing whale acoustic signals. Some scholars have used short-time Fourier transform (STFT), continuous wavelet transform (CWT), Chirplet transform and other methods, but these methods mostly use linear analysis methods and fail to address the multi-component and nonlinear characteristics of whale signals. Although the CWT method can process the nonlinear frequency modulation components in whistle signals, it reduces the time-frequency resolution.

[0005] Meanwhile, there are a large number of studies on the recognition method of whale signals. Some scholars have proposed Gaussian mixture model (GMM) classifier and Hidden Markov Model (HMM) to classify different species or different individuals of organisms, and a large number of scholars have used neural networks and deep learning for whale signal classification and recognition. However, the deep learning network currently used is obtained by training deep abstract features for classification, which can eliminate detailed features, resulting in the loss of some information in the whale signal time-frequency diagram and reducing the recognition accuracy. SUMMARY

[0006] In order to solve the above technical problems, the present application proposes a whale sound signal recognition method and system based on multi-scale time-frequency feature extraction according to the characteristics of whale signals.

[0007] According to the first aspect of the present application, a whale sound signal recognition method based on multi-scale time-frequency feature extraction is proposed, comprising the following steps:

[0008] S1, obtaining the sound signal emitted by the whale in the ocean;

[0009] S2, according to the nonlinear characteristics and the constantly changing chirp rate of the sound signal, using adaptive multi-scale chirplet transform to extract the multi-scale time-frequency parameter characteristics of the sound signal; and

[0010] S3, inputting the time-frequency parameter characteristics into an improved convolutional neural network framework for recognition, wherein the improved convolutional neural network framework specifically comprises: adding a context information extraction module at the highest layer of the top-down path of the feature pyramid network, and adding a multi-scale channel attention feature fusion module in the bottom-up path of the feature pyramid network.

[0011] Preferably, the step S2 specifically comprises:

[0012] S21, performing linear frequency modulation wavelet transform on the sound signal;

[0013] S22, according to the nonlinear characteristics and the chirp rate of the sound signal, respectively determining the window length and the window width of the Gaussian window to meet the condition that the sound signal is nearly stationary in the Gaussian window;

[0014] S23, according to the Gaussian window, obtaining the expression of the adaptive multi-scale chirplet transform, and obtaining the angle parameter according to the expression, so as to extract the time-frequency parameter characteristics containing the angle parameter and the window length.

[0015] Preferably, the determination of the window length in step S22 comprises calculating the standard deviation of the Gaussian window function at each time point, so as to determine the window length of the Gaussian window, so that the acoustic signal is close to stationary in the Gaussian window, and the expression of the Gaussian window function is specifically:

[0016]

[0017] wherein σ(t) is the standard deviation;

[0018] The expression of the window length is specifically:

[0019]

[0020] Preferably, the determination of the window width in step S22 comprises detecting the instantaneous frequency by detecting the wavelet transform ridge line of the acoustic signal, so as to obtain the chirp rate:

[0021]

[0022] wherein v(t) is the instantaneous frequency;

[0023] The conditional expression of the window width is specifically:

[0024]

[0025] wherein the threshold ξ is adjusted so that the acoustic signal is close to stationary in the Gaussian window.

[0026] Preferably, the expression of the adaptive multi-scale linear frequency modulation wavelet transform in step S23 is specifically:

[0027]

[0028] wherein S(t) is the acoustic signal, α m and β n are angle parameters, h(t-tc) is a two-dimensional Gaussian window function, the window length represents the span in the time domain, and the window width represents the span in the frequency domain.

[0029] Preferably, the extraction of the time-frequency parameter feature in step S23 further comprises: according to the time-frequency energy concentration measurement formula:

[0030]

[0031] wherein p is a parameter greater than 1, W is the window length, and L is the length of the acoustic signal;

[0032] The time-frequency energy concentration is maximized by minimizing a time-frequency energy concentration measurement formula CM, so as to obtain the optimal angle parameter and the optimal window length.

[0033] Preferably, the context information extraction module is composed of multiple ratio different multi-path dilated convolution layers, and the multiple dilated convolution layers are closely connected.

[0034] Preferably, the multi-scale channel attention feature fusion module uses a multi-scale channel attention block to fuse different features, and the expression is specifically:

[0035]

[0036] wherein X and Y are input features, M is a weight, Z is a fused feature output by the module, and represents an initial feature fusion operation.

[0037]

[0038] According to a second aspect of the present application, a whale sound signal recognition system based on multi-scale time-frequency feature extraction is provided, comprising:

[0039] A sound signal acquisition module configured to acquire sound signals emitted by whale organisms in the ocean;

[0040] A time-frequency parameter feature extraction module configured to extract multi-scale time-frequency parameter features of the sound signals by using adaptive multi-scale chirplet wavelet transform for nonlinear features and continuously changing chirp rates of the sound signals.

[0041] A sound signal recognition module configured to input the time-frequency parameter features into an improved convolutional neural network framework for recognition, wherein the improved convolutional neural network framework specifically comprises: adding a context information extraction module to the highest layer of a top-down path of a feature pyramid network, and adding a multi-scale channel attention feature fusion module to a bottom-up path of the feature pyramid network.

[0042] According to a third aspect of the present application, a computer readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the whale sound signal recognition method based on multi-scale time-frequency feature extraction according to the first aspect of the present application.

[0043] The application provides a whale sound signal recognition method and system based on multi-scale time-frequency feature extraction. By using an adaptive multi-scale linear frequency modulation wavelet transform method to extract the time-frequency parameter features of the whale sound signal, the nonlinear time-varying characteristics of the whale signal can be effectively addressed, and the information loss in the feature extraction process can be reduced. Meanwhile, a novel time-frequency convolutional neural network framework for the time-frequency parameter features of the sound signal is designed, so that the extracted time-frequency features can be better utilized. By using the novel time-frequency parameter feature extraction method and the classification network framework, the recognition accuracy of the whale sound signal under different distances, different signal-to-noise ratios and different Doppler frequency offsets is improved compared with existing recognition algorithms, and the recognition ability and robustness of the current whale sound signal are improved. The application has important significance for the current marine resource exploration, marine environment development and marine animal protection in China. BRIEF DESCRIPTION OF DRAWINGS

[0044] The accompanying drawings are included to provide a further understanding of embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and, together with the description, serve to explain principles of the application. Other embodiments and many of the intended advantages of the present application will be readily appreciated as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings. The elements of the drawings are not necessarily to scale relative to each other. Like reference numerals designate corresponding similar parts.

[0045] Figure 1 is a whale sound signal recognition method flowchart based on multi-scale time-frequency feature extraction according to an embodiment of the application;

[0046] Figure 2 is an explanation diagram of a linear frequency modulation wavelet transform according to a specific embodiment of the application;

[0047] Figure 3 is a time-frequency parameter feature convolutional neural network framework according to a specific embodiment of the application;

[0048] Figure 4 is a framework diagram of a context information extraction module according to a specific embodiment of the application;

[0049] Figure 5 is a framework diagram of a multi-scale channel attention feature fusion module according to a specific embodiment of the application;

[0050] Figure 6 is a recognition rate comparison diagram of three recognition algorithms under different distances according to a specific embodiment of the application;

[0051] Figure 7 is a recognition rate comparison diagram of three recognition algorithms under different signal-to-noise ratios according to a specific embodiment of the application;

[0052] Figure 8 This is a comparison chart of the recognition rates of three recognition algorithms according to specific embodiments of this application under different Doppler frequency offsets;

[0053] Figure 9 This is a block diagram of a whale acoustic signal recognition system based on multi-scale time-frequency feature extraction according to an embodiment of this application.

[0054] Explanation of reference numerals in the attached diagram: 1. Acoustic signal acquisition module; 2. Time-frequency parameter feature extraction module; 3. Acoustic signal recognition module. Detailed Implementation

[0055] The features and exemplary embodiments of various aspects of this application will now be described in detail. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain this application and are not configured to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.

[0056] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0057] According to the first aspect of this application, a method for whale acoustic signal recognition based on multi-scale time-frequency feature extraction is proposed. Figure 1 A flowchart of a cetacean acoustic signal recognition method based on multi-scale time-frequency feature extraction according to an embodiment of this application is shown, as follows: Figure 1 As shown, the method includes the following steps:

[0058] S1. Acquire acoustic signals emitted by whales in the ocean;

[0059] S2, for the nonlinear characteristics and the constantly changing chirp rate of the acoustic signal, an adaptive multi-scale chirplet transform is used to extract multi-scale time-frequency parameter characteristics of the acoustic signal;

[0060] S3, inputting the time-frequency parameter characteristics into an improved convolutional neural network framework for recognition, the improved convolutional neural network framework specifically comprising: adding a context information extraction module to the highest layer of the top-down path of the feature pyramid network, and adding a multi-scale channel attention feature fusion module to the bottom-up path of the feature pyramid network.

[0061] In a specific embodiment, step S2 comprises the following steps:

[0062] S21, performing a chirplet wavelet transform on the acoustic signal;

[0063] S22, according to the nonlinear characteristics and the chirp rate of the acoustic signal, respectively determining the window length and the window width of the Gaussian window to meet the requirement that the acoustic signal is close to stationary in the Gaussian window;

[0064] S23, according to the Gaussian window, obtaining an expression of the adaptive multi-scale chirplet transform, and obtaining an angle parameter according to the expression, thereby extracting time-frequency parameter characteristics containing the angle parameter and the window length.

[0065] In order to better introduce the adaptive multi-scale chirplet transform (AMSCT) method in step S2 and the improved convolutional neural network framework in step S3, the following will elaborate on these two aspects in a complete method flow.

[0066] 1. Adaptive multi-scale chirplet transform (AMSCT) method

[0067] Traditional linear time-frequency analysis (TFA) is based on the assumption that the signal is piecewise stationary in a short time, but for some strong modulation frequency signals, their instantaneous frequency (IF) is also changing in a short time, resulting in low time-frequency resolution when performing TFA on such signals. Chirplet transform (CT) is a method that uses prior knowledge of the signal to set the parameter value of the phase function, which can more effectively represent the time-frequency characteristics of the frequency modulation signal. The linear chirplet transform of a whale acoustic signal s(t)∈L 2 (R) can be expressed as:

[0068]

[0069] Where z(t) is the analytic signal of s(t).

[0070] Figure 2 An interpretation diagram of the linear frequency modulated wavelet transform according to a specific embodiment of this application is shown, such as... Figure 2 As shown, the solid line represents the actual instantaneous frequency trajectory of the sound signal ω0+λ0t, and the dotted line represents the transformed trajectory, α=tan(θ). It can be intuitively seen that the linear frequency modulated wavelet transform is equivalent to rotating the instantaneous frequency (IF) clockwise by an angle θ, then shifting it upwards by αt0, and finally performing a short-time Fourier transform on the IF using a window function h(t-t0) with a time width of σ. It can be seen that after rotation and shift, the bandwidth intercepted by the window function of the same time width becomes smaller; in other words, with the same time resolution, the frequency resolution of the signal is improved. When α equals the slope of the IF, i.e., the chirp rate of the signal, the frequency resolution reaches its maximum, and the energy of the time-frequency plot is most concentrated.

[0071] However, linear frequency modulated wavelet transform also has some shortcomings. For example, it is often difficult to obtain prior knowledge of actual signals, and the chirp rate and window length cannot be determined. For nonlinear frequency modulated signals, the chirp rate is not a constant, so it is impossible to use a kernel function with a constant chirp rate to effectively analyze the signal at different time periods. Within the same time window, the chirp rate is also fixed, and the kernel function cannot match every time point within the time window.

[0072] For nonlinear frequency modulated signals, the rate of change of instantaneous frequency changes continuously over time. Therefore, to achieve a reasonable balance between time resolution and frequency resolution for nonlinear frequency modulated signals, the time window used should continuously change with the rate of change of instantaneous frequency. Thus, the adaptive multi-scale linear frequency modulated wavelet transform method proposed in this application does not require manual setting of the analysis window length based on prior knowledge of the acoustic signal, but automatically selects the optimal window length based on the time-frequency energy concentration. The Gaussian window function of AMSCT is not fixed, but a function that changes over time. The standard deviation σ(t) of the Gaussian window function at each moment can be obtained through an algorithm, thereby determining the length W(t) of the Gaussian window, allowing the acoustic signal to approach stationarity within this Gaussian window. The specific expression for the Gaussian window function is:

[0073]

[0074] The expression for window length is as follows:

[0075]

[0076] The window width depends on the chirp rate of the acoustic signal, which is the first derivative of the instantaneous frequency. The instantaneous frequency v(t) is first estimated by detecting the ridge of the wavelet transform of the signal. Then, the chirp rate is obtained as:

[0077]

[0078] According to the chirp rate, the quasi-stationary window width X(t) must satisfy the condition (1):

[0079]

[0080] X(t) is adjusted by a threshold ξ so that the signal is quasi-stationary at every time t. For a discrete signal with sampling interval Δt, the discrete chirp rate at the kth time sampling point is:

[0081]

[0082] v[k] is the discrete form of the instantaneous frequency. For a discrete signal, the condition (1) for the above window width can be written as:

[0083]

[0084] After the Gaussian window is determined, the expression of the adaptive multi-scale chirplet transform (AMSCT) can be obtained:

[0085]

[0086] where S(t) is the acoustic signal, α m and β n are the angle parameters, and h(t-tc) is a two-dimensional Gaussian window function. The window length represents the span in the time domain, and the window width represents the span in the frequency domain.

[0087] In the preferred method, since the time-frequency distribution concentration measure can provide a quantitative criterion for evaluating the performance of different distributions, the appropriate time-frequency analysis parameters can be selected according to the value size, so as to obtain the window length that makes the time-frequency energy concentration the highest. The concentration measure formula is:

[0088]

[0089] where p is a parameter greater than 1, W is the window length, and L is the length of the acoustic signal. It is worth noting that the smaller the value of the above formula, the higher the time-frequency distribution concentration. By minimizing CM, the parameters α m and β n are obtained, and the angle parameters α m and β nThe parameter estimation expression for AMSCT can be obtained simultaneously with the window length W, thus achieving optimality.

[0090]

[0091] Angular parameter α m β n The specific calculation steps for the window length W are as follows:

[0092] (1) Input window shift length H, and parameters M, N and K, where M and N are the number of angle segments and K is the number of window lengths.

[0093] (2) Obtain a series of angle parameters and window length based on M and N.

[0094]

[0095]

[0096]

[0097] (3) Obtain a five-dimensional matrix sub-time frequency representation (TFR) for each time center under different angle parameters α, β and window length w.

[0098] (4) Obtain the optimal angle parameters α, β and window length w for each time center according to the AMSCT expression.

[0099] (5) Finally, the TFRs corresponding to the optimal angle parameters α, β and window length w of each time-frequency block are spliced ​​together to form a complete TFR.

[0100] (6) Output the final TFR.

[0101] Thus, the multi-scale time-frequency parameter features of the acoustic signal are extracted.

[0102] 2. Improved Convolutional Neural Network Framework

[0103] After obtaining the time-frequency parameter features estimated by AMSCT, identification can be performed using a convolutional neural network framework designed specifically for these features. Figure 3 A time-frequency parameter feature convolutional neural network framework diagram according to a specific embodiment of this application is shown, such as... Figure 3 As shown, this network is an optimization based on the Feature Pyramid Network (FPN), which consists of two branches: a top-down branch and a bottom-up branch.

[0104] Since the features learned by the deep network have low resolution, in order to avoid the small line component structure in the time-frequency diagram being erased and part of the information being lost, a context information extraction module (CEM) is added to the highest layer of the top-down path of the feature pyramid network to obtain multi-scale feature maps and effectively fuse them, and the ability of the network to extract features is enhanced. Figure 4 A framework diagram of the context information extraction module according to an embodiment of the present application is shown in FIG. 2, which Figure 4 As shown in FIG. 2, the module can extract a large amount of context information from various receptive fields, thereby generating objective features with good recognition effect. The CEM is composed of multiple path dilated convolution layers with different ratios, and the ratios are 3, 6, 12, 18 and 24, respectively. By using these convolution layers, features of multiple different sizes of receptive fields can be obtained. At the same time, a deformable convolution layer is added on each path to improve the performance of geometric transformation, so as to ensure that the CEM can obtain features that are invariant to transformation from the given data. In addition, in order to realize fine fusion of multi-scale information, the CEM uses a tight connection method to combine the output of each dilated convolution layer with the input features, and then sends them to the next dilated convolution layer.

[0105] In the feature pyramid network (FPN), the feature maps of the top-down and bottom-up paths are fused by simple addition operation, but this simple and rough linear fusion method is not suitable for the signals in the present application. Therefore, the present application uses a multi-scale channel attention feature fusion (AFF) module as a component of the bottom-up branch in the FPN network, so that the network can dynamically and adaptively fuse the received features in a context scale perception manner. Figure 5 A framework diagram of the multi-scale channel attention feature fusion module according to an embodiment of the present application is shown in FIG. 3, which Figure 5 As shown in FIG. 3, where C is the number of channels, and the feature map size is HxW. It uses a multi-scale channel attention module (MS-CAM) to fuse different features, and its expression is:

[0106]

[0107] Where X and Y are input features, M is a weight, Z is the fused feature output by the module, and is an initial feature fusion operation.

[0108]

[0109] Since the sum of weights M(X+Y) and 1-M(X+Y) is 1, it is equivalent to taking a weighted average of features X and Y.

[0110] At the end of the network, the output of AFF is passed through a fully connected (FC) network to obtain the classification result.

[0111] In a preferred embodiment, this application also compares the performance of different classification algorithms.

[0112] To demonstrate the performance of the proposed algorithm, the performance of the proposed AMSCT algorithm, the Short Time Fourier Transform (STFT) algorithm, and the Velocity Synchronous Linear Chirplet Transform (VSLCT) algorithm in terms of whale sound signal recognition rate was compared.

[0113] Regarding the data, the whale vocal signal dataset used was derived from the open-source Whale FM dataset. The data was collected by DTAG devices attached to whales. The dataset includes signals from 7 pilot whales and 9 orcas along the coasts of Iceland, Norway, and the Bahamas. The data format is MP3, with a length of 1-8 seconds, totaling approximately 10,000 samples. The dataset can be divided into four categories based on species: short-finned pilot whales in the Bahamas, long-finned pilot whales in Norway, orcas in Iceland, and orcas in Norway.

[0114] The AMSCT algorithm has four input parameters: M, N, K, and H. M, N, and K control the accuracy of parameter estimation, while H controls the time-frequency resolution. Larger values ​​for M, N, and K, and smaller values ​​for H, result in clearer time-frequency plots, but also increase computation time. These parameters can be flexibly adjusted. To balance time-frequency feature resolution and computational efficiency, M, N, K, and H were set to 20, 20, 20, and 160, respectively, during the experiment.

[0115] To investigate the robustness of the identification algorithm to underwater channels, samples with 10 different distances, 10 different signal-to-noise ratios, and 10 different Doppler frequency offsets were tested in the simulated signal test set.

[0116] Figure 6 The following diagram shows a comparison of the recognition rates of three recognition algorithms according to specific embodiments of this application at different distances: Figure 6 As shown, due to the increasing multipath delay and signal attenuation with increasing distance, the recognition rates of the three time-frequency algorithms all show a downward trend. However, AMSCT has a higher recognition rate than the other two algorithms, and the difference becomes more obvious with greater distance.

[0117] Figure 7 A recognition rate comparison chart of three recognition algorithms according to embodiments of the present application under different signal-to-noise ratios is shown in FIG. 3. Figure 7 As shown in FIG. 3, as the signal-to-noise ratio increases, the recognition accuracy of the signal is continuously rising, but the accuracy using the AMSCT algorithm is always higher than that of the other two methods under different signal-to-noise ratios.

[0118] Figure 8 A recognition rate comparison chart of three recognition algorithms according to embodiments of the present application under different Doppler frequency offsets is shown in FIG. 4. Figure 8 As shown in FIG. 4, the recognition rate does not change much under different Doppler frequency offsets, but the recognition performance of AMSCT is better overall.

[0119] In summary, the time-frequency parameter features extracted by the AMSCT method and the time-frequency convolutional neural network framework designed based on the time-frequency parameter features can fully extract and utilize the multi-scale nonlinear features in the whale sound signal, improve the recognition accuracy and robustness of the whale sound signal, and exhibit obvious advantages in different distance, different signal-to-noise ratio, and different Doppler frequency offset test environments.

[0120] According to a second aspect of the present application, a whale sound signal recognition system based on multi-scale time-frequency feature extraction is provided, which is built based on the above recognition method. Figure 9 A whale sound signal recognition system block diagram based on multi-scale time-frequency feature extraction according to embodiments of the present application is shown in FIG. 5. Figure 9 As shown in FIG. 5, the system includes:

[0121] A sound signal acquisition module 1 configured to acquire sound signals emitted by whale organisms in the ocean;

[0122] A time-frequency parameter feature extraction module 2 configured to extract multi-scale time-frequency parameter features of the sound signal using adaptive multi-scale chirplet transform for the nonlinear features and the continuously changing chirp rate of the sound signal;

[0123] A sound signal recognition module 3 configured to input the time-frequency parameter features into an improved convolutional neural network framework for recognition, and the improved convolutional neural network framework specifically includes: adding a context information extraction module to the highest layer of the path from top to bottom of the feature pyramid network, and adding a multi-scale channel attention feature fusion module to the path from bottom to top of the feature pyramid network.

[0124] According to a third aspect of the present application, a computer-readable storage medium storing a computer program is provided, and the computer program implements the whale sound signal recognition method based on multi-scale time-frequency feature extraction according to the first aspect of the present application when executed by a processor.

[0125] The application provides a method and system for recognizing whale sound signals based on multi-scale time-frequency feature extraction and a medium, which extracts time-frequency parameter features of the whale sound signals by using an adaptive multi-scale linear frequency modulation wavelet transform method, can more effectively reduce information loss in the feature extraction process in view of nonlinear time-varying characteristics of the whale signals, and designs a novel time-frequency convolutional neural network framework for the time-frequency parameter features of the sound signals, so that the extracted time-frequency features can be better utilized. By using the novel time-frequency parameter feature extraction method and the classification network framework, the application improves the recognition accuracy of the whale sound signals under different distances, different signal-to-noise ratios, and different Doppler frequency offsets compared with existing recognition algorithms, improves the recognition ability and robustness of the current whale sound signals, and has important significance for the current marine resource exploration, marine environment development, and marine animal protection of China.

[0126] In the embodiments of the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the above-described device / system / method embodiment is only illustrative, for example, the division of the units can be a logical function division, and actual implementation can have another division method, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0127] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place or distributed on multiple units. Part or all of the units can be selected to achieve the purpose of the embodiment according to actual needs.

[0128] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or software functional unit.

[0129] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0130] Obviously, those skilled in the art can make various modifications and changes to the embodiments of the present application without departing from the spirit and scope of the present application. In this way, if these modifications and changes are within the scope of the claims of the present application and their equivalents, the present application also aims to cover these modifications and changes. The word "comprises" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are described in mutually different dependent claims does not mean that the combination of these measures cannot be used to advantage. Any reference signs in the claims should not be considered as limiting the scope.

Claims

1. A method for recognizing a cetacean acoustic signal based on multi-scale time-frequency feature extraction, characterized in that, The method comprises the following steps: S1, acquiring a sound signal emitted by a whale in the ocean; S2, for the nonlinear characteristics and the constantly changing chirp rate of the sound signal, adopting an adaptive multi-scale chirplet wavelet transform to extract multi-scale time-frequency parameter characteristics of the sound signal, specifically comprising: S21, performing a chirplet wavelet transform on the sound signal; S22, according to the nonlinear characteristics and the chirp rate of the sound signal, respectively determining a window length and a window width of a Gaussian window to satisfy that the sound signal is nearly stationary in the Gaussian window; S23, according to the Gaussian window, obtaining an expression of the adaptive multi-scale chirplet wavelet transform, specifically being: ; wherein is an acoustic signal, is an angle parameter, is a two-dimensional Gaussian window function, the window length represents the span in the time domain, the window width represents the span in the frequency domain, and the angle parameter is obtained according to the expression, so as to extract the time-frequency parameter feature containing the angle parameter and the window length; and S3, inputting the time-frequency parameter characteristics into an improved convolutional neural network framework for recognition, the improved convolutional neural network framework specifically comprising: adding a context information extraction module to the highest layer of a path from top to bottom of a feature pyramid network, and adding a multi-scale channel attention feature fusion module to a path from bottom to top of the feature pyramid network.

2. The method of claim 1, wherein, The determination process of the window length in the step S22 specifically comprises: calculating a standard deviation of a Gaussian window function at each time to determine the window length of the Gaussian window, so that the sound signal is nearly stationary in the Gaussian window, and an expression of the Gaussian window function is specifically: ; wherein is the standard deviation; The expression of the window length is specifically: 。 3. The method of claim 1, wherein, The determination process of the window width in the step S22 specifically comprises: estimating an instantaneous frequency by detecting a wavelet transform ridge line of the sound signal to obtain a chirp rate: ; wherein is the instantaneous frequency; Thus, a conditional expression of the window width is specifically: ; wherein the threshold is adjusted such that the acoustic signal is nearly stationary within the Gaussian window.

4. The method of claim 1, wherein, The extraction of the time-frequency parameter characteristics in the step S23 further comprises: according to a time-frequency energy concentration measurement formula: ; wherein, a parameter greater than 1, is the window length, the length of the acoustic signal; By minimizing a time-frequency energy concentration measure formula maximizing the time-frequency energy concentration, thereby obtaining an optimal said angle parameter and an optimal said window length.

5. The method of claim 1, wherein, The context information extraction module is composed of multiple dilated convolution layers with different ratios, and the multiple dilated convolution layers are tightly connected.

6. The method of claim 1, wherein, The multi-scale channel attention feature fusion module uses a multi-scale channel attention module to fuse different features, and an expression thereof is specifically: ; wherein, is an input feature, is a weight, is a fused feature output by the module, denotes an initial feature fusion operation: 。 7. A system for recognition of cetacean acoustic signals based on multiscale time-frequency feature extraction, characterized in that, It comprises: An acoustic signal acquisition module configured to acquire a sound signal emitted by a whale in the ocean; A time-frequency parameter characteristic extraction module configured to, for nonlinear characteristics and constantly changing chirp rates of the sound signal, adopt an adaptive multi-scale chirplet wavelet transform to extract multi-scale time-frequency parameter characteristics of the sound signal, specifically comprising: S21, performing a chirplet wavelet transform on the sound signal; S22, according to the nonlinear characteristics and the chirp rate of the sound signal, respectively determining a window length and a window width of a Gaussian window to satisfy that the sound signal is nearly stationary in the Gaussian window; S23, according to the Gaussian window, obtaining an expression of the adaptive multi-scale chirplet wavelet transform, specifically being: ; wherein is an acoustic signal, is an angle parameter, is a two-dimensional Gaussian window function, the window length represents the span in the time domain, the window width represents the span in the frequency domain, and the angle parameter is obtained according to the expression, thereby extracting the time-frequency parameter feature containing the angle parameter and the window length; and An acoustic signal recognition module configured to input the time-frequency parameter characteristics into an improved convolutional neural network framework for recognition, the improved convolutional neural network framework specifically comprising: adding a context information extraction module to the highest layer of a path from top to bottom of a feature pyramid network, and adding a multi-scale channel attention feature fusion module to a path from bottom to top of the feature pyramid network.

8. A computer readable storage medium storing a computer program which, when executed by a processor, implements the method of any one of claims 1-6.