Method and system for detecting whale overlapped sound events based on adaptive multi-scale synthetic attention

Through the adaptive multi-scale synthetic attention network, the complex time-frequency structure in the marine environment is dynamically captured, the problem of identifying the subtle behavioral states of marine biological sound sources is solved, and high-precision detection of overlapping sound events of whales is achieved.

CN120673770AActive Publication Date: 2025-09-19ANHUI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511142084.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-19
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively distinguishing the subtle behavioral states of marine organisms in marine environments, and lack detection robustness in complex sound source environments, especially under low signal-to-noise ratios, making it difficult to accurately identify overlapping sound events of marine organisms such as whales.

Method used

Adopting the adaptive multi-scale synthetic attention method, by constructing a time-frequency-aware cross-offset feature extraction network, an adaptive window size prediction network and a multi-scale attention fusion network, the complex time-frequency structure and multi-scale features are dynamically captured to achieve high-precision detection of overlapping whale sound events.

Benefits of technology

It improves the detection accuracy and generalization ability of overlapping whale sound events in complex ocean environments, enhances the feature separation ability of weak targets and overlapping signals, dynamically matches the feature receptive fields of targets of different scales, and realizes the saliency enhancement of multi-scale features and reconstruction of complementary information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673770A_ABST
    Figure CN120673770A_ABST
Patent Text Reader

Abstract

The invention discloses a whale overlapped sound event detection method and system based on adaptive multi-scale synthetic attention, and belongs to the technical field of ocean engineering and ocean signals, and the method comprises the steps: collecting historical whale sound signals, and carrying out the preprocessing and ACT transformation of the historical whale sound signals, and obtaining a transformed data set; constructing a time-frequency perception cross-offset feature extraction network, and training the time-frequency perception cross-offset feature extraction network by using the transformed data to obtain an adaptive multi-scale synthesis attention network; constructing an adaptive window size prediction network, and predicting window parameters required for constructing a multi-scale attention fusion network according to a feature map output by the adaptive multi-scale synthesis attention network; constructing a multi-scale attention fusion network based on the window parameters, and carrying out iteration to obtain a whale overlapped sound event detection model; and acquiring a real-time whale sound signal, and detecting the real-time whale sound signal by using the detection model to obtain a detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of ocean engineering and ocean signal technology, and particularly relates to a method and system for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention. Background Art

[0002] Overlapping sound event detection is of great significance in intelligent monitoring of marine ecology, and can realize the autonomous identification and precise positioning of the sound sources of marine organisms such as whales and dolphins. However, the types of sound sources in the marine environment are complex, and different organisms often make sounds at the same time, accompanied by strong background noise, resulting in obvious time-frequency overlap, non-stationarity and energy differences in the received signals. In addition, the sounds of marine organisms are species-dependent, and small changes in frequency often correspond to different semantics or behavioral states. Convolution is shift-invariant and has difficulty capturing small differences in frequency. At the same time, the duration of different events varies significantly, and fixed-scale contexts are difficult to adapt to diverse temporal features.

[0003] In order to improve the frequency structure perception ability of the model, researchers have proposed methods such as frequency dynamic convolution, frequency pyramid module and two-dimensional separable convolution to overcome the translation invariance assumption and improve the detection performance to a certain extent. However, the existing methods still have many limitations. First, the semantic differences implied by small changes in frequency are insufficiently modeled, making it difficult to accurately distinguish the subtle behavioral states of organisms. Secondly, context modeling of a unified scale is difficult to adapt to sound events of different durations, affecting the complete representation of events of varying lengths. In addition, there is a lack of explicit modeling mechanism for the overlapping relationships between multiple events, and detection robustness is still insufficient, especially under low signal-to-noise ratios. To this end, there is an urgent need to construct a collaborative detection framework that integrates frequency and behavior perception, multi-scale context understanding and structural interpretability to more effectively improve detection accuracy and generalization capabilities in complex marine environments. Summary of the Invention

[0004] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:

[0005] A method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention includes the following steps:

[0006] collecting historical cetacean sound signals, performing preprocessing and ACT transformation on the historical cetacean sound signals, and obtaining a transformed data set;

[0007] Constructing a time-frequency-aware cross-offset feature extraction network, and using the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network;

[0008] Constructing an adaptive window size prediction network to predict window parameters required for constructing a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthetic attention network;

[0009] The multi-scale attention fusion network is constructed based on the window parameters and iterated to obtain a whale overlapping sound event detection model;

[0010] A real-time cetacean sound signal is acquired, and the real-time cetacean sound signal is detected using the cetacean overlapping sound event detection model to obtain a detection result.

[0011] Preferably, the method for obtaining the transformed data set includes:

[0012] collecting historical cetacean sound signals, and performing non-overlapping equal-length frame cropping on the historical cetacean sound signals to obtain cropped signals;

[0013] The foreground signal and the background sound in the cropped signal are synthesized using a soundscape synthesizer library to obtain a synthesized signal, and then the ACT feature of the synthesized signal is extracted and normalized to obtain the transformed data set:

[0014] ,

[0015] Among them, x N Indicates different ACT characteristics.

[0016] Preferably, the method for obtaining an adaptive multi-scale synthetic attention network includes:

[0017] Constructing a time-frequency-aware cross-offset feature extraction network, the time-frequency-aware cross-offset feature extraction network comprising: a first TFER convolution unit, a channel expansion unit, a frequency-domain dynamic convolution unit, a time-domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit, and a second TFER convolution unit;

[0018] Dividing the transformed data into a training set, a validation set, and a test set according to a preset ratio;

[0019] The time-frequency-aware cross-offset feature extraction network is trained using the training set, and the trained network is verified using the verification set to obtain an adaptive multi-scale synthetic attention network.

[0020] Preferably, the method for obtaining the window parameters includes:

[0021] Construct an adaptive window size prediction network, and convolve the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator to obtain the gradient map vector:

[0022] ,

[0023] in, Represents the gradient map obtained by convolving the Nth feature map with the Sobel operator;

[0024] The gradient map vector is squared to obtain the edge strength map vector:

[0025] ,

[0026] Among them, e N Represents the edge intensity map obtained after the Nth gradient map vector is squared;

[0027] Based on the edge intensity map vector, a global edge intensity vector, a maximum response intensity vector and a significant edge density vector are constructed:

[0028] ,

[0029] ,

[0030] ,

[0031] ,

[0032] ,

[0033] ,

[0034] Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map;

[0035] Based on the global edge intensity vector, the maximum response intensity vector and the significant edge density vector, a large window complexity vector and a small window complexity vector are constructed:

[0036] ,

[0037] ,

[0038] in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Indicates the complexity parameter of the Nth small window;

[0039] Based on the large window complexity vector and the small window complexity vector, a large window vector and a small window vector are constructed, and the window parameters are obtained:

[0040] ,

[0041] ,

[0042] in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

[0043] Preferably, the method for obtaining the cetacean overlapping sound event detection model comprises:

[0044] Constructing the multi-scale attention fusion network based on the window parameters, the multi-scale attention fusion network comprising: a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function;

[0045] The multi-scale attention fusion network is iterated, and the iterated model is tested using the test set. If the preset target is met, the whale overlapping sound event detection model is obtained.

[0046] The present invention also provides a whale overlapping sound event detection system based on adaptive multi-scale synthetic attention, wherein the detection system applies any of the above-mentioned detection methods and comprises: a historical data acquisition module, a first model construction module, a window parameter acquisition module, a second model construction module, and a detection module;

[0047] The historical data acquisition module is used to collect historical whale sound signals, perform preprocessing and ACT transformation on the historical whale sound signals, and obtain a transformed data set;

[0048] The first model building module is used to build a time-frequency-aware cross-offset feature extraction network, and use the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network;

[0049] The window parameter acquisition module is used to construct an adaptive window size prediction network, and predict the window parameters required to construct a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthesis attention network;

[0050] The second model construction module constructs the multi-scale attention fusion network based on the window parameters and iterates to obtain a whale overlapping sound event detection model;

[0051] The detection module is used to obtain real-time cetacean sound signals and detect the real-time cetacean sound signals using the cetacean overlapping sound event detection model to obtain a detection result.

[0052] Preferably, the workflow of the historical data collection module includes:

[0053] collecting historical cetacean sound signals, and performing non-overlapping equal-length frame cropping on the historical cetacean sound signals to obtain cropped signals;

[0054] The foreground signal and the background sound in the cropped signal are synthesized using a soundscape synthesizer library to obtain a synthesized signal, and then the ACT feature of the synthesized signal is extracted and normalized to obtain the transformed data set:

[0055] ,

[0056] Among them, x N Indicates different ACT characteristics.

[0057] Preferably, the workflow of the first model building module includes:

[0058] Constructing a time-frequency-aware cross-offset feature extraction network, the time-frequency-aware cross-offset feature extraction network comprising: a first TFER convolution unit, a channel expansion unit, a frequency-domain dynamic convolution unit, a time-domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit, and a second TFER convolution unit;

[0059] Dividing the transformed data into a training set, a validation set, and a test set according to a preset ratio;

[0060] The time-frequency-aware cross-offset feature extraction network is trained using the training set, and the trained network is verified using the verification set to obtain an adaptive multi-scale synthetic attention network.

[0061] Preferably, the workflow of the window parameter acquisition module includes:

[0062] Construct an adaptive window size prediction network, and convolve the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator to obtain the gradient map vector:

[0063] ,

[0064] in, Represents the gradient map obtained by convolving the Nth feature map with the Sobel operator;

[0065] The gradient map vector is squared to obtain the edge strength map vector:

[0066] ,

[0067] Among them, e N Represents the edge intensity map obtained after the Nth gradient map vector is squared;

[0068] Based on the edge intensity map vector, a global edge intensity vector, a maximum response intensity vector and a significant edge density vector are constructed:

[0069] ,

[0070] ,

[0071] ,

[0072] ,

[0073] ,

[0074] ,

[0075] Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map;

[0076] Based on the global edge intensity vector, the maximum response intensity vector and the significant edge density vector, a large window complexity vector and a small window complexity vector are constructed:

[0077] ,

[0078] ,

[0079] in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Indicates the complexity parameter of the Nth small window;

[0080] Based on the large window complexity vector and the small window complexity vector, a large window vector and a small window vector are constructed, and the window parameters are obtained:

[0081] ,

[0082] ,

[0083] in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

[0084] Preferably, the workflow of the second model building module includes:

[0085] Constructing the multi-scale attention fusion network based on the window parameters, the multi-scale attention fusion network comprising: a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function;

[0086] The multi-scale attention fusion network is iterated, and the iterated model is tested using the test set. If the preset target is met, the whale overlapping sound event detection model is obtained.

[0087] Compared with the prior art, the present invention has the following beneficial effects:

[0088] (1) The time-frequency-aware cross-offset feature extraction network model of the present invention is combined with the ACT transformation to obtain a high-resolution time-frequency map. By using the cross-offset mechanism and the time-frequency dynamic convolution strategy, it can adaptively capture the key feature areas in the complex time-frequency structure, enhance the spatial sensitivity and time-frequency coupling expression ability of the feature map, and improve the feature separation ability of weak targets or overlapping signals.

[0089] (2) The adaptive window size prediction network model of the present invention combines the semantic information of the deep feature map to dynamically predict the optimal window size, match the feature receptive field of targets of different scales, avoid information loss or redundancy caused by fixed windows, and improve the focusing ability of the multi-scale attention module and the time-frequency structure adaptability of the network.

[0090] (3) The multi-scale attention fusion module model of the present invention fuses the deep time-frequency feature maps extracted by the multi-branch network and combines it with the local dense synthetic attention mechanism to achieve saliency enhancement of multi-scale features and reconstruction of complementary information, effectively improving the recognition accuracy of signals with drastic changes in time-frequency distribution and complex component intersections. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0092] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0093] Figure 2 Schematic diagram of the structure of the TFER convolution unit according to an embodiment of the present invention;

[0094] Figure 3 A schematic diagram of the structure of a time-frequency-aware cross-offset feature extraction network according to an embodiment of the present invention;

[0095] Figure 4 Schematic diagram of the structure of a multi-scale attention fusion network according to an embodiment of the present invention;

[0096] Figure 5 Schematic diagram of the confusion matrix of the signal classification structure in Example 4 of the present invention;

[0097] Figure 6 Schematic diagram of PSDS score changes under different signal-to-noise ratios for the detection model in Example 6 of the present invention. DETAILED DESCRIPTION

[0098] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0099] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0100] Example 1:

[0101] In this embodiment, if Figure 1 As shown, a method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention includes the following steps:

[0102] S1. Collect historical whale sound signals, perform preprocessing and ACT transformation on the historical whale sound signals, and obtain a transformed data set.

[0103] In this embodiment, the method for obtaining the transformed data set includes:

[0104] Historical cetacean sound signals were collected and processed by cropping different types of historical cetacean sound signals with non-overlapping frames of equal length, eliminating invalid or low-energy segments to obtain valid foreground signals of varying lengths, i.e., the cropped signals. The foreground signals in the cropped signals were synthesized with the background sounds using the soundscape synthesizer library to obtain synthetic signals. The synthetic audio data was generated under four different signal-to-noise ratio (SNR) conditions: Very Low (−10dB to −5dB), Low (−5dB to 0dB), Medium (0dB to 5dB), and High (5dB to 10dB). The ACT features of the synthetic signals were then extracted and normalized to obtain the transformed dataset:

[0105] ,

[0106] Among them, x N Indicates different ACT characteristics.

[0107] S2. Construct a time-frequency-aware cross-offset feature extraction network and use the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network.

[0108] The method of obtaining the adaptive multi-scale synthetic attention network includes:

[0109] Construct a time-frequency-aware cross-offset feature extraction network. The time-frequency-aware cross-offset feature extraction network is as follows: Figure 3As shown, it includes: a first TFER convolution unit, a channel expansion unit, a frequency domain dynamic convolution unit, a time domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit and a second TFER convolution unit; the transformed data is divided into a training set, a validation set and a test set according to a preset ratio. In this embodiment, the preset ratio is 7:1.5:1.5; the time-frequency perception cross-shift feature extraction network is trained using the training set, and the trained network is verified using the validation set to obtain an adaptive multi-scale synthetic attention network. The TFER convolution unit is shown as Figure 2 shown.

[0110] In this embodiment, the workflow of the adaptive multi-scale synthetic attention network includes: the data of the training set is extracted through the first TFER convolution unit to extract the deep features of all ACT features X, and the output deep feature map is:

[0111] ,

[0112] Among them, f N Represents the deep features extracted from the transformed dataset during the Nth training; after being expanded 3 times by the channel expansion unit, the deep feature map F is expanded to F c :

[0113] ,

[0114] Among them, F1, F2 and F3 represent the feature data after the channel is expanded 3 times respectively; after F1 passes through the frequency domain dynamic convolution unit and the first cross shift unit, the output feature for:

[0115] ,

[0116] in, It represents the deep features extracted by the Nth feature data F1 after the frequency domain dynamic convolution unit and the first cross displacement unit; after F3 passes through the time domain dynamic convolution unit and the second cross displacement unit, the output feature for:

[0117] ,

[0118] in, Indicates that the Nth feature data F3 is extracted through the time domain dynamic convolution unit and the second cross displacement unit. Then the three branches are , F2 and After inputting into the weight fusion unit, we get F w :

[0119] ,

[0120] in, Indicates the Nth feature data , F2 and The characteristics after the corresponding elements are fused; the final F w The deep features after fusion are extracted by the second TFER convolution unit to obtain F o for:

[0121] ,

[0122] in, Indicates the Nth feature data F w The deep features extracted by the second TFER convolutional unit.

[0123] S3. Construct an adaptive window size prediction network to predict the window parameters required to construct a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthetic attention network.

[0124] The method for obtaining window parameters includes: building an adaptive window size prediction network, convolving the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator, and obtaining the gradient map vector:

[0125] ,

[0126] in, Represents the gradient map obtained by convolving the Nth feature map with the Sobel operator; square the gradient map vector to obtain the edge strength map vector:

[0127] ,

[0128] Among them, e N Represents the edge intensity map obtained after the Nth gradient map vector is squared; based on the edge intensity map vector, the global edge intensity vector, the maximum response intensity vector and the significant edge density vector are constructed:

[0129] ,

[0130] ,

[0131] ,

[0132] ,

[0133] ,

[0134] ,

[0135] Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map; based on the global edge intensity vector, the maximum response intensity vector, and the significant edge density vector, a large window complexity vector and a small window complexity vector are constructed:

[0136] ,

[0137] ,

[0138] in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Represents the Nth small window complexity parameter; based on the large window complexity vector and the small window complexity vector, construct the large window vector and the small window vector, and obtain the window parameters:

[0139] ,

[0140] ,

[0141] in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

[0142] S4. A multi-scale attention fusion network is constructed based on window parameters and iterated to obtain a whale overlapping sound event detection model.

[0143] The method for obtaining a whale overlapping sound event detection model includes: constructing a multi-scale attention fusion network based on window parameters, the multi-scale attention fusion network is as follows: Figure 4As shown, it includes: a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function; the multi-scale attention fusion network is iterated, and the iterated model is tested using a test set. If the preset goals are met, a whale overlapping sound event detection model is obtained.

[0144] In this embodiment, the workflow of the multi-scale attention fusion network includes: extracting the deep features F o The large and small scale feature vectors are obtained by inputting them into the network through the input unit and passing through the multi-scale attention fusion unit. and for:

[0145] ,

[0146] ,

[0147] in, Represents the Nth deep feature F o The large-scale feature vector obtained after the multi-scale large-scale attention fusion unit, Represents the Nth deep feature F o The small-scale feature vector is obtained after the multi-scale small-scale attention fusion unit; the large-scale feature vector With small-scale eigenvectors The deep feature vector Y obtained by weight fusion through the weight fusion unit is:

[0148] ,

[0149] Among them, α represents The weight of The weight of the output fusion enhanced deep feature vector Y is obtained by weight fusion, passes through the GRU unit, then passes through the linear layer, and finally passes through the loss function to obtain the detection result.

[0150] S5. Acquire real-time cetacean sound signals, and detect the real-time cetacean sound signals using a cetacean overlapping sound event detection model to obtain detection results.

[0151] In this example, the detection results cover four signal-to-noise ratio (SNR) levels: very low (-10 dB to -5 dB), low (-5 dB to 0 dB), medium (0 dB to 5 dB), and high (5 dB to 10 dB). The detection results include four typical North Atlantic right whale acoustic signal types: upcall, gunshot, scream, and moancall. Each audio file is accurately timestamped, indicating the start and end time of each acoustic event and its category.

[0152] like Figure 5 Figure 2 shows the classification confusion matrix for four types of cetacean acoustic events. As can be seen, all types of events exhibit high classification accuracy. Moancalls achieved the highest recognition accuracy, reaching 93.6%, while screams achieved an accuracy of 81.8%. Although some events exhibited some degree of confusion, such as a high misidentification rate between upcalls and moancalls, the model generally performed well in distinguishing between the various categories.

[0153] like Figure 6 As shown in the figure, the PSDS scores of Conformer, CRNN, PANNs, CNN-Transformer, TDNN-LSTM and the model AWMSA of this embodiment are changed at different signal-to-noise ratios. The model is in the leading position at all four signal-to-noise ratios.

[0154] Example 2:

[0155] In this embodiment, a whale overlapping sound event detection system based on adaptive multi-scale synthetic attention includes: a historical data acquisition module, a first model construction module, a window parameter acquisition module, a second model construction module and a detection module.

[0156] The historical data acquisition module is used to collect historical whale sound signals, preprocess and perform ACT transformation on the historical whale sound signals, and obtain the transformed data set.

[0157] The workflow of the historical data acquisition module includes: collecting historical cetacean sound signals, cropping them into non-overlapping frames of equal length to obtain a cropped signal; synthesizing the foreground signal and background sound in the cropped signal using the soundscape synthesizer library to obtain a synthesized signal; extracting the ACT features of the synthesized signal, and normalizing the features to obtain a transformed dataset:

[0158] ,

[0159] Among them, x NIndicates different ACT characteristics.

[0160] The first model building module is used to construct a time-frequency-aware cross-offset feature extraction network, and use the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network.

[0161] The workflow of the first model construction module includes: constructing a time-frequency-aware cross-offset feature extraction network, which includes: a first TFER convolution unit, a channel expansion unit, a frequency domain dynamic convolution unit, a time domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit and a second TFER convolution unit; dividing the transformed data into a training set, a validation set and a test set according to a preset ratio; using the training set to train the time-frequency-aware cross-offset feature extraction network, and using the validation set to verify the trained network to obtain an adaptive multi-scale synthetic attention network.

[0162] The window parameter acquisition module is used to build an adaptive window size prediction network and predict the window parameters required to build a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthetic attention network.

[0163] The workflow of the window parameter acquisition module includes: building an adaptive window size prediction network, convolving the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator, and obtaining the gradient map vector:

[0164] ,

[0165] in, Represents the gradient map obtained by convolving the i-th feature map with the Sobel operator, i=1,2,...,N, where N is a natural number; square the gradient map vector to obtain the edge strength map vector:

[0166] ,

[0167] Among them, e i Represents the edge intensity map obtained after the ith gradient map vector is squared, i=1,2,...,N, where N represents a natural number. Based on the edge intensity map vector, the global edge intensity vector, the maximum response intensity vector, and the significant edge density vector are constructed:

[0168] ,

[0169] ,

[0170] ,

[0171] ,

[0172] ,

[0173] ,

[0174] Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map;

[0175] Based on the global edge intensity vector, the maximum response intensity vector, and the significant edge density vector, the large window complexity vector and the small window complexity vector are constructed:

[0176] ,

[0177] ,

[0178] in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Indicates the complexity parameter of the Nth small window;

[0179] Based on the large window complexity vector and the small window complexity vector, construct the large window vector and the small window vector, and obtain the window parameters:

[0180] ,

[0181] ,

[0182] in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

[0183] The second model building module constructs a multi-scale attention fusion network based on window parameters and iterates to obtain a whale overlapping sound event detection model.

[0184] The workflow of the second model construction module includes: constructing a multi-scale attention fusion network based on window parameters, the multi-scale attention fusion network consists of a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function; iterating the multi-scale attention fusion network and testing the iterated model using a test set. If the preset goals are met, a whale overlapping sound event detection model is obtained.

[0185] The detection module is used to obtain real-time whale sound signals and use the whale overlapping sound event detection model to detect the real-time whale sound signals to obtain detection results.

[0186] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention, characterized in that: The following steps are involved: collecting historical cetacean sound signals, performing preprocessing and ACT transformation on the historical cetacean sound signals, and obtaining a transformed data set; Constructing a time-frequency-aware cross-offset feature extraction network, and using the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network; Constructing an adaptive window size prediction network to predict window parameters required for constructing a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthetic attention network; The multi-scale attention fusion network is constructed based on the window parameters and iterated to obtain a whale overlapping sound event detection model; A real-time cetacean sound signal is acquired, and the real-time cetacean sound signal is detected using the cetacean overlapping sound event detection model to obtain a detection result.

2. The method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention according to claim 1, characterized in that: The method for obtaining the transformed data set includes: collecting historical cetacean sound signals, and performing non-overlapping equal-length frame cropping on the historical cetacean sound signals to obtain cropped signals; The foreground signal and the background sound in the cropped signal are synthesized using a soundscape synthesizer library to obtain a synthesized signal, and then the ACT feature of the synthesized signal is extracted and normalized to obtain the transformed data set: , Among them, x N Indicates different ACT characteristics.

3. The method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention according to claim 1, characterized in that: The method of obtaining the adaptive multi-scale synthetic attention network includes: Constructing a time-frequency-aware cross-offset feature extraction network, the time-frequency-aware cross-offset feature extraction network comprising: a first TFER convolution unit, a channel expansion unit, a frequency-domain dynamic convolution unit, a time-domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit, and a second TFER convolution unit; Dividing the transformed data into a training set, a validation set, and a test set according to a preset ratio; The time-frequency-aware cross-offset feature extraction network is trained using the training set, and the trained network is verified using the verification set to obtain an adaptive multi-scale synthetic attention network.

4. The method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention according to claim 3, characterized in that: The method for obtaining the window parameters includes: Construct an adaptive window size prediction network, and convolve the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator to obtain the gradient map vector: , in, Represents the gradient map obtained by convolving the Nth feature map with the Sobel operator; The gradient map vector is squared to obtain the edge strength map vector: , Among them, e N Represents the edge intensity map obtained after the Nth gradient map vector is squared; Based on the edge intensity map vector, a global edge intensity vector, a maximum response intensity vector and a significant edge density vector are constructed: , , , , , , Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map; Based on the global edge intensity vector, the maximum response intensity vector and the significant edge density vector, a large window complexity vector and a small window complexity vector are constructed: , , in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Indicates the complexity parameter of the Nth small window; Based on the large window complexity vector and the small window complexity vector, a large window vector and a small window vector are constructed, and the window parameters are obtained: , , in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

5. The method for detecting overlapping cetacean sound events based on adaptive multi-scale synthetic attention according to claim 4, characterized in that: The method for obtaining the cetacean overlapping sound event detection model includes: Constructing the multi-scale attention fusion network based on the window parameters, the multi-scale attention fusion network comprising: a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function; The multi-scale attention fusion network is iterated, and the iterated model is tested using the test set. If the preset target is met, the whale overlapping sound event detection model is obtained.

6. A cetacean overlapping sound event detection system based on adaptive multi-scale synthetic attention, wherein the detection system applies the detection method according to any one of claims 1 to 5, characterized in that: include: Historical data acquisition module, first model construction module, window parameter acquisition module, second model construction module and detection module; The historical data acquisition module is used to collect historical whale sound signals, perform preprocessing and ACT transformation on the historical whale sound signals, and obtain a transformed data set; The first model building module is used to build a time-frequency-aware cross-offset feature extraction network, and use the transformed data to train the time-frequency-aware cross-offset feature extraction network to obtain an adaptive multi-scale synthetic attention network; The window parameter acquisition module is used to construct an adaptive window size prediction network, and predict the window parameters required to construct a multi-scale attention fusion network based on the feature map output by the adaptive multi-scale synthesis attention network; The second model construction module constructs the multi-scale attention fusion network based on the window parameters and iterates to obtain a whale overlapping sound event detection model; The detection module is used to obtain real-time cetacean sound signals and detect the real-time cetacean sound signals using the cetacean overlapping sound event detection model to obtain a detection result.

7. The cetacean overlapping sound event detection system based on adaptive multi-scale synthetic attention according to claim 6, characterized in that: The workflow of the historical data acquisition module includes: collecting historical cetacean sound signals, and performing non-overlapping equal-length frame cropping on the historical cetacean sound signals to obtain cropped signals; The foreground signal and the background sound in the cropped signal are synthesized using a soundscape synthesizer library to obtain a synthesized signal, and then the ACT feature of the synthesized signal is extracted and normalized to obtain the transformed data set: , Among them, x N Indicates different ACT characteristics.

8. The cetacean overlapping sound event detection system based on adaptive multi-scale synthetic attention according to claim 6, characterized in that: The workflow of the first model building module includes: Constructing a time-frequency-aware cross-offset feature extraction network, the time-frequency-aware cross-offset feature extraction network comprising: a first TFER convolution unit, a channel expansion unit, a frequency-domain dynamic convolution unit, a time-domain dynamic convolution unit, a first cross-shift unit, a second cross-shift unit, a first weight fusion unit, and a second TFER convolution unit; Dividing the transformed data into a training set, a validation set, and a test set according to a preset ratio; The time-frequency-aware cross-offset feature extraction network is trained using the training set, and the trained network is verified using the verification set to obtain an adaptive multi-scale synthetic attention network.

9. The cetacean overlapping sound event detection system based on adaptive multi-scale synthetic attention according to claim 8, characterized in that: The workflow of the window parameter acquisition module includes: Construct an adaptive window size prediction network, and convolve the feature map output by the adaptive multi-scale synthetic attention network with the Sobel operator to obtain the gradient map vector: , in, Represents the gradient map obtained by convolving the Nth feature map with the Sobel operator; The gradient map vector is squared to obtain the edge strength map vector: , Among them, e N Represents the edge intensity map obtained after the Nth gradient map vector is squared; Based on the edge intensity map vector, a global edge intensity vector, a maximum response intensity vector and a significant edge density vector are constructed: , , , , , , Where Ω represents the global edge strength vector, Ψ represents the maximum response strength vector, and Φ represents the significant edge density vector. represents the Nth global edge strength, represents the Nth maximum response intensity, Indicates the Nth significant edge density, n f Represents the frequency dimension of the edge strength map, n t represents the time dimension of the edge intensity map, i represents the i-th frequency dimension of the edge intensity map, j represents the j-th time dimension of the edge intensity map, I represents the threshold function, and L represents a pixel value of the edge intensity map; Based on the global edge intensity vector, the maximum response intensity vector and the significant edge density vector, a large window complexity vector and a small window complexity vector are constructed: , , in, represents the large window complexity vector, represents the small window complexity vector, Indicates the complexity parameter of the Nth large window, Indicates the complexity parameter of the Nth small window; Based on the large window complexity vector and the small window complexity vector, a large window vector and a small window vector are constructed, and the window parameters are obtained: , , in, represents a large window vector, represents a small window vector, Indicates the Nth largest window parameter, Indicates the Nth small window parameter.

10. The cetacean overlapping sound event detection system based on adaptive multi-scale synthetic attention according to claim 9, characterized in that: The workflow of the second model building module includes: Constructing the multi-scale attention fusion network based on the window parameters, the multi-scale attention fusion network comprising: a small-scale input unit, a large-scale input unit, a small-scale attention fusion unit, a large-scale attention fusion unit, a second weight fusion unit, a GRU unit, a linear layer and a loss function; The multi-scale attention fusion network is iterated, and the iterated model is tested using the test set. If the preset target is met, the whale overlapping sound event detection model is obtained.

Citation Information

Patent Citations

  • Multi-scale time-frequency feature extraction-based whale sound signal identification method and system

    CN115547347A

  • OSA detection method and device based on moving window self-attention model

    CN117598662A

  • Heart sound signal noise detection method based on multi-scale convolutional network

    CN119207466A

  • Multivariable time series prediction method based on Patching and multi-scale feature extraction

    CN120316427A

  • Discrimination of components of audio signals based on multiscale spectro-temporal modulations

    US20060025989A1