Semantic segmentation detection method and device for marine mammal tick signal and medium

By introducing Bayesian convolution and wavelet transform convolution into the encoder-decoder architecture, acoustic image semantic segmentation network model is constructed, which solves the problem of false detection and missed detection in the acoustic signal detection in marine mammals, and achieves high-precision tick signal segmentation detection.

CN120279554AActive Publication Date: 2025-07-08FIRST INSTITUTE OF OCEANOGRAPHY MNR
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510712473.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-08
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the marine environment, the acoustic signal detection of marine mammals is easily disturbed by environmental noise, resulting in the problems of misdetection and missed detection, especially the poor effect of local small soundprint detection.

Method used

Bayesian convolution and wavelet transform convolution are introduced into the encoder-decoder architecture to construct a semantic segmentation network model of acoustic images, process the original acoustic signals of marine mammals, and multi-scale feature extraction and fusion are achieved through Bayesian convolution pyramid and wavelet decomposition convolution to improve detection accuracy.

Benefits of technology

It effectively enhances the model's ability to model uncertainty, can accurately characterize the local details of the sound signals of marine mammals, avoid missed detection, and improves the accuracy of the tick signal segmentation detection in marine mammals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279554A_ABST
    Figure CN120279554A_ABST
Patent Text Reader

Abstract

The invention discloses a marine mammal tick signal semantic segmentation detection method and device and a medium, and relates to the field of speech recognition, and the method comprises the steps: obtaining an original sound signal of a marine mammal; processing the original sound signal to obtain a time-frequency diagram; bayesian convolution and wavelet transform convolution are introduced into an encoder-decoder architecture, and an acoustic image semantic segmentation network model is constructed; and inputting the time-frequency diagram into the acoustic image semantic segmentation network model to obtain a semantic segmentation detection result of the marine mammal tick signal. According to the method, the local fine voiceprint contained in the marine mammal sound signal can be detected, the problems of missing detection and the like are avoided, and then the accuracy of segmentation detection of the marine mammal tick signal can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition, and particularly to a method, device, and medium for semantic segmentation detection of click signals of marine mammals. Background Art

[0002] Marine mammals are a key component of the marine ecosystem and play an irreplaceable role in maintaining ecological balance and protecting biodiversity. The vocal signals of marine mammals are mainly divided into three categories: Click, Burst-pulse, and Whistle. Due to the overlap of environmental noise and the acoustic signals of marine mammals in terms of frequency and time, problems such as misjudgment or missed detection are likely to occur during the signal extraction and recognition of the acoustic signals of marine mammals. Therefore, the detection of the acoustic signals of marine mammals faces extremely high requirements.

[0003] To address the challenges brought by marine noise interference to signal detection, the inventors have proposed various feature extraction and classification methods for the acoustic signals of marine mammals in the known technology. Traditional methods can effectively detect signals under high signal-to-noise ratios, but are prone to false detection when there is interference. Although methods using machine learning and the like are currently used for the detection of the acoustic signals of marine mammals, these methods have problems such as missed detection for local and fine acoustic patterns. Summary of the Invention

[0004] To solve the above problems existing in the prior art, the present application provides a method, device, and medium for semantic segmentation detection of click signals of marine mammals.

[0005] To achieve the above object, the present application provides the following solutions: In a first aspect, the present application provides a method for semantic segmentation detection of click signals of marine mammals, including: Obtaining an original acoustic signal of a marine mammal; Processing the original acoustic signal to obtain a time-frequency diagram; Introducing Bayesian convolution and wavelet transform convolution into an encoder-decoder architecture to construct an acoustic image semantic segmentation network model; Inputting the time-frequency diagram into the acoustic image semantic segmentation network model to obtain a semantic segmentation detection result of the click signal of the marine mammal.

[0006] Optionally, obtaining an original acoustic signal of a marine mammal includes: Collecting the original acoustic signal of the marine mammal using a hydrophone at a set sampling rate.

[0007] Optionally, processing the original acoustic signal to obtain a time-frequency diagram includes: Filter the original acoustic signal to obtain a filtered signal; Obtain audio segments based on the filtered signal; Perform time-frequency conversion on the audio segments to obtain the time-frequency diagram.

[0008] Optionally, filtering the original acoustic signal to obtain a filtered signal includes: Obtain the cut-off frequency and sampling rate of the original acoustic signal; Determine the normalized cut-off frequency according to the cut-off frequency and the sampling rate; Based on the normalized cut-off frequency, use the inverse discrete-time Fourier transform to obtain the ideal impulse response; Use a Hamming window to smoothly truncate the ideal impulse response to obtain filter coefficients; Use convolution operation to apply the filter coefficients to the original acoustic signal to obtain the filtered signal.

[0009] Optionally, performing time-frequency conversion on the audio segments to obtain the time-frequency diagram includes: Use a Hamming window to window each audio segment; Perform a fast Fourier transform on the filtered signal in each windowed audio segment to obtain a frequency-domain signal to generate the time-frequency diagram.

[0010] Optionally, introduce Bayesian convolution and wavelet transform convolution in the encoder-decoder architecture to construct an acoustic image semantic segmentation network model, including: In the encoder part of the encoder-decoder architecture, construct a Bayesian convolution pyramid and a multi-scale mapping module to obtain an initial encoder; In the decoder part of the encoder-decoder architecture, construct a wavelet decomposition convolution, and combine the inverse wavelet transform and the alignment filtering module to achieve multi-scale feature alignment and fusion to obtain an initial decoder; Obtain a MobileNetV2 model, and combine the initial encoder and the initial decoder to form an initial semantic segmentation network model; the MobileNetV2 model is used to extract local and global features of the input image; Obtain a sample data set, and divide the sample data set into a training set and a test set according to a set ratio; Use the training set to train the initial semantic segmentation network model, and use the test set to test the trained initial semantic segmentation network model during the training process until the trained initial semantic segmentation network model meets the set requirements to obtain the trained initial semantic segmentation network model; Use the trained initial semantic segmentation network model as the acoustic image semantic segmentation network model.

[0011] Optionally, the Bayesian convolutional pyramid consists of multi-scale Bayesian convolutional layers with different convolutional kernel sizes.

[0012] Optionally, the process of inputting the time-frequency map into the acoustic image semantic segmentation network model to obtain the semantic segmentation detection result of the marine mammal click signal includes: Input the time-frequency map into the MobileNetV2 model in the acoustic image semantic segmentation network model to obtain global features and local features; Input the local features into the encoder in the acoustic image semantic segmentation network model. The Bayesian convolutional pyramid in the encoder performs multi-scale feature extraction on the local features to obtain first multi-scale features; after the multi-scale mapping module in the encoder performs feature mapping on the first multi-scale features, a dimensionality reduction process is performed using an image pooling layer to obtain encoded features; Input the global features into the decoder in the acoustic image semantic segmentation network model. The wavelet decomposition convolution in the decoder decomposes the global features to obtain second multi-scale features; the second multi-scale features are processed using Bayesian convolution, and an inverse wavelet transform is performed on the processed second multi-scale features to obtain reconstructed features; in the decoder, the reconstructed features and the global features processed by Bayesian convolution are fused to obtain first fusion features; Input the first fusion features and the encoded features into the alignment filtering module in the decoder. After performing multi-scale fusion and alignment filtering processing, second fusion features are obtained; after sequentially performing Bayesian convolution processing and segmentation processing on the second fusion features, output features are obtained; Based on the output features, obtain the semantic segmentation detection result of the marine mammal click signal.

[0013] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the steps of the above-provided method for semantic segmentation detection of marine mammal click signals.

[0014] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-provided method for semantic segmentation detection of marine mammal click signals are implemented.

[0015] According to the specific embodiments provided by the present application, the present application has the following technical effects: The method, device, and medium for semantic segmentation detection of marine mammal click signals provided by this application can effectively enhance the uncertainty modeling ability of the constructed model by introducing Bayesian convolution into the encoder-decoder architecture, and can fully exploit the context information in marine mammal acoustic signals. By introducing wavelet transform convolution into the encoder-decoder architecture, the local details of marine mammal acoustic signals can be accurately characterized, enabling the detection of local fine sound patterns in marine mammal acoustic signals and avoiding problems such as missed detections. Based on this, by using an acoustic image semantic segmentation network model for semantic segmentation detection of marine mammal click signals, the accuracy of marine mammal click signal segmentation detection can be further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of the method for semantic segmentation detection of marine mammal click signals provided by an embodiment of this application; Figure 2 It is a schematic diagram of the implementation processing flow of the method for semantic segmentation detection of marine mammal click signals provided by an embodiment of this application; Figure 3 It is a schematic diagram of an example of the result of manual marking provided by an embodiment of this application; Figure 4 It is a schematic diagram of the structure of the acoustic image semantic segmentation network model provided by an embodiment of this application; Figure 5 It is a schematic flowchart of the wavelet decomposition convolution provided by an embodiment of this application; Figure 6 It is a schematic diagram of the structure of a computer device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0019] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] In an exemplary embodiment, the present application provides a method for semantic segmentation detection of marine mammal click signals. This method is executed by a computer device, which can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, this method is described by taking its application to a server as an example. As Figure 1 shown, this method includes: Step 100: Obtain the original sound signal of the marine mammal.

[0021] Step 101: Process the original sound signal to obtain a time-frequency diagram.

[0022] Step 102: Introduce Bayesian convolution and wavelet transform convolution into the encoder-decoder architecture to construct an acoustic image semantic segmentation network model.

[0023] Step 103: Input the time-frequency diagram into the acoustic image semantic segmentation network model to obtain the semantic segmentation detection result of the marine mammal click signal.

[0024] By implementing the above steps 100 - 103, the present application can detect the local and fine voiceprints contained in the marine mammal sound signal, avoid problems such as missed detection, and thus can significantly improve the accuracy of the segmentation detection of the marine mammal click signal.

[0025] In another exemplary embodiment of the present application, in order to achieve the comprehensive acquisition of the original sound signal of the marine mammal, in this embodiment, the above step 100 may be: using a hydrophone to collect the original sound signal of the marine mammal at a set sampling rate.

[0026] Among them, the hydrophone used to collect the marine mammal sound signal is generally deployed in seawater in the form of a mooring buoy and operates in a self-contained manner. Based on this, in the actual application process, the response frequency of the hydrophone used can be 20 Hz - 150 kHz, the sensitivity is -190 dB, and the sampling rate is 400 kHz. The hydrophone works continuously to achieve continuous collection of the sound signal to obtain the original sound signal of the marine mammal.

[0027] In another exemplary embodiment of the present application, in order to provide a data basis for processing for the acoustic image semantic segmentation network model, in this embodiment, the implementation process of the above step 101 can be replaced by the following steps 200 - 202.

[0028] Step 200: Filter the original sound signal to obtain a filtered signal.

[0029] In the actual application process, taking the implementation of high-pass filtering as an example, the implementation manner of the above step 200 may include: (1) Obtain the cut-off frequency and sampling rate of the original acoustic signal, and calculate the normalized cut-off frequency according to the cut-off frequency and sampling rate. Among them, the normalized cut-off frequency is expressed as: .

[0030] In the formula, is the normalized cut-off frequency, is the cut-off frequency, is the sampling rate.

[0031] (2) Based on the normalized cut-off frequency, use the inverse discrete-time Fourier transform to obtain the ideal impulse response of the ideal high-pass filter (which can also be called the infinite impulse response). Among them, the ideal impulse response is expressed as: .

[0032] In the formula, is the ideal impulse response, is the discrete-time index, is the order of the filter, is the discrete-time unit impulse function.

[0033] (3) Use the Hamming window to smoothly truncate the ideal impulse response to obtain the filter coefficients (which can also be called the filter coefficients). Among them, the filter coefficients are expressed as: .

[0034] In the formula, is the filter coefficient, is the Hamming window function.

[0035] (4) Use the convolution operation to apply the filter coefficients to the original acoustic signal to obtain the filtered signal. Among them, the filtered signal is expressed as: .

[0036] In the formula, is the filtered signal, is the value of the impulse response at position k, is the value of the discrete signal at the summation variable position, is the length of the impulse response.

[0037] Step 201: Obtain the audio segment based on the filtered signal.

[0038] Step 202: Perform time-frequency conversion on the audio segment to obtain a time-frequency diagram. Specifically, each audio segment is windowed using a Hamming window. Then, the windowed audio segment is subjected to a fast Fourier transform to convert it from the time domain to the frequency domain, thereby converting the acoustic signal into a time-frequency diagram.

[0039] In another exemplary embodiment of the present application, a high-quality dataset can be constructed based on a tick detection method with multi-parameter constraints to support the superior performance of the acoustic image semantic segmentation network (BayesWaveNet) model in a complex noise environment.

[0040] In the encoder-decoder architecture, the encoder enhances the uncertainty modeling ability through smooth Bayesian convolution, and the decoder accurately extracts local details through wavelet decomposition convolution. A multi-scale feature fusion is achieved by combining an alignment filtering module. Based on this, in this embodiment, the specific construction process of the acoustic image semantic segmentation network model in step 102 is described in conjunction with the Figure 2 process shown as follows, including: Step 1: In the encoder part of the encoder-decoder architecture, construct a Bayesian convolution pyramid and a multi-scale mapping module to obtain an initial encoder. The Bayesian convolution pyramid consists of multi-scale Bayesian convolutional layers with different convolutional kernel sizes, which are used to capture multi-scale features of the image. Bayesian convolution uses a probability distribution to describe the weight parameters.

[0041] For example, as Figure 4 shown, the Bayesian convolution pyramid includes Bayesian convolutional layers (BayesConv) with four convolutional kernel sizes of 1×1, 3×3, 5×5, and 7×7 to capture details and patterns of the image at different scales and achieve comprehensive extraction of image features.

[0042] Step 2: In the decoder part of the encoder-decoder architecture, construct a wavelet decomposition convolution, and combine the inverse wavelet transform and the alignment filtering module to achieve multi-scale feature alignment and fusion to obtain an initial decoder.

[0043] Based on Step 1 and Step 2, in the encoder part, a feature extraction network based on a Bayesian convolution pyramid and a multi-scale mapping module is constructed. Among them, a smooth Bayesian convolution is proposed in the encoder part and combined with the multi-scale mapping module. In the decoder part, a wavelet decomposition convolution (WDConv) is proposed, combined with the inverse wavelet transform (IWT) and the alignment filtering module (MAF) to achieve multi-scale feature alignment and fusion.

[0044] Step 3: Obtain a MobileNetV2 model, and combine the initial encoder and the initial decoder to form an initial semantic segmentation network model. The MobileNetV2 model is used to extract local and global features of the input image.

[0045] Step 4: Obtain the sample data set and divide the sample data set into a training set and a test set according to a set ratio. In this step, the acoustic signals of marine mammals collected can be subjected to high-pass filtering and time-frequency conversion to obtain a time-frequency sample map, and then a high-quality data set WhaleDataSet (i.e., the sample data set) can be constructed through multi-parameter screening and data annotation.

[0046] Among them, the method of multi-parameter screening and data annotation can be: determine whether there is a click signal of marine mammals in the time-frequency sample map according to the anti-interference real-time click sound detection method based on multi-parameter constraints. Based on this, all the time-frequency sample maps in the data set WhaleDataSet are manually annotated, and the obtained annotation results are as Figure 3 shown. Figure 3 In, part (a) is a schematic diagram of the manual annotation result of the finless porpoise in the Yangtze River, part (b) is a schematic diagram of the manual annotation result of the Chinese white dolphin, and part (c) is a schematic diagram of the manual annotation result of the noise. Figure 3 In the abscissa is the sampling time; the ordinate is the click signal of marine mammals, and the unit is decibel (dB).

[0047] Step 5: Train the initial semantic segmentation network model with the training set, and use the test set to test the trained initial semantic segmentation network model during the training process until the trained initial semantic segmentation network model meets the set requirements (for example, reaches the maximum number of iterations or the model evaluation index reaches the set value), and obtain the trained initial semantic segmentation network model.

[0048] For example, use the time-frequency sample map as the feature input of the initial semantic segmentation network model to train the initial semantic segmentation network model. The experimental environment for training and testing is implemented based on Python 3.6.5, the hardware configuration uses an NVIDIA GTX 3090 graphics card, and the CUDA 10.0 toolkit is used to accelerate the calculation. The data set WhaleDataset is randomly divided into a training set and a test set according to a ratio of 7:3. The model training is carried out for 200 epochs (i.e., the set requirements), and the initial learning rate is set to 1×10 -4 . The learning rate adopts a fixed decay strategy, remains unchanged in the first 100 epochs, and then decays by 0.01 every 10 epochs. The optimizer selects the stochastic gradient descent method (SGD). The finally obtained BayesWaveNet model achieves an average intersection over union of 91.16% and an average precision of 95.27% in a complex noise environment, which is significantly better than the existing methods.

[0049] Step 6: Use the trained initial semantic segmentation network model as the acoustic image semantic segmentation network model. Among them, the obtained acoustic image semantic segmentation network model is as Figure 4as shown Figure 4 In this case, the input global features are subjected to wavelet transform, that is, the global features are decomposed into a low-frequency sub-band (L) and high-frequency sub-bands (H, V, D) through a set of specific filters. Among them, the low-frequency sub-band (L) is obtained through a low-pass filter, while the high-frequency sub-bands (H, V, D) correspond to high-frequency information in the horizontal, vertical, and diagonal directions respectively, and are obtained through high-pass filters f H 、f V and f D obtained

[0050] In another exemplary embodiment of the present application, in order to accurately identify the click signals of marine mammals using the finally obtained acoustic image semantic segmentation network model, the original acoustic signals outside the data set can be processed according to the above steps 200-step 202 to obtain a time-frequency diagram, which is input into the acoustic image semantic segmentation network model, and the acoustic image semantic segmentation network model gives the corresponding recognition result, realizing the prediction of the original acoustic signal (i.e., new data). Based on this, the implementation process of the above step 103 may include: Step 1), input the time-frequency diagram into the MobileNetV2 model in the acoustic image semantic segmentation network model to obtain global features and local features. This step mainly uses the MobileNetV2 model to extract global and local features of the input time-frequency diagram.

[0051] Step 2), input the local features into the encoder in the acoustic image semantic segmentation network model. The Bayesian convolutional pyramid in the encoder performs multi-scale feature extraction on the local features to obtain the first multi-scale features. After the multi-scale mapping module in the encoder performs feature mapping on the first multi-scale features, a dimensionality reduction process is performed using an image pooling layer to obtain encoded features for subsequent global feature extraction.

[0052] Step 3), input the global features into the decoder in the acoustic image semantic segmentation network model. The wavelet decomposition convolution in the decoder decomposes the global features to obtain the second multi-scale features. The second multi-scale features are processed using Bayesian convolution, and an inverse wavelet transform is performed on the processed second multi-scale features to obtain reconstructed features. In the decoder, the reconstructed features and the global features processed by Bayesian convolution are fused to obtain the first fused features.

[0053] Step 4), input the first fused features and the encoded features into the alignment filtering module of the decoder. After performing multi-scale fusion and alignment filtering processing, the second fused features are obtained. After sequentially performing Bayesian convolution processing and segmentation processing on the second fused features, output features are obtained.

[0054] Step 5), based on the output features, obtain the semantic segmentation detection result of the click signals of marine mammals.

[0055] Based on the above description, in steps 3) - 5), the data processing process of the acoustic image semantic segmentation network model can be described as follows: First, the MobileNetV2 model is used to extract global and local features from the input time-frequency map. Then, the global features are decomposed into multi-scale features through wavelet transform, and then Bayesian convolution is applied to process the features, and the signal is reconstructed through inverse wavelet transform (for example, the reconstructed signal is as Figure 5 shown). The wavelet decomposition convolution performs wavelet decomposition on the global features of the MobileNetV2 model, decomposes these feature maps into components of different frequencies, and performs Bayesian convolution operations on the sub-bands, obtaining a large receptive field and multi-frequency responses, so as to further extract the local features and texture information of the time-frequency map. Subsequently, the global features and the processed local features are jointly input into the alignment filter (MAF) in the alignment filtering module. The alignment filter precisely aligns and deeply fuses features of different scales to ensure information consistency between multi-scale features. Finally, the refined detection is performed on the fused feature map to generate high-precision detection and classification results of marine mammal acoustic signals.

[0056] In summary, the method provided in this application constructs an acoustic image semantic segmentation network (BayesWaveNet) model based on Bayesian convolution and wavelet transform for the detection and classification of marine mammal acoustic signals, which can provide an efficient and reliable technical means for marine exploration and protection.

[0057] The BayesWaveNet model adopts an encoder-decoder architecture. In the encoder part, a smooth Bayesian convolution is proposed to effectively enhance the uncertainty modeling ability of the BayesWaveNet model, and it is combined with the multi-scale mapping module to fully exploit the context information of the global features. In the decoder part, this application adopts wavelet decomposition convolution, which can accurately depict and extract local details, and realizes the precise alignment and fusion of multi-scale features by introducing the alignment filtering module, supporting the BayesWaveNet model to achieve high-precision detection and classification of marine mammal acoustic signals in a complex noise environment.

[0058] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data for semantic segmentation detection of marine mammal click signals. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for semantic segmentation detection of marine mammal click signals.

[0059] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0060] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.

[0061] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0062] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0063] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0064] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memory (RRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0065] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0066] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0067] In this article, specific examples are used to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. To sum up, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A method for semantic segmentation detection of the click signal of marine mammals, characterized in that, Including: Obtaining the original acoustic signal of marine mammals; Processing the original acoustic signal to obtain a time-frequency diagram; Introducing Bayesian convolution and wavelet transform convolution into the encoder-decoder architecture to construct an acoustic image semantic segmentation network model; Inputting the time-frequency diagram into the acoustic image semantic segmentation network model to obtain the semantic segmentation detection result of the click signal of marine mammals.

2. The method for semantic segmentation detection of the click signal of marine mammals according to claim 1, wherein Obtaining the original acoustic signal of marine mammals, including: Collecting the original acoustic signal of marine mammals using a hydrophone at a set sampling rate.

3. The method for semantic segmentation detection of the click signal of marine mammals according to claim 1, wherein Processing the original acoustic signal to obtain a time-frequency diagram, including: Performing filtering processing on the original acoustic signal to obtain a filtered signal; Obtaining audio segments based on the filtered signal; Performing time-frequency conversion on the audio segments to obtain the time-frequency diagram.

4. The method for semantic segmentation detection of the click signal of marine mammals according to claim 3, characterized in that, Performing filtering processing on the original acoustic signal to obtain a filtered signal, including: Obtaining the cut-off frequency and sampling rate of the original acoustic signal; Determining the normalized cut-off frequency according to the cut-off frequency and the sampling rate; Based on the normalized cut-off frequency, obtaining the ideal impulse response using the inverse discrete-time Fourier transform; Smoothing and truncating the ideal impulse response using a Hamming window to obtain a filter coefficient; Performing convolution operation and applying the filter coefficient to the original acoustic signal to obtain the filtered signal.

5. The method for semantic segmentation detection of the click signal of marine mammals according to claim 3, wherein Performing time-frequency conversion on the audio segments to obtain the time-frequency diagram, including: Using a Hamming window to window each audio segment; Performing fast Fourier transform on the filtered signal in each windowed audio segment to obtain a frequency-domain signal to generate the time-frequency diagram.

6. The method for semantic segmentation detection of the click signal of marine mammals according to claim 1, wherein Introducing Bayesian convolution and wavelet transform convolution into the encoder-decoder architecture to construct an acoustic image semantic segmentation network model, including: In the encoder part of the encoder-decoder architecture, constructing a Bayesian convolution pyramid and a multi-scale mapping module to obtain an initial encoder; In the decoder part of the encoder-decoder architecture, constructing a wavelet decomposition convolution, and combining the inverse wavelet transform and an alignment filtering module to achieve multi-scale feature alignment and fusion to obtain an initial decoder; Obtaining a MobileNetV2 model, and combining the initial encoder and the initial decoder to form an initial semantic segmentation network model; the MobileNetV2 model is used to extract local features and global features of the input image; Obtaining a sample data set and dividing the sample data set into a training set and a test set according to a set ratio; Training the initial semantic segmentation network model using the training set, and testing the trained initial semantic segmentation network model using the test set during the training process until the trained initial semantic segmentation network model meets the set requirements to obtain the trained initial semantic segmentation network model; Taking the trained initial semantic segmentation network model as the acoustic image semantic segmentation network model.

7. The method for semantic segmentation detection of marine mammal click signals according to claim 6, wherein The Bayesian convolution pyramid consists of multi-scale Bayesian convolution layers with different convolution kernel sizes.

8. The method for semantic segmentation detection of marine mammal click signals according to claim 6, characterized in that The process of inputting the time-frequency diagram into the acoustic image semantic segmentation network model to obtain the semantic segmentation detection result of the click signal of marine mammals includes: Input the time-frequency map into the MobileNetV2 model in the acoustic image semantic segmentation network model to obtain global features and local features; Input the local features into the encoder in the acoustic image semantic segmentation network model. The Bayesian convolutional pyramid in the encoder performs multi-scale feature extraction on the local features to obtain first multi-scale features; after the multi-scale mapping module in the encoder performs feature mapping on the first multi-scale features, a dimensionality reduction process is performed using an image pooling layer to obtain encoded features; Input the global features into the decoder in the acoustic image semantic segmentation network model. The wavelet decomposition convolution in the decoder decomposes the global features to obtain second multi-scale features; the second multi-scale features are processed using Bayesian convolution, and an inverse wavelet transform is performed on the processed second multi-scale features to obtain reconstructed features; in the decoder, the reconstructed features and the global features processed by Bayesian convolution are fused to obtain first fusion features; Input the first fusion features and the encoded features into the alignment filtering module of the decoder. After multi-scale fusion and alignment filtering processing, second fusion features are obtained; after the second fusion features are sequentially processed by Bayesian convolution and segmentation processing, output features are obtained; Based on the output features, obtain the semantic segmentation detection result of the marine mammal click signal.

9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the marine mammal click signal semantic segmentation detection method according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the marine mammal click signal semantic segmentation detection method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Image semantic segmentation method and device based on codec

    CN111292330A

  • Detection signal correction method, device and equipment for sensor on autonomous vehicle

    CN116164784A

  • SAR image radio frequency interference suppression method based on semantic segmentation

    CN117192490A

  • Semantic segmentation method and device for ocean remote sensing image and electronic equipment

    CN117522884A

  • Radar target detection method based on camera supervision feature enhancement

    CN119064925A