An Automatic Synthesis Method for Cetacean Calls Based on Time-Frequency Spectrum
Through the automatic synthesis method based on time spectrum, the PFMB signal model and sub-signal model are used to identify and segmentally imitate the time spectrum profile of cetaceans' calls, which solves the problem of cetacean sound that is difficult to synthesize complex time frequency structures in the prior art, and achieves high similarity bionic signal synthesis.
Patent Information
- Application Number
- CN202111193685.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-13
AI Technical Summary
It is difficult to efficiently synthesize cetacean call bionic signals with complex time-frequency structures in the prior art, especially when quickly identifying and imitating cetacean sounds with nonlinear frequency modulation characteristics, the degree of freedom is low and it is difficult to achieve high similarity imitation.
The automatic synthesis method of cetacean call based on time spectrum is used to construct tonal sound through four sub-signal models in the PFMB signal model, and the time spectrum profile of the fundamental frequency and harmonics is identified and extracted, the sound is segmented and each sound is imitated using bionic signal segments, and finally the synthesized bionic signal is superimposed on the time domain.
It realizes automatic synthesis of high similarity to cetacean calls, can quickly identify and imitate cetacean sounds with complex time-frequency structures, and improves the flexibility and adaptability of bionic signals.
Smart Images

Figure CN114023298B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of underwater acoustic signal processing, and particularly relates to a method for automatically synthesizing the calls of cetaceans. Background Art
[0002] Tonal sound is an important component of the sounds of cetaceans. This sound is produced by toothed whales and baleen whales, which are sister branches that include all extant whales. Among cetaceans, the acoustic characteristics of tonal sound vary greatly. Therefore, the tonal sound of cetaceans is broadly defined as a frequency-modulated signal, and this sound is usually classified according to their time-frequency spectral profiles. Due to the wide distribution of tonal sound and their diverse acoustic characteristics in terms of duration, frequency distribution, and time-frequency distribution, tonal sound has attracted the attention of many researchers in the past few decades.
[0003] In biological research, tonal sound is considered as a carrier for cetaceans to transmit information, and is closely related to their behaviors such as foraging, courtship, and playing. The collection of tonal sound is generally achieved through passive acoustic monitoring technology. In order to accurately identify and classify sounds, a large amount of sound data is usually required to train acoustic detectors and classifiers, such as neural networks and support vector machines. In addition, during the research process of environmental monitoring and environmental protection, biologists often use tonal sound and bionic signals to study the effects of different sounds on the behaviors of cetaceans; in non-biological research, tonal sound also plays an important role. For example, underwater bionic acoustic detection and communication, which have received increasing attention in recent years, often use tonal sound for underwater target detection or as a carrier of communication information.
[0004] The call characteristics of cetaceans in nature are rich and diverse. They will emit calls during the processes of predation and communication. Different species and different ages of cetaceans will emit different calls at different times and locations. However, due to the vastness of the ocean, the whereabouts of cetaceans are usually difficult to determine, and researchers usually need to spend months or even years to search for and track cetaceans. Therefore, compared with the rich and diverse calls of cetaceans in nature, the call data obtained by researchers is usually scarce, and the calls of some categories may not be in the database.
[0005] The bionic signal waveform design based on existing tonal sound data is an important method to solve the above problems. The synthesized bionic signal can be used to expand the tonal sound database and evaluate the effects of different acoustic characteristics (duration, spectral distribution, time-frequency distribution, etc.) on cetacean behavior. In addition, good detection and communication performance can be obtained by modifying the parameters of the bionic signal. Traditional bionic signal waveform design methods can be mainly divided into two categories. The first category synthesizes bionic signals through the weighted superposition technology of sine signals. However, this method can only extract the contour of tonal sound and it is difficult to modify the time-frequency spectrum contour of the synthesized bionic signal. The second category of methods is based on traditional frequency modulation signal models, such as linear frequency modulation signals and double-tone frequency modulation signal models. However, due to the complex and diverse time-frequency structures of tonal sound, the traditional frequency modulation signal models have low degrees of freedom and can only synthesize cetacean sounds with simple time-frequency structures, making it difficult to achieve a high degree of similarity imitation of cetacean sounds with complex non-linear frequency modulation characteristics.
[0006] In the patent document with the publication number CN111431625A, a method for synthesizing and modifying the calls of cetaceans is disclosed. This method uses two bionic signal models, namely the Power Frequency Modulation Bionic (PFMB) signal model and the Sinusoidal Frequency Modulation Bionic (SFMB) signal model, to segmentally imitate the tonal sound of cetaceans. However, this method requires manual identification of whether there are harmonics in the tonal sound, splitting the fundamental frequency and harmonics into multiple segments, and adjusting the parameters of each bionic signal segment to match the time-frequency spectrum contour of the original tonal sound, resulting in a very time-consuming synthesis process. In addition, this method requires two models (including eight sub-bionic signal models) to synthesize tonal sounds, leading to a very high logical complexity in model selection. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for automatically synthesizing the calls of cetaceans, which can automatically synthesize bionic signals with a high degree of similarity to the original calls of cetaceans.
[0008] The purpose of the present invention is achieved through the following technical solutions:
[0009] An automatic synthesis method for cetacean calls based on the time-frequency spectrum constructs tonal sound through four sub-signal models included in the PFMB signal model. For the bionic signal of constructing tonal sound, first, identify and extract the time-frequency spectrum profiles of the fundamental frequency and harmonics in the tonal sound. For the fundamental frequency or harmonics, according to the extreme points and inflection points of their time-frequency spectrum profiles, divide the tonal sound into M segments, where M≥1 and M is an integer. Then, use M bionic signal segments to imitate these M segments of tonal sound respectively, and splice the M bionic signal segments end to end to synthesize the fundamental frequency or harmonics. Finally, superimpose the fundamental frequency and all harmonics in the time domain to synthesize the bionic signal of the tonal sound. Specifically, it includes the following steps:
[0010] (1) Initially extract the time-frequency spectrum profile of the fundamental frequency of the tonal sound;
[0011] (2) Filter out the outliers generated by ocean noise during the process of step (1);
[0012] (3) When the outliers are filtered out, the time-frequency points at the positions of the previous outliers in the filtered time-frequency spectrum profile of the fundamental frequency are vacant. Fill in the time-frequency points at the vacant positions to obtain the final time-frequency spectrum profile of the fundamental frequency;
[0013] (4) Identify the harmonics and construct the time-frequency spectrum profiles of the harmonics;
[0014] (5) Construct the time-frequency expressions of each bionic signal segment in the bionic signal;
[0015] (6) Construct the segmented bionic signal;
[0016] (7) Extract the envelope of the tonal sound;
[0017] (8) Synthesize the bionic signal.
[0018] Further, step (1) specifically includes the following steps:
[0019] Let ST(t) represent the time-continuous signal of the tonal sound, and the discrete signal generated after sampling is ST[n]; the time-frequency spectrum of ST[n] is represented as X[l, h], and X[l, h] is obtained by the short-time Fourier transform, where l represents the data block number, 1≤l≤L and l is an integer, and h represents the frequency index, 1≤h≤H and h is an integer; the spectrum of the l-th data block is represented as X l [h], and X l [h] is obtained by the discrete Fourier transform; the peak value of X l [h] is P l and is obtained by the following formula
[0020] [P l,F l = max(X l [h]) (12)
[0021] where F l is the frequency index corresponding to P l , that is, X[l, F l = P l ; In the tonal sound, the energy of the fundamental frequency is the strongest, and the energy of the harmonics gradually decreases as the number of times increases; Therefore, the spectral contour of the fundamental frequency is expressed as f1[l], which is initially obtained from expression (13);
[0022] f1[l]= argmax(X l [h]) (13)
[0023] When extracting the spectral contour of the fundamental frequency, it is necessary to determine whether each time-frequency point extracted is a point on the second harmonic; Assume that at l = l x , the peak value of X lx [h] is P lx , and its corresponding frequency index is F lx . If the condition in expression (14) is satisfied, then at l = l x , correct f1[l x to the frequency corresponding to the frequency index F lx / 2;
[0024]
[0025] where λ is the fundamental frequency energy threshold;
[0026] Furthermore, step (2) specifically includes the following steps:
[0027] (201) Aggregate the time-frequency points in the corrected spectral contour f COR [l] of the fundamental frequency into several connected domains; Select a time-frequency point f COR [l] in f COR [l o as the seed point, and the first time-frequency point f COR [1] as the first seed, and f COR [1] belongs to the first connected domain CC1; If the time-frequency point f COR [l o adjacent to the seed point f COR [l1] satisfies the condition in expression (15), then add f COR [l1] to the connected domain to which f COR [l o belongs;
[0028] fCOR [l1] - f COR [l0] | < T CC1 (15)
[0029] Where T CC1 is the time - frequency point frequency threshold; then, take the time - frequency point f COR [l1] as a seed point. If f COR [l1] does not satisfy the condition in expression (15), also take it as a seed point; then repeat the process of this step until no time - frequency point can be used as a seed point; finally, the time - frequency points in f COR [l] will be aggregated into several connected domains {CC1, CC2,...};
[0030] (202) Filter out the connected domain where the wild point is located; Take the i - th connected domain CC i as the seed connected domain, and compare the remaining connected domains with CC i respectively; If the frequency difference between CC i and the j - th connected domain satisfies the condition in expression (16), then splice CC i and CC j into one connected domain; i ≥ 1 and i is an integer, j > i and j is a positive integer;
[0031] |CC j [1] - CC i [end]| < T CC2 (16)
[0032] Where CC j [1] is the first time - frequency point of CC j , CC i [end] is the last time - frequency point of CC i , and T CC2 is the connected domain frequency threshold; then take CC j as the seed connected domain. If CC j does not satisfy the condition in expression (16), also take it as the seed connected domain;
[0033] (203) Repeat the process of the previous step until no connected domain can be used as the seed connected domain; finally, {CC1, CC2,...} are spliced into several larger connected domains, and the connected domain containing the most time - frequency points forms the filtered fundamental frequency time - frequency spectrum profile, which is denoted as f FILT [l].
[0034] Furthermore, step (3) specifically includes the following steps:
[0035] (301) For the filtered fundamental frequency time - frequency spectrum profile f FILT[l]Perform smoothing filtering to obtain the smoothed fundamental frequency time-frequency spectrum profile f BS1 [l];
[0036] (302) In the time-frequency spectrum of the tonal sound, with f BS1 [l] as the center, within the frequency range of ±Δf1 near f BS1 [l], the time-frequency point with the maximum energy is considered to be the missing time-frequency point, that is, the data block number l and frequency index h corresponding to the missing time-frequency point satisfy the conditions in expression (17).
[0037]
[0038] When all the missing time-frequency points are filled, the fundamental frequency time-frequency spectrum profile can be obtained.
[0039] Furthermore, step (4) specifically includes the following steps:
[0040] (401) Since the fundamental frequency time-frequency spectrum profile has been extracted, the maximum possible harmonic number in the tonal sound is:
[0041] R max = f S / 2 / (minf1[l]) (18)
[0042] where f S is the sampling frequency of the tonal sound, the maximum possible harmonic number R max is an integer and R max ≥ 1. If R max = 1, it means that there are no second harmonics and higher harmonics in the tonal sound;
[0043] (402) In the spectrum X l [h], the spectral peaks within the frequency range [rf1[l] - Δf2, rf1[l] + Δf2] are considered to be the candidate spectral peaks of the r-th harmonic; detect all the candidate spectral peaks of the r-th harmonic. If a candidate spectral peak is larger than its left and right adjacent candidate spectral peaks by T P (in dB), then select it for the next comparison; among all the selected candidate spectral peaks, the candidate spectral peak with the maximum peak value is the effective spectral peak;
[0044] (403) If the number L r of the effective spectral peaks of the r-th harmonic satisfies the conditions in expression (19), then it is determined that the r-th harmonic exists;
[0045] L r > T L ·L (19)
[0046] Where T L (0 < T L <1) is the length threshold;
[0047] (404) Sequentially detect whether there is a harmonic from the second harmonic to the R max th harmonic in the tonal sound, and the number of harmonics present in the tonal sound can be obtained;
[0048] (405) Since the frequency of the harmonic is an integer multiple of the frequency of the fundamental frequency, the spectral profile f r [l] at the rth harmonic is expressed as
[0049] f r [l] = r·f1[l] (20).
[0050] Further, step (5) specifically includes the following steps:
[0051] (501) Perform smoothing filtering on the spectral profile f1[l] at the fundamental frequency to obtain the smoothed spectral profile f BS2 [l];
[0052] (502) The first derivative and the second derivative of f BS2 [l] are respectively expressed as f BS2 ′[l] and f BS2 ″[l]; The extreme points and inflection points of f BS2 [l] respectively satisfy the conditions f BS2 ′[l] = 0 and f BS2 ″[l] = 0;
[0053] (503) Divide the spectral profile of the harmonic into several spectral profile segments in the same way as above;
[0054] (504) For each spectral profile segment, select one from the four sub-signal models in the PFMB signal model to construct a bionic signal segment; for the spectral profiles of each spectral profile segment and the corresponding bionic signal segment, the values of the three parameters of the bandwidth B, the duration T, and the carrier frequency f C are the same;
[0055] (505) In the spectral profile of the bionic signal segment, the two parameters of the curvature adjustment factor α and the slope adjustment factor k should satisfy the conditions in expression (21)
[0056] argmin∑(f(α,k) - f r,m [l]) 2 (21)
[0057] where f(α,k) is the spectral profile of the bionic signal segment.
[0058] Further, step (6) specifically includes the following steps:
[0059] (601) The m-th bionic signal segment of the r-th harmonic is SB r,m (t) is expressed as
[0060]
[0061] where f r,m (t) is the time-frequency spectrum profile of SB r,m (t);
[0062] (602) According to expression (10), all bionic signal segments are spliced end to end in the time domain to obtain the segmented bionic signal SN r (t);
[0063] (603) All segmented bionic signals SN r (t) are superimposed in the time domain to obtain the normalized bionic signal SN(t), and its expression is
[0064]
[0065] Further, step (7) specifically includes the following steps: X l [h] is the spectrum of the l-th data block, and the time corresponding to the first sampling point in the l-th data block is τ l ; within the frequency range [rf1[l] - Δf2, rf1[l] + Δf2], find the r-th harmonic spectrum peak as P r,l = max(X l [h]);
[0066] Then the envelope of the extracted r-th harmonic is
[0067] A r (τ l ) = KP r,l (24)
[0068] where K is the amplitude correction factor, and by modifying K, the envelope of SB(t) can be made close to the envelope of ST(t);
[0069] Substitute the segmented bionic signal SN r (t) and the envelope A r (t) into the expression of the bionic signal SB(t) containing the R-th harmonic to obtain the bionic signal SB(t).
[0070] Further, the time-frequency expressions f PO (t), f PX (t), f PY(t) and f PZ (t) are respectively:
[0071] f PO f(t) = f P f(t) = (B - kT)(t / T) α + kt + f C (1)
[0072] f PX f(t) = -f P f(t)+B + 2f C = (kT - B)(t / T) α -kt + B + f C (2)
[0073] f PY f(t) = f P [-(t - T)] = (B - kT)(1 - t / T) α -k(t - T)+f C (3)
[0074] f PZ f(t) = -f P [-(t - T)]+B + 2f C = (kT - B)(1 - t / T) α +k(t - T)+B + f C (4)
[0075] where 0 ≤ t ≤ T, T represents the signal duration, B is the signal bandwidth, f C is the signal carrier frequency. The role of the curvature adjustment factor α (α > 0) is to adjust the curvature of the time - frequency expression, and the role of the slope adjustment factor k (0 ≤ k ≤ B / T) is to adjust the slopes of the time - frequency expression at t = 0 and t = T.
[0076] The expressions of the four sub - PFMB signal models are:
[0077]
[0078]
[0079]
[0080]
[0081] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:
[0082] (1) The present invention proposes an automatic extraction algorithm for the time - frequency spectrum profile of cetacean calls, which can automatically identify and extract the time - frequency spectrum profiles of the fundamental frequency and harmonics, and effectively filter out wild points.
[0083] (2) The present invention proposes an automatic synthesis algorithm for the time-frequency spectrum profile of cetacean calls, which can automatically segment the extracted time-frequency spectrum profile and select appropriate biomimetic signal models and parameters to match each segment of the profile.
[0084] (3) The present invention proposes an automatic synthesis algorithm for the envelope of cetacean calls, which can automatically extract the envelopes of the fundamental frequency and harmonics of cetacean calls, and the envelope of the finally synthesized biomimetic signal has a very high similarity to the envelope of the original cetacean call.
[0085] (4) The present invention proposes a new automatic synthesis strategy for biomimetic signals, which has strong versatility. Regardless of the shape of the time-frequency spectrum profile of cetacean calls and whether the calls have harmonics, it can automatically synthesize the time-frequency spectrum profile and the corresponding biomimetic signal waveform, and can achieve a high degree of similarity imitation of various cetacean calls and some cetacean calls with complex time-frequency spectrum profiles. Description of the Drawings
[0086] Figure 1 Show the waveforms and time-frequency diagrams of tonal sounds of six types of cetaceans in the present invention.
[0087] Figure 2a Show the function graphs of f PO (t) corresponding to different curvature adjustment factors α in the present invention.
[0088] Figure 2b Show the function graphs of f PX (t) corresponding to different curvature adjustment factors α in the present invention.
[0089] Figure 2c Show the function graphs of f PY (t) corresponding to different curvature adjustment factors α in the present invention.
[0090] Figure 2d Show the function graphs of f PZ (t) corresponding to different curvature adjustment factors α in the present invention.
[0091] Figure 3a Show the waveform of the sine-type tonal sound in the present invention.
[0092] Figure 3b Show the time-frequency diagram of the sine-type tonal sound in the present invention.
[0093] Figure 4a Show the time-frequency spectrum profile of the fundamental frequency before correction in the present invention.
[0094] Figure 4b Show the time-frequency spectrum profile of the fundamental frequency after correction in the present invention.
[0095] Figure 5a Show the time-frequency spectrum profile of the filtered fundamental frequency and the missing time-frequency points in the present invention.
[0096] Figure 5b Show the time-frequency spectrum profile of the smoothed fundamental frequency in the present invention.
[0097] Figure 5c Show the time-frequency spectrum profile of the smoothed fundamental frequency and the time-frequency diagram of the sinusoidal tonal sound in the present invention.
[0098] Figure 5d Show the time-frequency spectrum profile of the fundamental frequency in the present invention.
[0099] Figure 6a Show a spectrogram of the sinusoidal tonal sound and an effective spectral peak on the second harmonic in the present invention.
[0100] Figure 6b Show all the effective spectral peaks of the fundamental frequency and the second harmonic in the present invention.
[0101] Figure 7a Show the extreme points and inflection points extracted based on the smoothed time-frequency spectrum profile in the present invention.
[0102] Figure 7b Show the segmentation of the smoothed time-frequency spectrum profile according to the extreme points and inflection points in the present invention
[0103] Figure 8a Show the waveform of the normalized bionic signal in the present invention.
[0104] Figure 8b Show the time-frequency diagram of the normalized bionic signal in the present invention.
[0105] Figure 9a Show the waveform of the sinusoidal bionic signal in the present invention.
[0106] Figure 9b Show the time-frequency diagram of the sinusoidal bionic signal in the present invention.
[0107] Figure 10a Show the waveform of the constant-frequency tonal sound in the present invention.
[0108] Figure 10b Show the waveform of the constant-frequency bionic signal in the present invention.
[0109] Figure 10c Show the time-frequency diagram of the constant-frequency tonal sound in the present invention.
[0110] Figure 10d Show the time-frequency diagram of the constant-frequency bionic signal in the present invention.
[0111] Figure 11a Show the waveform of the up - frequency - modulated tonal sound in the present invention.
[0112] Figure 11b Show the waveform of the up - frequency - modulated bionic signal in the present invention.
[0113] Figure 11c Show the time - frequency diagram of the up - frequency - modulated tonal sound in the present invention.
[0114] Figure 11d Show the time - frequency diagram of the up - frequency - modulated bionic signal in the present invention.
[0115] Figure 12a Show the waveform of the down - frequency - modulated tonal sound in the present invention.
[0116] Figure 12b Show the waveform of the down - frequency - modulated bionic signal in the present invention.
[0117] Figure 12c Show the time - frequency diagram of the down - frequency - modulated tonal sound in the present invention.
[0118] Figure 12d Show the time - frequency diagram of the down - frequency - modulated bionic signal in the present invention.
[0119] Figure 13a Show the waveform of the concave - shaped tonal sound in the present invention.
[0120] Figure 13b Show the waveform of the concave - shaped bionic signal in the present invention.
[0121] Figure 13c Show the time - frequency diagram of the concave - shaped tonal sound in the present invention.
[0122] Figure 13d Show the time - frequency diagram of the concave - shaped bionic signal in the present invention.
[0123] Figure 14a Show the waveform of the convex - shaped tonal sound in the present invention.
[0124] Figure 14b Show the waveform of the convex - shaped bionic signal in the present invention.
[0125] Figure 14c Show the time - frequency diagram of the convex - shaped tonal sound in the present invention.
[0126] Figure 14d Show the time - frequency diagram of the convex - shaped bionic signal in the present invention.
[0127] Figure 15aShows the waveform of the signature whistle in the present invention.
[0128] Figure 15b Shows the waveform of the signature whistle biomimetic signal in the present invention.
[0129] Figure 15c Shows the time-frequency diagram of the signature whistle in the present invention.
[0130] Figure 15d Shows the time-frequency diagram of the signature whistle biomimetic signal in the present invention. Detailed implementation manners
[0131] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0132] The present invention provides a method for automatically synthesizing cetacean calls, which can automatically synthesize biomimetic signals with high similarity to the original cetacean calls.
[0133] Figure 1 In the present invention, the waveforms and time-frequency diagrams of six types of tonal sounds of cetaceans are given. According to the time-frequency spectrum profile of the tonal sound, the tonal sound can be divided into six types: constant frequency type 1, upward frequency modulation type 2, downward frequency modulation type 3, concave type 4, convex type 5, and sine type 6. Among them, the frequency of the constant frequency type 1 basically does not change with time; the frequency of the upward frequency modulation type 2 increases with time; the frequency of the downward sweep frequency type 3 decreases with time; the frequency of the concave type 4 first decreases and then increases; the frequency of the convex type 5 first increases and then decreases; the frequency of the sine type 6 first increases and then decreases, and then increases again, and so on, or first decreases and then increases, and then decreases again, and so on.
[0134] In the patent document with the publication number CN111431625A, the PFMB signal model and the 4 sub-signal models it contains are disclosed.
[0135] In the present invention, the tonal sound is constructed by using the 4 sub-signal models included in the PFMB signal model. The time-frequency expressions f PO (t), f PX (t), f PY (t) and f PZ (t) are respectively:
[0136] f PO (t) = f P (t) = (B - kT)(t / T)α +kt + f C (1)
[0137] f PX f(t)= - f P f(t)+B + 2f C =(kT - B)(t / T) α - kt + B + f C (2)
[0138] f PY f(t)=f P [ - (t - T)]=(B - kT)(1 - t / T) α - k(t - T)+f C (3)
[0139] f PZ f(t)= - f P [ - (t - T)]+B + 2f C =(kT - B)(1 - t / T) α +k(t - T)+B + f C (4)
[0140] where 0 ≤ t ≤ T, T represents the signal duration, B is the signal bandwidth, f C is the signal carrier frequency, the role of the curvature adjustment factor α(α > 0) is to adjust the curvature of the time - frequency expression, and the role of the slope adjustment factor k(0 ≤ k ≤ B / T) is to adjust the slopes of the time - frequency expression at t = 0 and t = T. f PO The function graph of f(t) is as Figure 2a shown, f PX The function graph of f(t) is as Figure 2b shown, f PY The function graph of f(t) is as Figure 2c shown, f PZ The function graph of f(t) is as Figure 2d shown.
[0141] Furthermore, the expressions of the four sub - PFMB signal models are:
[0142]
[0143]
[0144]
[0145]
[0146] In the present invention, the bionic signal corresponding to the cetacean sound tonal sound ST(t) is SB(t). To construct SB(t), first, identify and extract the time-frequency spectrum profiles of the fundamental frequency and harmonics in the sound. For the fundamental frequency (or harmonics), according to the extreme points and inflection points of its time-frequency spectrum profile, the sound is divided into M segments (M≥1 and M is an integer). Then, use M bionic signal segments to imitate these M segments of cetacean tonal sound calls respectively, and splice the M bionic signal segments end to end to synthesize the fundamental frequency (or harmonics). Finally, superimpose the fundamental frequency and all harmonics in the time domain to synthesize SB(t). The expression of SB(t) containing R harmonics is:
[0147]
[0148] Among them, A r (t) is the envelope, SN r (t) is the segmented bionic signal, and its amplitude has been normalized. R is the maximum harmonic number of the tonal sound, r≥1 and r is a positive integer. When r = 1, A r (t)·SN r (t) represents the fundamental frequency. When 2≤r≤R, A r (t)·SN r (t) represents the rth harmonic. The expression of SN r (t) is:
[0149] SN r (t) = SB r,1 (t - TD r,1 ) +... + SB r,m (t - TD r,m )
[0150] +.. + SB r,M (t - TD r,M ) (10)
[0151] Among them, SB r,m (t) is the mth bionic signal segment in SN r (t). The duration of SB r,m (t) is TB r,m . The time delay TD r,m is the sum of the durations of all bionic signal segments in the signal set {SB r,1 (t),..., SB r,m-1}, that is
[0152] TD r,m = TB r,1 +... + TB r,m-1 (11)
[0153] Among them, the time delay of the first bionic signal segment SB r,1 (t) is set to T D,1 = 0, that is, SB r,1 (t - TD r,1 ) = SB r,1 (t).
[0154] Specifically, the specific implementation steps for the automatic synthesis of cetacean tonal sound are as follows:
[0155] The first step is the preliminary extraction of the fundamental frequency time-frequency spectrum profile 13.
[0156] (101) The tonal sound ST(t) is a time-continuous signal, and the discrete signal generated after sampling is ST[n]. The time-frequency spectrum of ST[n] is expressed as X[l, h], and X[l, h] is obtained by the Short-time Fourier Transform (STFT), where l (1 ≤ l ≤ L and l is an integer) represents the data block number, and h (1 ≤ h ≤ H and h is an integer) represents the frequency index. The spectrum of the l-th data block is expressed as X l [h], and X l [h] is obtained by the Discrete Fourier Transform (DFT). The peak value of X l [h] is P l and is obtained by the following formula
[0157] [P l , F l = max(X l [h]) (12)
[0158] where F l is the frequency index corresponding to P l , that is, X[l, F l = P l . Generally, the energy of the fundamental frequency in tonal sound is the strongest, and with the increase of the harmonic order, the energy gradually decreases. Therefore, the fundamental frequency time-frequency spectrum profile 13 is expressed as f1[l], which is preliminarily obtained by the expression (13);
[0159] f1[l] = argmax(X l [h]) (13)
[0160] (102) Sometimes there are some high-energy time-frequency points on the second harmonic, and their energy is higher than that of the fundamental frequency, which will cause these high-energy time-frequency points to be misidentified as points on the fundamental frequency time-frequency spectrum profile 13. Therefore, when extracting the fundamental frequency time-frequency spectrum profile 13, it is necessary to judge whether each extracted time-frequency point is a point on the second harmonic. Assume that at l = lx At X lx The peak of [h] is P lx and its corresponding frequency index is F lx If the condition in expression (14) is satisfied, then at l = l x the correction of f1[l x is the frequency corresponding to the frequency index F lx / 2.
[0161]
[0162] where λ is the fundamental frequency energy threshold. In the simulation of this embodiment, Figure 3a as shown is the waveform diagram of a sinusoidal tonal sound. Based on Figure 3b the time-frequency spectrum profile 7 of the sinusoidal tonal sound shown, without performing the correction operation, the uncorrected fundamental frequency time-frequency spectrum profile 8 as shown in Figure 4a is extracted. After setting the fundamental frequency energy threshold λ = 0.2 and re-extracting, the corrected fundamental frequency time-frequency spectrum profile 9 as shown in Figure 4b is obtained, which is denoted as f COR [l].
[0163] Second step, filter out outliers.
[0164] (201) Since the tonal sound may contain ocean noise, during the process of extracting the fundamental frequency time-frequency spectrum profile 13, these ocean noises will generate outliers, causing mutations in the corrected fundamental frequency time-frequency spectrum profile 9. Therefore, these outliers need to be filtered out;
[0165] Cluster the time-frequency points in f COR [l] into multiple connected domains. Select a time-frequency point f COR [l] in f COR [l o as the seed point. The first time-frequency point f COR [1] is used as the first seed, and f COR [1] belongs to the first connected domain CC1. If the adjacent time-frequency point f COR [l o to the seed point f COR [l1] satisfies the condition in expression (15), then add f COR [l1] to the connected domain to which f COR [l o belongs.
[0166] |f COR [l1] - f COR [l0]| < T CC1 (15)
[0167] where T CC1 is the frequency threshold of the time-frequency point. Then, take the time-frequency point f COR [l1] as the seed point. If f COR [l1] does not satisfy the condition in Expression (15), it is also taken as the seed point. Next, repeat the process of this step until no time-frequency point can be used as the seed point. Finally, the time-frequency points in f COR [l] will be aggregated into multiple connected domains {CC1, CC2,...};
[0168] (202) Filter out the connected domain where the wild point is located. Take the i-th (i≥1 and i is an integer) connected domain CC i as the seed connected domain, and compare the remaining connected domains with CC i respectively. If the frequency difference between CC i and the j-th (j>i and j is a positive integer) connected domain satisfies the condition in Expression (16), then splice CC i and CC j into one connected domain.
[0169] |CC j [1] - CC i [end]| < T CC2 (16)
[0170] where CC j [1] is the first time-frequency point of CC j , CC i [end] is the last time-frequency point of CC i , and T CC2 is the connected domain frequency threshold. Then take CC j as the seed connected domain. If CC j does not satisfy the condition in Expression (16), it is also taken as the seed connected domain;
[0171] (203) Repeat the process of the previous step until no connected domain can be used as the seed connected domain. Finally, {CC1, CC2,...} will be spliced into multiple larger connected domains, and the connected domain containing the most time-frequency points forms the filtered fundamental frequency time-frequency spectrum profile 10, which is denoted as f FILT [l]. In the simulation of this embodiment, as Figure 4a and Figure 4b show, the fundamental frequency time-frequency spectrum profile 8 before correction and the fundamental frequency time-frequency spectrum profile 9 after correction without filtering wild points have large mutations at the wild point locations. Set the time-frequency point frequency threshold T CC1 = 200 Hz and the connected domain frequency threshold T CC2 = 1 kHz. After filtering out the wild points, the filtered fundamental frequency time-frequency spectrum profile 10 shown in Figure 5a is obtained.
[0172] Step 3, final extraction of the fundamental frequency time-frequency spectrum profile 13.
[0173] (301) After the outliers are filtered, the positions of these outliers in the filtered fundamental frequency time-frequency spectrum profile 10 are vacant. Therefore, the missing time-frequency points 11 at these positions need to be filled to obtain the fundamental frequency time-frequency spectrum profile 13;
[0174] (302) Perform smoothing filtering on f FILT [l] to obtain the smoothed fundamental frequency time-frequency spectrum profile 12, which is denoted as f BS1 [l]. In the simulation of this embodiment, as Figure 5b shown is the smoothed fundamental frequency time-frequency spectrum profile 12;
[0175] (303) In the time-frequency spectrum of the tonal sound, with f BS1 [l] as the center, within the frequency range of ±Δf1 near f BS1 [l], the time-frequency point with the maximum energy is considered to be the missing time-frequency point 11, that is, the data block number l and frequency index h corresponding to the missing time-frequency point 11 satisfy the conditions in expression (17).
[0176]
[0177] When all the missing time-frequency points 11 are filled, the fundamental frequency time-frequency spectrum profile 13 can be obtained. As Figure 5c shown, in the time-frequency diagram of the tonal sound, the missing time-frequency points are searched with the smoothed fundamental frequency time-frequency spectrum profile 12 as the center. When Δf1 = 200 Hz is set, the obtained fundamental frequency time-frequency spectrum profile is as Figure 5d shown.
[0178] Step 4, identify the harmonics and construct the time-frequency spectrum profiles of the harmonics.
[0179] (401) Since the fundamental frequency time-frequency spectrum profile 13 has been extracted, the maximum possible harmonic order in the tonal sound is:
[0180] R max = f S / 2 / (minf1[l]) (18)
[0181] where f S is the sampling frequency of the tonal sound, the maximum possible harmonic order R max is an integer and R max ≥1. If R max = 1, it means that there are no second harmonics and higher harmonics in the tonal sound;
[0182] (402) Spectrum X l [h], within the frequency range [rf1[l] - Δf2, rf1[l] + Δf2], the spectral peak is considered as the candidate spectral peak of the r-th harmonic. Detect all the candidate spectral peaks of the r-th harmonic. If there is a candidate spectral peak that is larger than its adjacent candidate spectral peaks on the left and right by T P (in units of dB), then select it for the next comparison. Among all the selected candidate spectral peaks, the candidate spectral peak with the largest peak value is the effective spectral peak 14. In the simulation of this embodiment, it is set that Δf2 = 200Hz and T P = 10dB. As Figure 6a shown, there is an effective spectral peak 14 of the second harmonic. As Figure 6b shown, there is the spectral profile 13 at the fundamental frequency and all the effective spectral peaks 15 of the second harmonic;
[0183] (403) If the number of effective spectral peaks 14 of the r-th harmonic satisfies the condition of expression (19), then it is determined that the r-th harmonic exists.
[0184] L r > T L ·L (19)
[0185] where T L (0 < T L < 1) is the length threshold. In the simulation of this embodiment, when the length threshold T L = 0.3, the number of effective spectral peaks 14 in the second harmonic exceeds the length threshold T L , thus it is determined that the second harmonic exists;
[0186] (404) Detect in sequence whether there are harmonics from the second harmonic to the R max -th harmonic in the tonal sound, and then the number of harmonics existing in the tonal sound can be obtained. In the simulation of this embodiment, the final determination result is that this sinusoidal tonal sound only has the second harmonic and no higher harmonics;
[0187] (405) Since the frequency of the harmonic is an integer multiple of the frequency of the fundamental frequency, therefore, the spectral profile f r [l] at the r-th harmonic is expressed as
[0188] f r [l] = r·f1[l] (20).
[0189] Fifth step, construct the time-frequency expression of each bionic signal segment.
[0190] (501) Perform smoothing filtering on f1[l] to obtain the smoothed spectral profile 16, which is expressed as f BS2 [l];
[0191] (502) f BS2 The first derivative and the second derivative of [l] are respectively denoted as f BS2 ′[l] and f BS2 ″[l]. f BS2 The extreme points 17 and the inflection points 18 of [l] respectively satisfy the conditions f BS2 ′[l] = 0 and f BS2 ″[l] = 0. In the simulation of this embodiment, as Figure 7a shown, 5 extreme points 17 and 5 inflection points 18 are extracted based on f BS2 [l]. As Figure 7b shown, according to the extreme points 17 and the inflection points 18, the fundamental frequency time-frequency spectrum profile 13 is divided into 11 segments;
[0192] (503) The harmonic time-frequency spectrum profile is also divided into multiple time-frequency spectrum profile segments in the same manner as above;
[0193] (504) For each time-frequency spectrum profile segment, one is selected from the four sub-PFMB signal sub-models to construct a bionic signal segment. For the time-frequency spectrum profiles of each time-frequency spectrum profile segment and the corresponding bionic signal segment, the values of the three parameters of the bandwidth B, the duration T, and the carrier frequency f C are the same. For the m-th time-frequency spectrum profile segment f r,m [l] of the r-th harmonic, its starting frequency is f r,m [1], and the termination frequency is f r,m [end]. If f r,m [1] < f r,m [end], it means that this time-frequency spectrum profile segment is frequency-up modulated, and s PO (t) or s PZ (t) in the four sub-PFMB signal models is selected to construct a bionic signal segment. If f r,m [1] > f r,m [end], it means that this time-frequency spectrum profile segment is frequency-down modulated, and s PX (t) or s PY (t) is selected to construct a bionic signal segment;
[0194] (505) In the time-frequency spectrum profile of the bionic signal segment, the two parameters of the curvature adjustment factor α and the slope adjustment factor k should satisfy the conditions in the expression (21)
[0195] argmin∑(f(α, k) - f r,m [l]) 2 (21)
[0196] where f(α, k) is the time-frequency spectrum profile of the bionic signal segment.
[0197] Step 6: Construct the segmented bionic signal SN r (t).
[0198] (601) The m-th bionic signal segment of the r-th harmonic is SB r,m (t) is expressed as
[0199]
[0200] where f r,m (t) is the time-frequency spectrum profile of SB r,m (t);
[0201] (602) According to expression (10), all the bionic signal segments are spliced end to end in the time domain to obtain the segmented bionic signal SN r (t);
[0202] (603) All the segmented bionic signals SN r (t) are superimposed in the time domain to obtain the normalized bionic signal SN(t), and its expression is
[0203]
[0204] In the simulation of this embodiment, the waveform of the normalized bionic signal is as shown in Figure 8a , and the time-frequency diagram of the normalized bionic signal is as shown in Figure 8b .
[0205] Step 7: Extract the envelope A r (t).
[0206] First, X l [h] is the spectrum of the l-th data block, and the time corresponding to the first sampling point in the l-th data block is τ l . In the frequency range [rf1[l] - Δf2, rf1[l] + Δf2], find the peak value of the r-th harmonic spectrum as P r,l = max(X l [h]).
[0207] Furthermore, the envelope of the extracted r-th harmonic is
[0208] A r (τ l ) = KP r,l (24)
[0209] where K is the amplitude correction factor. By modifying K, the envelope of SB(t) can be made close to the envelope of ST(t). In the simulation of this embodiment, the Hamming window with N = 256 points is used in the STFT, and K = N / 0.27.
[0210] Step 8: Synthesize the bionic signal SB(t).
[0211] Substitute the segmented bionic signal SN r (t) and the envelope A r (t) into Expression (9), and the bionic signal SB(t) can be obtained. In the simulation of this embodiment, the waveform of the sine-type bionic signal is as shown in Figure 9a and the time-frequency diagram of the sine-type bionic signal is as shown in Figure 9b .
[0212] In addition, this embodiment also mimics the other 5 types of tonal sound, including constant frequency type 1, upward frequency modulation type 2, downward frequency modulation type 3, concave type 4, and convex type 5;
[0213] Figure 10a is the waveform of the constant frequency type tonal sound, 10b is the waveform of the constant frequency type bionic signal, Figure 10c is the time-frequency diagram of the constant frequency type tonal sound, Figure 10d is the time-frequency diagram of the constant frequency type bionic signal;
[0214] Figure 11a is the waveform of the upward frequency modulation type tonal sound, Figure 11b is the waveform of the upward frequency modulation type bionic signal, Figure 11c is the time-frequency diagram of the upward frequency modulation type tonal sound, Figure 11d is the time-frequency diagram of the upward frequency modulation type bionic signal;
[0215] Figure 12a is the waveform of the downward frequency modulation type tonal sound, Figure 12b is the waveform of the downward frequency modulation type bionic signal, Figure 12c is the time-frequency diagram of the downward frequency modulation type tonal sound, Figure 12d is the time-frequency diagram of the downward frequency modulation type bionic signal;
[0216] Figure 13a is the waveform of the concave type tonal sound, Figure 13b is the waveform of the concave type bionic signal, Figure 13c is the time-frequency diagram of the concave type tonal sound, Figure 13d is the time-frequency diagram of the concave type bionic signal;
[0217] Figure 14a is the waveform of the convex type tonal sound, Figure 14b is the waveform of the convex type bionic signal, Figure 14c is the time-frequency diagram of the convex type tonal sound, Figure 14d is the time-frequency diagram of the convex type bionic signal;
[0218] Furthermore, in addition to tonal sound, cetaceans also produce some frequency-modulated sounds with longer durations and more complex time-frequency spectral profiles, such as "signature whistle". Figure 15a is the waveform of the signature whistle, Figure 15b is the waveform of the signature whistle biomimetic signal, Figure 15c is the time-frequency diagram of the signature whistle, Figure 15d is the time-frequency diagram of the signature whistle biomimetic signal.
[0219] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit and scope of the present invention as protected by the claims, those of ordinary skill in the art can make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. An automatic synthesis method for cetacean calls based on the time-frequency spectrum, characterized in that Construct tonal sound through four sub-signal models included in the PFMB signal model. For the bionic signal of constructing tonal sound, first identify and extract the time-frequency spectrum profiles of the fundamental frequency and harmonics in the tonal sound. For the fundamental frequency or harmonics, divide the tonal sound into M segments according to the extreme points and inflection points of their time-frequency spectrum profiles, where M ≥ 1 and M is an integer. Then use M bionic signal segments to imitate these M segments of tonal sound respectively, and splice the M bionic signal segments end to end to synthesize the fundamental frequency or harmonics. Finally, superimpose the fundamental frequency and all harmonics in the time domain to synthesize the bionic signal of the tonal sound. Specifically, it includes the following steps: (1) Initially extract the time-frequency spectrum profile of the fundamental frequency of the tonal sound. Specifically: The bionic signal corresponding to the cetacean call tonal sound ST(t) is SB(t). To construct SB(t), first identify and extract the time-frequency spectrum profiles of the fundamental frequency and harmonics in the call. For the fundamental frequency or harmonics, divide the call into M segments according to the extreme points and inflection points of their time-frequency spectrum profiles, where M ≥ 1 and M is an integer. Then use M bionic signal segments to imitate these M segments of cetacean tonal sound calls respectively, and splice the M bionic signal segments end to end to synthesize the fundamental frequency or harmonics. Finally, superimpose the fundamental frequency and all harmonics in the time domain to synthesize SB(t). The expression of SB(t) containing R harmonics is: Among them, A r (t) is the envelope, and SN r (t) is the segmented bionic signal; R is the maximum harmonic order of the tonal sound, r ≥ 1 and r is a positive integer; when r = 1, A r (t)·SN r (t) represents the fundamental frequency. When 2 ≤ r ≤ R, A r (t)·SN r (t) represents the r-th harmonic; the expression of SN r (t) is: SN r S(t) = SB r,1 (t - TD r,1 ) +... + SB r,m (t - TD r,m ) +..+SB r,M (t-TD r,M )(10) Among them, SB r,m (t) is the m-th bionic signal segment in SN r (t), the duration of SB r,m (t) is TB r,m , the time delay TD r,m is the sum of the durations of all bionic signal segments in the signal set {SB r,1 (t),..., SB r,m-1},that is TD r,m = TB r,1 +...+ TB r,m-1 (11) Among them, the time delay of the first bionic signal segment SB r,1 (t) is set to T D,1 = 0, that is, SB r,1 (t - TD r,1 ) = SB r,1 (t); ST(t) represents a time - continuous signal of tonal sound, and the discrete signal generated after sampling is ST[n]; the time - frequency spectrum of ST[n] is denoted as X[l, h], and X[l, h] is obtained by short - time Fourier transform, where l represents the data - block number, 1 ≤ l ≤ L and l is an integer, h represents the frequency index, 1 ≤ h ≤ H and h is an integer; the spectrum of the l - th data - block is denoted as X l [h], X l [h] is obtained by discrete Fourier transform; X l [h]'s peak value is P l and is obtained by the following formula [P l ,F l = max(X l [h]) (12) where F l is the l corresponding frequency index, i.e., X[l, F l = P l ; in the tonal sound, the energy of the fundamental frequency is the strongest, and as the harmonic order increases, the energy gradually decreases; therefore, the spectral profile at the fundamental frequency is represented as f1[l], which is initially obtained from expression (13); f1[l] = argmax(X l [h]) (13) When extracting the spectral profile of the fundamental frequency, it is necessary to determine whether each time-frequency point extracted is a point on the second harmonic; assume that at l = l x where the peak of X lx [h] is P lx and its corresponding frequency index is F lx , if the condition in expression (14) is satisfied, then at l = l x , correct f1[l x to the frequency corresponding to the frequency index F lx / 2; where λ is the fundamental frequency energy threshold; (2) Filter out the outliers generated by ocean noise in the process of step (1). It includes: (201) Aggregate the time-frequency points in the corrected fundamental frequency time-frequency spectrum profile f COR [l] into several connected domains; Select a time-frequency point f COR [l] in f COR [l o as a seed point, the first time-frequency point f COR [1] as the first seed, and f COR [1] belongs to the first connected domain CC1; If the time-frequency point f COR [l o adjacent to the seed point f COR [l1] satisfies the condition in expression (15), then add f COR [l1] to the connected domain to which f COR [l o belongs; f COR [l1]-f COR [l0]|<T CC1 (15) where T CC1 is the frequency threshold of the time-frequency point; then, taking the time-frequency point f COR [l1] as a seed point, if f COR [l1] does not satisfy the condition in expression (15), it is also taken as a seed point; next, repeat the process of this step until there is no time-frequency point that can be used as a seed point; finally, the time-frequency points in f COR [l] will be aggregated into several connected domains {CC1, CC2,...}; (202) Filter the connected component where the outlier is located; Take the i-th connected component CC i as the seed connected component, and compare the remaining connected components with CC i respectively; If the frequency difference between CC i and the j-th connected component satisfies the condition in Expression (16), then splice CC i and CC j into one connected component; i ≥ 1 and i is an integer, j > i and j is a positive integer; |CC j [1]-CC i [end]|<T CC2 (16) Among them, CC j [1] is the first time-frequency point of CC j , and CC i [end] is the last time-frequency point of CC i , where T CC2 is the connected component frequency threshold; then CC j is used as the seed connected component. If CC j does not meet the conditions in Expression (16), it is also used as the seed connected component; (203) Repeat the process of the previous step until there are no connected components that can be used as seed connected components; finally, {CC1, CC2,...} are stitched together into several larger connected components, and the connected component with the most time-frequency points forms the filtered fundamental frequency time-frequency spectrum profile, which is denoted as f FILT [l]; (3) When the outliers are filtered out, the time-frequency points at the positions of the previous outliers in the filtered time-frequency spectrum profile of the fundamental frequency are vacant. Fill in the time-frequency points at the vacant positions to obtain the final time-frequency spectrum profile of the fundamental frequency. It includes: (301) Smoothly filter the spectral contour f of the fundamental frequency after filtering FILT [l] to obtain the smoothed spectral contour f of the fundamental frequency BS1 [l]; (302) In the time-frequency spectrum of the tonal sound, centered at f BS1 [l], within the frequency range of ±Δf1 near f BS1 [l], the time-frequency point with the maximum energy is considered as the missing time-frequency point, that is, the data block number l and the frequency index h corresponding to the missing time-frequency point satisfy the conditions in expression (17); When all the missing time-frequency points are filled in, the time-frequency spectrum profile of the fundamental frequency can be obtained; (4) Identify the harmonics and construct the time-frequency spectrum profiles of the harmonics. It includes: (401) Since the time-frequency spectrum profile of the fundamental frequency has been extracted, the maximum possible harmonic order in the tonal sound is: R max = f S / 2 / (minf1[l]) (18) where f S is the sampling frequency of the tonal sound, and the maximum possible harmonic number R max is an integer and R max ≥ 1. If R max = 1, it means that there are no second harmonics or higher harmonics in the tonal sound; (402) Spectrum X l In [h], the spectral peak within the frequency range [rf1[l] - Δf2, rf1[l] + Δf2] is considered as the candidate spectral peak of the r-th harmonic; detect all candidate spectral peaks of the r-th harmonic. If there is a candidate spectral peak that is larger than its adjacent candidate spectral peaks on the left and right by T P (in units of dB), then select it for the next comparison; among all the selected candidate spectral peaks, the candidate spectral peak with the largest peak value is the effective spectral peak; (403) If the number L of the effective spectral peaks of the r-th harmonic r satisfies the condition of the expression (19), it is determined that the r-th harmonic exists; L r > T L ·L (19) where T L (0 < T L < 1) is the length threshold; (404) Detect in sequence whether there is a harmonic from the second harmonic to the R-th harmonic in the tonal sound, and then the number of harmonics present in the tonal sound can be obtained; max Whether there is a harmonic from the second harmonic to the R-th harmonic in the tonal sound, and then the number of harmonics present in the tonal sound can be obtained; (405) Since the frequency of the harmonic is an integer multiple of the fundamental frequency, the spectral profile f at the r-th harmonic r [l] is expressed as f r [l] = r·f1[l] (20); (5) Construct the time-frequency expressions of each bionic signal segment in the bionic signal. It includes: (501) Smoothly filter the fundamental frequency spectrum profile f1[l] to obtain the smoothed spectrum profile f BS2 [l]; (502)f BS2 The first derivative and the second derivative of [l] are respectively denoted as f BS2 ′[l] and f BS2 ″[l]; f BS2 The extreme points and inflection points of [l] respectively satisfy the conditions f BS2 ′[l] = 0 and f BS2 ″[l] = 0; (503) Divide the harmonic time-frequency spectrum profile into several time-frequency spectrum profile segments in the same way as above; (504) For each time-frequency spectrum profile segment, select one from the four sub-signal models in the PFMB signal model to construct a bionic signal segment; for the time-frequency spectrum profiles of each time-frequency spectrum profile segment and the corresponding bionic signal segment, the values of the three parameters of bandwidth B, duration T, and carrier frequency f C are the same for these three parameters; (505) In the time-frequency spectrum profile of the bionic signal segment, the two parameters of the curvature adjustment factor α and the slope adjustment factor k should satisfy the conditions in expression (21) argmin∑(f(α,k)-f r,m [l]) 2 (21) where f(α,k) is the time-frequency spectrum profile of the bionic signal segment; (6) Construct the segmented bionic signal. It includes: (601) The m-th bionic signal segment of the r-th harmonic is SB r,m is expressed as (t) where f r,m (t) is the time-frequency spectrum profile of SB r,m (t); (602), according to the expression (10), by concatenating all the bionic signal segments head-to-tail in the time domain, the segmented bionic signal SN r (t) can be obtained; (603) Superimpose all segmented bionic signals SN r (t) in the time domain to obtain the normalized bionic signal SN(t), and its expression is (7) Extract the envelope of the tonal sound. It includes: X l If [h] is the spectrum of the l-th data block, then the time corresponding to the first sampling point in the l-th data block is τ l ; within the frequency range [rf1[l] - Δf2, rf1[l] + Δf2], find that the peak value of the r-th harmonic spectrum is P r,l = max(X l [h]); Then the envelope of the rth harmonic extracted is A r (τ l ) = KP r,l (24) where K is the amplitude correction factor. By modifying K, the envelope of SB(t) can be made close to the envelope of ST(t); Substitute the segmented bionic signal SN r (t) and the envelope A r (t) into the expression of the bionic signal SB(t) containing R harmonics, and the bionic signal SB(t) can be obtained; (8) Synthesize the bionic signal.
2. The automatic synthesis method of cetacean calls based on time-frequency spectrum according to claim 1, wherein The time-frequency expressions \(f\) of the four sub-signal models included in the PFMB signal model PO \((t)\), \(f\) PX \((t)\), \(f\) PY \((t)\) and \(f\) PZ (t) are respectively: f PO f(t) = f P f(t) = (B - kT)(t / T) α + kt + f C (1) f PX f(t) = -f P f(t) + B + 2f C = (kT - B)(t / T) α -kt + B + f C (2) f PY f(t) = f P f[-(t - T)] = (B - kT)(1 - t / T) α -k(t - T) + f C (3) f PZ f(t)= -f P [-(t - T)] + B + 2f C =(kT - B)(1 - t / T) α +k(t - T)+B + f C (4) where \(0\leq t\leq T\), \(T\) represents the signal duration, \(B\) is the signal bandwidth, and \(f\) C is the signal carrier frequency. The role of the curvature adjustment factor \(\alpha\) is to adjust the curvature of the time-frequency expression, and the role of the slope adjustment factor \(k\) is to adjust the slopes of the time-frequency expression at \(t = 0\) and \(t = T\), where \(\alpha>0\); \(0\leq k\leq B / T\); The expressions of the four sub-PFMB signal models are:
Citation Information
Patent Citations
Cetacean cry synthesis and modification method
CN111431625A
Dolphin whistle signal spectrum contour extraction method
CN104217722A
CPM modulation-based identification method for whale-imitating animal whish camouflage communication signals
CN112491765A