Pseudo-ear language generation method, device and equipment

By acquiring the effort and intelligibility features from normal speech data, calculating the vocal effort and adjusting multi-domain acoustic parameters, the problem of insufficient naturalness and interpretability in existing pseudo-whisper generation methods is solved, and pseudo-whisper generation and recognition with high naturalness and strong interpretability are achieved.

CN121708897APending Publication Date: 2026-03-20SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511925575.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing pseudo-whisper generation methods cannot generate pseudo-whispers with high naturalness and strong interpretability, resulting in static loudness of the generated speech, coarse spectral envelope processing, and distortion of the excitation source, which affects the naturalness and realism of the speech.

Method used

By acquiring acoustic features related to effort and intelligibility in normal speech data, the vocal effort is calculated, and the multi-domain acoustic parameters of the speech frame, including the time domain, frequency domain, and excitation domain, are adjusted based on this effort. The acoustic features of the pseudo-whisper are dynamically adjusted to simulate the dynamic fluctuations and high-frequency hissing of real whispers.

Benefits of technology

It improves the naturalness and interpretability of pseudo-whispers, enabling the generated pseudo-whispers to reflect the dynamic changes in semantic emphasis, emotional stress, and breathing rhythm, thereby improving the accuracy of the whisper recognition model and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708897A_ABST
    Figure CN121708897A_ABST
Patent Text Reader

Abstract

The invention discloses a pseudo-ear language generation method, device and equipment. The pseudo-ear language generation method comprises the following steps: acquiring normal voice; extracting multi-scale acoustic features of the normal voice; based on the multi-scale acoustic features, determining sounding effort capable of representing sounding force and clearness; and performing joint modulation of the cross-domain acoustic parameters based on the sounding effort, wherein the joint modulation comprises at least two of a time domain, a frequency domain and an excitation domain. The acoustic parameters of the time domain are adjusted based on the sounding effort, so that dynamic fluctuation caused by semantic key points, emotion emphasis or breathing rhythm in real ear language can be reflected; the acoustic parameters of the frequency domain are adjusted based on the sounding effort, so that the phenomenon that the sounding is clearer by force can be simulated; the acoustic parameters of the excitation domain are adjusted based on the sounding effort, and the effect of enhancing the high-frequency hoarseness sound during forced blowing can be simulated; therefore, the naturalness and interpretability of the pseudo-ear language can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice processing, in particular to a pseudo-whisper generation method and device, a whisper recognition model construction method and device, a whisper recognition method and device, and an electronic device. BACKGROUND

[0002] Whisper is a typical abnormal vocalization mode, and its acoustic characteristics are as follows: missing fundamental frequency, excitation source being non-periodic airflow noise, resonance peak structure being fuzzy and overall energy being low. Pseudo-whisper is not recorded by real human vocalization, but is artificially synthesized through technical means to imitate the acoustic characteristics of real whisper, and is mainly used to expand the whisper voice data set. Since real whisper corpus is difficult to collect, the labeling cost is high, and individual differences are significant, therefore, data enhancement generally relies on parameterized vocoders to generate pseudo-whisper.

[0003] At present, a typical pseudo-whisper generation method is fixed parameter transformation method, which is based on macroscopic observation of the differences in acoustic characteristics between whisper and normal speech, and performs simple and fixed rule transformation on the acoustic parameters extracted by the vocoder. The fixed parameter transformation method usually adopts the following strategies: 1) fundamental frequency (F0) processing: the fundamental frequency of all speech frames is forcibly set to zero (F0 = 0) to eliminate the periodic vibration of the vocal cords; 2) spectral envelope (SP) processing: a smoothing processing of a fixed intensity (for example, a low-pass filter with a fixed cutoff frequency or a smoothing filter with a fixed window width) is applied to the spectral envelope of the entire speech. The purpose is to blur the resonance peak and simulate the effect of weakened resonance of the vocal tract in whisper; 3) aperiodicity (AP) processing: the aperiodicity parameter of all frequency bands and all time frames is forcibly set to the maximum value (usually 1.0), which intends to use pure white noise as the excitation source to simulate pure white noise excitation.

[0004] However, the inventors of the present application found that the existing scheme at least has the following problems: 1) the strategy of forcibly setting the fundamental frequency to zero will cause the loudness to be static, that is, the global loudness of the generated speech is constant, and cannot reflect the dynamic fluctuations in real whisper caused by semantic emphasis, emotional emphasis or breathing rhythm; 2) the strategy of applying a smoothing filter of a fixed intensity to the spectral envelope will cause the spectral envelope processing to be rough, that is, a uniform smoothing intensity is used for all phonemes, without distinguishing between vowels (which require strong smoothing to simulate fuzzy resonance peaks) and clear consonants (which should retain high-frequency details to maintain intelligibility), resulting in mechanical and distorted listening experience; 3) the strategy of setting the aperiodicity parameter to all 1 will cause the excitation source to be distorted, that is, pure white noise or fixed AP is used as the excitation, which cannot reflect the spectral structure characteristics (such as high-frequency "sibilance" sound varying with the intensity of vocalization) of airflow noise in real whisper, seriously affecting the naturalness and realism of the generated speech. In summary, how to generate pseudo-whisper with high naturalness and strong interpretability is a problem that needs to be researched and tackled. SUMMARY

[0005] The present application provides a pseudo-whisper generation method to solve the problem that the prior art cannot generate pseudo-whisper with high naturalness and strong explainability. The present application further provides a pseudo-whisper generation device, a whisper recognition model construction method and device, a whisper recognition method and device, and an electronic device.

[0006] The present application provides a pseudo-whisper generation method, comprising: obtaining normal speech data; for a normal speech frame in the normal speech data, obtaining a force-related acoustic feature and a clarity-related acoustic feature of the normal speech frame; obtaining a vocal effort degree of the normal speech frame according to the force-related acoustic feature and the clarity-related acoustic feature, the vocal effort degree being positively correlated with the force and the clarity; adjusting an acoustic parameter of a multi-domain representation of the normal speech frame according to the vocal effort degree, the multi-domain representation comprising at least two of the following: time domain, frequency domain, and excitation domain; synthesizing a pseudo-whisper speech frame corresponding to the normal speech frame according to the adjusted acoustic parameter of the multi-domain representation of the normal speech frame, a plurality of pseudo-whisper speech frames corresponding to a plurality of normal speech frames in the normal speech data forming pseudo-whisper speech data.

[0007] Optionally, the obtaining of the vocal effort degree of the normal speech frame according to the force-related acoustic feature and the clarity-related acoustic feature comprises: performing weighted operation processing according to a force weight and the force-related acoustic feature, and a clarity weight and the clarity-related acoustic feature, and taking a weighted operation value as the vocal effort degree.

[0008] Optionally, the obtaining of the vocal effort degree of the normal speech frame according to the force-related acoustic feature and the clarity-related acoustic feature comprises: obtaining simulated perception data of the human ear on the force-related acoustic feature; obtaining the vocal effort degree according to the simulated perception data and the clarity-related acoustic feature.

[0009] Optionally, the obtaining of the vocal effort degree of the normal speech frame according to the force-related acoustic feature and the clarity-related acoustic feature comprises: performing normalization processing on the force-related acoustic feature and the clarity-related acoustic feature; obtaining the vocal effort degree according to the normalized force-related acoustic feature and the normalized clarity-related acoustic feature; The vocal effort degree is normalized.

[0010] Optionally, the time-domain acoustic parameter comprises a loudness-related parameter as the first parameter. The time-domain acoustic parameter is adjusted according to the vocal effort degree, comprising: A first parameter peak-to-peak value of the normal speech frame is obtained, and a first parameter corresponding to low voice or a first parameter corresponding to high voice is obtained. A fluctuation value of the first parameter is obtained according to the first parameter peak-to-peak value and the vocal effort degree. An adjusted first parameter is obtained according to the first parameter corresponding to low voice or the first parameter corresponding to high voice and the fluctuation value.

[0011] Optionally, the time-domain acoustic parameter is adjusted according to the vocal effort degree, further comprising: A first parameter contrast is obtained. A vocal effort degree affected by the first parameter contrast is obtained according to the vocal effort degree and the first parameter contrast. The fluctuation value of the first parameter is obtained according to the first parameter peak-to-peak value and the vocal effort degree, comprising: The fluctuation value is obtained according to the first parameter peak-to-peak value and the vocal effort degree affected by the first parameter contrast.

[0012] Optionally, the frequency-domain acoustic parameter comprises a smoothing bandwidth of a spectral envelope. The frequency-domain acoustic parameter is adjusted according to the vocal effort degree, comprising: A target reduction amount of the smoothing bandwidth is obtained according to the vocal effort degree. The smoothing bandwidth is adjusted according to the target reduction amount.

[0013] Optionally, the frequency-domain acoustic parameter is adjusted according to the vocal effort degree, further comprising: A spectral centroid is obtained. A target increment of the smoothing bandwidth is obtained according to the spectral centroid; the target increment corresponding to a clear consonant is less than the target increment corresponding to a vowel. The smoothing bandwidth is adjusted according to the target increment.

[0014] Optionally, the excitation-domain acoustic parameter comprises a non-periodicity parameter. The excitation-domain acoustic parameter is adjusted according to the vocal effort degree, comprising: White noise and the non-periodicity parameter are mixed according to the vocal effort degree, so that the vocal effort degree is positively correlated with the mixed non-periodicity parameter.

[0015] Optionally, the white noise and the aperiodic parameter are mixed according to the vocal effort degree, so that the vocal effort degree is positively correlated with the mixed aperiodic parameter, and the method comprises the following steps of: According to the vocal effort degree, a first weight and a second weight are determined, the vocal effort degree is positively correlated with the first weight, and the vocal effort degree is negatively correlated with the second weight; According to the first weight and the second weight, the white noise and the aperiodic parameter are weighted and summed.

[0016] The present application provides a whisper recognition model construction method, comprising: Obtain a plurality of normal voice data; For normal voice frames in the normal voice data, obtain force degree related acoustic features and intelligibility related acoustic features of the normal voice frames; According to the force degree related acoustic features and the intelligibility related acoustic features, the vocal effort degree of the normal voice frame is obtained, and the vocal effort degree is positively correlated with the force degree and the intelligibility; According to the vocal effort degree, the acoustic parameters of the multi-domain representation of the normal voice frame are adjusted; the multi-domain representation comprises at least two of the following: time domain, frequency domain, excitation domain; According to the adjusted acoustic parameters of the multi-domain representation of the normal voice frame, pseudo-whisper voice frames corresponding to the normal voice frames are synthesized, and a plurality of pseudo-whisper voice frames corresponding to a plurality of normal voice frames in the normal voice data form pseudo-whisper voice data corresponding to the normal voice data; According to a plurality of pseudo-whisper voice data corresponding to a plurality of normal voice data, a whisper recognition model is constructed.

[0017] The present application provides a whisper recognition method, comprising: Collecting a whisper signal; According to the whisper signal, the whisper content is obtained through the whisper recognition model; The whisper recognition model is constructed in the following manner: a plurality of normal speech data is acquired; for a normal speech frame in the normal speech data, force-related acoustic features and articulation-related acoustic features of the normal speech frame are acquired; according to the force-related acoustic features and the articulation-related acoustic features, a vocal effort degree of the normal speech frame is acquired, the vocal effort degree being positively correlated with force and articulation; according to the vocal effort degree, acoustic parameters of multi-domain representation of the normal speech frame are adjusted, the multi-domain representation including at least two of the following: a time domain, a frequency domain, and an excitation domain; according to the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, a pseudo-whisper speech frame corresponding to the normal speech frame is synthesized, and a plurality of pseudo-whisper speech frames corresponding to a plurality of normal speech frames in the normal speech data form pseudo-whisper speech data corresponding to the normal speech data; and the whisper recognition model is constructed according to a plurality of pseudo-whisper speech data corresponding to a plurality of normal speech data.

[0018] The application provides a pseudo-whisper generation device, comprising: a normal speech acquisition unit configured to acquire normal speech data; an acoustic feature extraction unit configured to acquire, for a normal speech frame in the normal speech data, force-related acoustic features and articulation-related acoustic features of the normal speech frame; a vocal effort degree calculation unit configured to acquire, according to the force-related acoustic features and the articulation-related acoustic features, a vocal effort degree of the normal speech frame, the vocal effort degree being positively correlated with force and articulation; a cross-domain parameter joint adjustment unit configured to adjust, according to the vocal effort degree, acoustic parameters of multi-domain representation of the normal speech frame, the multi-domain representation including at least two of the following: a time domain, a frequency domain, and an excitation domain; a pseudo-whisper synthesis unit configured to synthesize, according to the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, a pseudo-whisper speech frame corresponding to the normal speech frame, and a plurality of pseudo-whisper speech frames corresponding to a plurality of normal speech frames in the normal speech data form pseudo-whisper speech data.

[0019] The application provides an electronic device, comprising: a processor; and a memory configured to store a program for implementing any of the above methods, and the device is powered on and runs the program of the method through the processor.

[0020] The application also provides a computer-readable storage medium, which stores instructions, when running on a computer, causes the computer to execute the above various methods.

[0021] The application also provides a computer program product comprising instructions which, when executed on a computer, cause the computer to perform the various methods described above.

[0022] Compared with the prior art, the application has the following advantages: The pseudo-whisper generation method provided by the embodiment of the application comprises the following steps: acquiring normal speech data; acquiring force degree related acoustic features and intelligibility related acoustic features of a normal speech frame in the normal speech data; acquiring a vocal effort degree of the normal speech frame according to the force degree related acoustic features and the intelligibility related acoustic features, the vocal effort degree being positively correlated with the force degree and the intelligibility; adjusting acoustic parameters of multi-domain representation of the normal speech frame according to the vocal effort degree; the multi-domain representation comprises at least two of the following: a time domain, a frequency domain and an excitation domain; and synthesizing a pseudo-whisper speech frame corresponding to the normal speech frame according to the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, a plurality of pseudo-whisper speech frames corresponding to a plurality of normal speech frames in the normal speech data forming pseudo-whisper speech data. By using this processing method, the vocal effort degree representing the degrees of “force” and “intelligibility” of voice emission is determined based on multi-scale acoustic features of normal speech, and cross-domain parameter joint modulation is performed based on the vocal effort degree to model the physiological mechanism of whisper voice emission. Since the acoustic parameters of the time domain are adjusted based on the vocal effort degree, the acoustic parameters of the time domain of the synthesized pseudo-whisper change dynamically, which can reflect the dynamic fluctuations caused by semantic emphasis, emotional emphasis or breathing rhythm in real whisper; the acoustic parameters of the frequency domain are adjusted based on the vocal effort degree, which can simulate the phenomenon that forced pronunciation is clearer; the acoustic parameters of the excitation domain are adjusted based on the vocal effort degree, which can simulate the effect of enhanced high-frequency hissing sound when forced blowing; therefore, the naturalness of the pseudo-whisper can be effectively improved. In addition, this processing method also constructs a “white box” pseudo-whisper generation model with strong interpretability, so the interpretability of the pseudo-whisper can be effectively improved.

[0023] The whisper recognition model construction method provided in the embodiments of the present application comprises the following steps: acquiring multiple normal speech data; acquiring force-related acoustic features and intelligibility-related acoustic features of normal speech frames in the normal speech data; acquiring a vocal effort degree of the normal speech frames according to the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort degree is positively correlated with the force and the intelligibility; adjusting acoustic parameters of multi-domain representations of the normal speech frames according to the vocal effort degree, wherein the multi-domain representations comprise at least two of the following: a time domain, a frequency domain, and an excitation domain; synthesizing pseudo-whisper speech frames corresponding to the normal speech frames according to the adjusted acoustic parameters of the multi-domain representations of the normal speech frames, wherein multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data form pseudo-whisper speech data corresponding to the normal speech data; and constructing a whisper recognition model according to multiple pseudo-whisper speech data corresponding to multiple normal speech data. In this way, pseudo-whisper with high naturalness and strong interpretability is generated, and the whisper recognition model is trained according to the pseudo-whisper corpus with high naturalness, so that the accuracy of the whisper recognition model can be effectively improved.

[0024] The whisper recognition method provided in the embodiments of the present application comprises the following steps: acquiring a whisper signal; acquiring whisper content according to the whisper signal through a whisper recognition model, wherein the whisper recognition model is constructed in the following way: acquiring multiple normal speech data; acquiring force-related acoustic features and intelligibility-related acoustic features of normal speech frames in the normal speech data; acquiring a vocal effort degree of the normal speech frames according to the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort degree is positively correlated with the force and the intelligibility; adjusting acoustic parameters of multi-domain representations of the normal speech frames according to the vocal effort degree, wherein the multi-domain representations comprise at least two of the following: a time domain, a frequency domain, and an excitation domain; synthesizing pseudo-whisper speech frames corresponding to the normal speech frames according to the adjusted acoustic parameters of the multi-domain representations of the normal speech frames, wherein multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data form pseudo-whisper speech data corresponding to the normal speech data; and constructing a whisper recognition model according to multiple pseudo-whisper speech data corresponding to multiple normal speech data. In this way, pseudo-whisper with high naturalness and strong interpretability is generated, and the whisper recognition model with high accuracy is trained according to the pseudo-whisper corpus with high naturalness, and the whisper is recognized according to the whisper recognition model with high accuracy; therefore, the whisper recognition accuracy can be effectively improved, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is a flowchart of an embodiment of the pseudo-whisper generation method provided in the present application; Figure 2 This is a schematic diagram of a scenario of an embodiment of the whisper recognition model construction method provided in this application. Detailed Implementation

[0026] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0027] This application provides a method and apparatus for generating pseudo-whispering, a method and apparatus for constructing a whispering recognition model, a method and apparatus for whispering recognition, and an electronic device. The various solutions are described in detail below in each embodiment.

[0028] First Embodiment Please refer to Figure 1 This is a flowchart of the pseudo-whisper generation method provided in this application embodiment. In this embodiment, the method may include the following steps: Step S101: Obtain normal voice data.

[0029] The pseudo-whisper generation method provided in this application generates pseudo-whisper by parametrically processing and transforming normal speech. Therefore, the first step is to acquire normal speech data. Normal speech is the default, normal conversational speech, with volume and clarity suitable for daily communication. Whisper is a typical abnormal mode of speech, a deliberate lowering of the voice to speak softly, such as when afraid of disturbing others. Its acoustic characteristics include: missing fundamental frequency, excitation source being non-periodic airflow noise, blurred formant structure, and overall low energy. Pseudo-whisper is speech synthesized through technical means to imitate the acoustic characteristics of real whisper.

[0030] Step S103: For normal speech frames in normal speech data, obtain the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames.

[0031] The method provided in this application embodiment involves parameterizing and transforming normal speech frames in normal speech through steps S103 to S109, thereby generating pseudo whisper speech frames corresponding to normal speech frames. Multiple pseudo whisper speech frames corresponding to multiple normal speech frames in normal speech form pseudo whisper speech corresponding to normal speech.

[0032] A speech segment typically consists of multiple speech frames. A normal speech segment is called a normal speech frame, and a normal speech segment comprises multiple normal speech frames. The duration of a speech frame is called the frame length, which commonly ranges from 10 milliseconds to 50 milliseconds. For example, a 25-millisecond speech frame corresponds to a duration of 0.025 seconds. A suitable frame length ensures signal stability within the frame while also meeting the requirements of frequency analysis. Step S103 involves acquiring the force-related acoustic features and intelligibility-related acoustic features of any given normal speech frame within the normal speech segment.

[0033] Force-dependent acoustic features include those that reflect the force exerted by speech, such as temporal energy (RMS) and loudness (e.g., 60 dB). Temporal energy (RMS) reflects the physical intensity of speech. Some force-dependent acoustic features are positively correlated with force, while others are negatively correlated. For example, energy is a force-dependent acoustic feature positively correlated with force; energy is the energy transferred when an object vibrates, and the greater the energy, the greater the force exerted.

[0034] Speech intelligibility-related acoustic features include those reflecting speech intelligibility, such as spectral flatness (SFM), spectral centrality, signal intelligibility, noise interference level, spectral centroid (SC), and harmonic noise ratio (HNR). The spectral centroid (SC) reflects the center of spectral energy distribution; a high spectral centroid typically corresponds to voiceless consonants (e.g., / s / ), while a low spectral centroid corresponds to vowels (e.g., / a / ). The harmonic noise ratio (HNR) measures the ratio of periodic components to noise components in speech. Spectral flatness (SFM) measures the flatness of the spectrum; a lower value indicates a steeper spectrum, usually indicating the presence of clear formants (e.g., vowels). Some speech intelligibility-related acoustic features are positively correlated with intelligibility, while others are negatively correlated. For example, spectral flatness, which describes the uniformity of the spectrum distribution, is negatively correlated with intelligibility; a lower spectral flatness indicates higher intelligibility.

[0035] In one example, step S103 can be implemented as follows: obtain the basic acoustic parameters of a normal speech frame through a vocoder; and obtain the force-related acoustic features and intelligibility-related acoustic features of the normal speech frame based on the basic acoustic parameters.

[0036] Basic acoustic parameters include, but are not limited to: fundamental frequency (F0), spectral envelope (SP), and aperiodicity (AP). The fundamental frequency determines the pitch of a sound and is the lowest harmonic frequency in a periodic sound. The spectral envelope (SP) describes the smooth shape of the speech signal's spectrum, primarily reflecting the resonance characteristics of the vocal tracts (oral cavity and nasal cavity), and determines the timbre of vowels (i.e., formant structure). In whispered speech, the formant structure becomes blurred. The aperiodicity (AP) describes the proportion of aperiodic components (such as airflow noise) to periodic components (vocal cord vibration) in the speech excitation source. For whispered speech, the excitation source is primarily airflow noise.

[0037] In practice, vocoders (such as WORLD and Synthesizer V) can be used to analyze the input normal speech data, extracting fundamental acoustic parameters such as fundamental frequency, spectral envelope, and aperiodic parameters frame by frame. Speech processing libraries (such as librosa) can be used to calculate multi-scale acoustic features of the normal speech data based on these fundamental acoustic parameters, including force-related acoustic features and profilometry-related acoustic features. A vocoder is a speech analysis and synthesis system. It can decompose a speech signal into a set of fundamental parameters describing its acoustic characteristics (such as fundamental frequency, spectral envelope, and aperiodic parameters), and can also resynthesize speech based on these fundamental parameters.

[0038] Step S105: Obtain the vocal effort of the normal speech frame based on the force-related acoustic features and the clarity-related acoustic features.

[0039] Unified Vocal Effort does not correspond to a single physical quantity, but rather integrates force-related acoustic features (such as energy) and intelligibility-related acoustic features (such as spectral flatness) of speech, representing the degree of "effort" and "clarity" in vocalization. Vocal effort is positively correlated with force and intelligibility; the higher the vocal effort, the more "effort" and "clarity" the vocalization. Vocal effort is a dynamically changing data point calculated through modeling; different speech frames within a normal speech segment may correspond to different vocal effort levels.

[0040] In one example, step S105 can be implemented as follows: A weighted calculation is performed based on the force weight and the force-related acoustic features, as well as the intelligibility weight and the intelligibility-related acoustic features, and the weighted value is used as the vocal effort. This processing method allows the contribution of the force-related acoustic features and the intelligibility-related acoustic features to the vocal effort to be controlled by the force weight and intelligibility weight; therefore, the interpretability of the vocal effort can be effectively improved.

[0041] In one example, the vocal effort is positively correlated with the force-related acoustic feature and positively correlated with the clarity-related acoustic feature; step S105 can be implemented as follows: according to the force weight and the clarity weight, the force-related acoustic feature and the clarity-related acoustic feature are weighted and summed, and the weighted sum is used as the vocal effort.

[0042] In another example, the vocal effort is positively correlated with the force-related acoustic feature, and negatively correlated with the intelligibility-related acoustic feature. For example, the force-related acoustic feature is speech energy; the intelligibility-related acoustic feature is spectral flatness, which is negatively correlated with intelligibility—the smaller the spectral flatness, the greater the intelligibility. Since vocal effort is positively correlated with intelligibility, spectral flatness is also negatively correlated with vocal effort—the smaller the spectral flatness, the greater the vocal effort. Accordingly, step S105 can be implemented as follows: obtain intelligibility data based on the spectral flatness; perform a weighted summation of the force-related acoustic feature and the intelligibility data based on the force weight and intelligibility weight, and use the weighted summation value as the vocal effort.

[0043] In one example, step S105 can be implemented as follows: acquiring simulated perception data of the human ear on the force-related acoustic features; and acquiring the vocal effort based on the simulated perception data and the intelligibility-related acoustic features. Force-related acoustic features are objective physical quantities of sound, while simulated perception data of the human ear on these features are simulated data on subjective human perception. For example, the force-related acoustic feature is time-domain energy (RMS), and simulated perception data of time-domain energy is loudness; energy is the intensity of the sound itself, and loudness is the perceived loudness by the human ear; the unit of energy can be joules or watts, and the unit of loudness can be decibels (dB). This processing method simulates the human ear's perception of vocal effort; therefore, it can effectively improve the accuracy of vocal effort.

[0044] In specific implementation, when the force-related acoustic feature is time-domain energy (RMS), the simulated perception data of the human ear on the force-related acoustic feature can be obtained in the following way: the logarithm of the energy is used as the simulated perception data.

[0045] In one example, step S105 can be implemented as follows: normalizing the force-related acoustic features and the intelligibility-related acoustic features; obtaining the vocal effort of the normal speech frame based on the normalized force-related acoustic features and the normalized intelligibility-related acoustic features; and normalizing the vocal effort. This processing method makes the vocal effort a scalar value (between 0 and 1), making the vocal effort, force-related acoustic features, and intelligibility-related acoustic features more comparable; therefore, it can effectively improve the accuracy of the vocal effort.

[0046] In one example, the force-related acoustic features of a normal speech frame include speech energy, and the intelligibility-related acoustic features of a normal speech frame include spectral flatness. Step S105 can be implemented as follows: normalize the speech energy and the spectral flatness; use the logarithm of the normalized speech energy as simulated perceptual data of human ear loudness; perform weighted summation on the simulated perceptual data and the complement of the spectral flatness according to the force weight and the intelligibility weight; normalize the weighted summation result as the vocal effort of the normal speech frame. This processing method allows the vocal effort to be obtained by combining the speech energy (force) and spectral flatness (intelligence). This method of obtaining vocal effort can be expressed by the following formula: Effort(t) = σ[ w_E · log(1 + α · E_norm(t)) + w_SFM · (1 - SFM_norm(t)) ] Here, E_norm(t) and SFM_norm(t) are the normalized (between 0 and 1) values ​​of frame-level (normal speech frame) energy and spectral flatness, respectively. log(1 + α · E_norm(t)) is the energy term; the larger the energy, the larger this term value, representing more effort. The log is used to simulate the logarithmic perception of loudness by the human ear. (1 - SFM_norm(t)) is the normalized complement of spectral flatness, representing the intelligibility term; the lower the spectral flatness (i.e., the larger 1 - SFM), the clearer the formants, the more distinct the pronunciation, representing more effort in trying to make the other party hear clearly. w_E and w_SFM are the force weight and intelligibility weight, respectively, used to balance the contributions of "force" and "intelligence" in the final effort. They can be preset or learned through a small amount of data. σ(·) is the Sigmoid function, which compresses the final weighted summation result to the [0, 1] interval, making it a standard, normalized control signal.

[0047] Step S107: Adjust the acoustic parameters of the multi-domain representation of the normal speech frame according to the vocal effort.

[0048] Vocal effort is used as a high-level, unified control signal to coordinate the acoustic parameters of the multi-domain representation of normal speech frames. The multi-domain representation includes at least two of the time domain, frequency domain, and excitation domain to simulate the coupling relationship between various physiological components during actual vocalization.

[0049] The time domain, frequency domain, and excitation domain are the core dimensions for describing sound signals in sound signal processing. The time domain of a sound signal uses time as its horizontal axis, directly recording the state of the sound signal as it changes over time. Acoustic parameters in the time domain include, but are not limited to, loudness-related parameters, which affect the intensity of the sound, such as loudness, energy, and vibration amplitude. The frequency domain of a sound signal uses frequency (pitch) as a reference, describing which different frequencies of pure tones constitute the sound and the intensity of each pure tone. Acoustic parameters in the frequency domain include, but are not limited to, spectral envelope (SP), bandwidth, and spectral density. The spectral envelope refers to the amplitude variation trend of frequency components. Acoustic parameters in the frequency domain affect the clarity of the sound. The excitation domain of a sound signal uses the input parameters that generate the sound as a reference, describing the sound output response of the sound-generating system under different inputs (excitations). Acoustic parameters in the excitation domain include, but are not limited to, aperiodic parameter (AP), signal-to-noise ratio (SNR), and distortion. Acoustic parameters in the excitation domain affect the naturalness and emotional expression of the sound.

[0050] This step employs cross-domain parameter joint modulation, which no longer modifies individual acoustic parameters in isolation and statically. Instead, it coordinates and dynamically adjusts parameters from different domains, such as loudness (time domain), spectral envelope (frequency domain), and excitation source (excitation domain), through a unified control signal (i.e., vocal effort).

[0051] In one example, the acoustic parameters in the time domain include loudness-related parameters as the first parameter; adjusting the acoustic parameters in the time domain according to the vocal effort may include the following sub-steps: obtaining the peak-to-peak value of the first parameter of the normal speech frame, obtaining the first parameter corresponding to low sound or high sound; obtaining the fluctuation value of the first parameter according to the peak-to-peak value of the first parameter and the vocal effort; obtaining the adjusted first parameter according to the first parameter corresponding to low sound or high sound and the fluctuation value.

[0052] Loudness-related parameters affect the intensity of sound, such as loudness, energy, and vibration amplitude. For ease of description, this application refers to loudness-related parameters as the first parameter. The peak-to-peak value of the first parameter is the difference between the first parameter of a high-pitched sound and the first parameter of a low-pitched sound within the same speech frame, used to describe the fluctuation range of the first parameter within the speech frame. Based on the peak-to-peak value of the first parameter in a normal speech frame and the vocal effort in the normal speech frame, the fluctuation value of the first parameter in the normal speech frame can be obtained. For example, if the first parameter is speech loudness, the peak-to-peak value of the first parameter is the peak-to-peak value of loudness, and the fluctuation value of the first parameter can be the product of the peak-to-peak value of loudness and the vocal effort.

[0053] This processing method ensures that the adjusted first parameter is based on the low-pitched first parameter of the normal speech frame, plus the fluctuation value of the first parameter affected by the vocal effort; the greater the vocal effort, the greater the fluctuation of the first parameter; if the vocal effort is zero, the adjusted first parameter is the low-pitched first parameter; if the vocal effort is one, the adjusted first parameter is the high-pitched first parameter; thus, a pseudo-whisper loudness is achieved based on dynamic adjustment of vocal effort, and the pseudo-whisper loudness is related to the force and clarity of speech.

[0054] For example, the loudness-related parameters include speech loudness; adjusting the acoustic parameters in the time domain according to the vocal effort may include the following sub-steps: obtaining the peak-to-peak loudness and low-frequency loudness of the normal speech frame; obtaining the product of the peak-to-peak loudness and the vocal effort as the loudness fluctuation value; and using the sum of the low-frequency loudness and the loudness fluctuation value as the adjusted speech loudness. Peak-to-peak loudness is the loudness difference between high-frequency and low-frequency loudness within the same speech frame, used to describe the range of loudness fluctuation within the speech frame. The product of the peak-to-peak loudness and the vocal effort of the normal speech frame is the loudness fluctuation term. This processing method ensures that the adjusted speech loudness is based on the low loudness of the normal speech frame, plus a loudness fluctuation value affected by vocal effort; the greater the vocal effort, the greater the loudness fluctuation; if the vocal effort is zero, the adjusted speech loudness is low loudness; if the vocal effort is one, the adjusted speech loudness is high loudness; thus, a pseudo-whisper loudness is achieved based on dynamic adjustment of vocal effort, which is related to the force and clarity of the speech.

[0055] For example, adjusting the acoustic parameters in the time domain based on the vocal effort can be achieved as follows: obtain the peak-to-peak loudness and pitch loudness of the normal speech frame; obtain the product of the peak-to-peak loudness and the vocal effort as the loudness fluctuation value; and obtain the adjusted speech loudness based on the pitch loudness and the loudness fluctuation value.

[0056] In practice, the speech loudness in the above processing method can be replaced with other loudness-related parameters such as speech energy. The specific processing procedure is the same as the principle of the above processing method, and will not be repeated here.

[0057] In one example, adjusting the acoustic parameters in the time domain based on the vocal effort may further include the following sub-steps: obtaining a first parameter contrast; obtaining the vocal effort affected by the first parameter contrast based on the vocal effort and the first parameter contrast; obtaining the fluctuation value of the first parameter based on the peak-to-peak value of the first parameter and the vocal effort includes: obtaining the fluctuation value based on the peak-to-peak value of the first parameter and the vocal effort affected by the first parameter contrast. This processing method ensures that the fluctuation value of the first parameter is controlled not only by the vocal effort but also by the first parameter contrast. A first parameter contrast greater than one will cause more drastic changes in the first parameter, while a contrast less than one will result in smoother changes. Therefore, it can effectively improve the controllability and interpretability of pseudo-whispers.

[0058] For example, adjusting the acoustic parameters in the time domain based on the vocal effort may further include the following sub-steps: obtaining loudness contrast; obtaining a power derived from the vocal effort as the base and the loudness contrast as the exponent, as the vocal effort affected by the loudness contrast; obtaining the product of the loudness peak-to-peak value and the vocal effort includes obtaining the product of the loudness peak-to-peak value and the power. This processing method ensures that the loudness fluctuation is controlled not only by the vocal effort but also by the loudness contrast. A loudness contrast greater than one results in more dramatic loudness changes, while a contrast less than one results in smoother changes. Therefore, it can effectively improve the controllability and interpretability of pseudo-whispers. This effort-driven dynamic loudness mapping method can be expressed by the following formula: L_target(t) = L_min + [Effort(t)]^β · (L_max - L_min) This formula directly controls the target loudness L_target(t) based on the vocal effort Effort(t). When Effort(t) is 0, the loudness is the low-frequency loudness L_min; when it is 1, the loudness is the high-frequency loudness L_max. (L_max - L_min) is the peak-to-peak loudness. The loudness contrast β parameter controls the "contrast" of the loudness change. β > 1 will make the loudness change more drastic, while β < 1 will make it more gradual.

[0059] In one example, the acoustic parameters in the frequency domain include the smoothing bandwidth of the spectral envelope. Adjusting the acoustic parameters in the frequency domain according to the vocal effort includes: obtaining a target reduction in the smoothing bandwidth based on the vocal effort; and adjusting the smoothing bandwidth based on the target reduction. Vocal effort is positively correlated with the target reduction in smoothing bandwidth; that is, the greater the vocal effort, the greater the target reduction, and the smaller the adjusted smoothing bandwidth. This processing method determines the reduction in the smoothing bandwidth of the smoothing filter based on the vocal effort, effectively "reducing smoothing" and simulating the phenomenon of "more forceful vocalization resulting in clearer sound."

[0060] In one example, adjusting the acoustic parameters in the frequency domain according to the vocal effort further includes: obtaining the spectral centroid; obtaining the target increment of the smoothing bandwidth according to the spectral centroid; the target increment corresponding to voiceless consonants is less than the target increment corresponding to vowels; and adjusting the smoothing bandwidth according to the target increment.

[0061] The spectral envelope (SP) describes the smooth shape of the speech signal spectrum, primarily reflecting the resonance characteristics of the vocal tract (oral cavity and nasal cavity), and determining the timbre (i.e., formant structure) of vowels. In whispered speech, the formant structure becomes blurred. The method provided in this application simulates this phenomenon by adaptively smoothing the spectral envelope. The spectral centroid is the center of gravity of the sound spectrum and reflects the brightness of the sound. When voiceless consonants are pronounced, the vocal cords do not vibrate, and the energy is mainly concentrated in the high-frequency region, so the spectral centroid value is high, making it sound bright and sharp; while when vowels are pronounced, the vocal cords vibrate, and the energy is mainly concentrated in the low-frequency region, so the spectral centroid value is low, making it sound thick and full. This processing method makes the target increment corresponding to voiceless consonants smaller than the target increment corresponding to vowels. This preserves the detailed changes in the high-frequency region of voiceless consonants, accurately presents their spectral shape, has the texture of breathy sound, and makes the characteristics of voiceless consonants more prominent; at the same time, it increases the blurring of the formants; thus, adaptive spectral envelope smoothing is achieved.

[0062] In practical implementation, the adaptive spectral envelope smoothing method can be expressed by the following formula: B(t) = B_base + ΔB_phoneme(f_c(t)) + ΔB_effort(Effort(t)) Where B(t) is the instantaneous bandwidth of the smoothing filter; the larger the bandwidth, the stronger the smoothing and the more blurred the formants. B_base is the initial smoothing amount, serving as a base smoothing amount. ΔB_phoneme(f_c(t)) is the target increment, determined by the spectral centroid f_c(t). If f_c(t) is negative (corresponding to voiceless consonants), the increment of the smoothing bandwidth is smaller, weakening high-frequency glitches and highlighting core features, making the acoustic differences between different voiceless consonants clearer. If f_c(t) is positive (corresponding to vowels), the increment of the smoothing bandwidth is larger, making the low-frequency formant contours softer and the value of the spectral centroid more stable. ΔB_effort(Effort(t)) = -ΔB_E · Effort(t) indicates that the target reduction amount is determined by effort; ΔB_E is a set coefficient; the higher Effort(t), the larger the negative value of this term, the weaker the smoothing, simulating the phenomenon of "more forceful pronunciation is clearer".

[0063] In one example, the acoustic parameters of the excitation domain include the aperiodic parameters (AP) of a normal speech frame. A normal speech frame is a mixture of multiple frequency components. The speech excitation source includes aperiodic components (such as airflow noise) and periodic components (vocal cord vibration). Vocal cord vibration dominates the low-frequency fundamental band, corresponding to vowels, voiced consonants, and other major sounds; airflow noise dominates the mid-to-high frequency / high-frequency band, corresponding to voiceless consonants. The superposition of airflow noise and vocal cord vibration constitutes a complete speech frame. The aperiodic parameters (AP) are a general term for the aperiodic signal characteristics in speech, a set of parameters including, but not limited to: the aperiodic component ratio (ACR), aperiodic component energy (sound pressure level), frequency concentration (bandwidth / peak value), noise fluctuation, etc. The aperiodic component ratio can be the proportion of aperiodic components (such as airflow noise) to all components (including airflow noise and vocal cord vibration) in the speech excitation source, or the proportion of aperiodic components to periodic components (vocal cord vibration) in the speech excitation source.

[0064] In this embodiment, adjusting the aperiodic parameters of the normal speech frame according to the vocal effort can be achieved in the following way: mixing white noise and aperiodic parameters according to the vocal effort, so that the vocal effort is positively correlated with the mixed aperiodic parameters.

[0065] In practical implementation, white noise can be pure white noise. Pure white noise has a completely constant power spectral density across the entire frequency range (from 0 to infinity), and the phases of each frequency component are random. White noise covers the audible frequency band of 20Hz to 20kHz, with a uniform frequency distribution, and is an engineering simplification of pure white noise. For whispering, the excitation source is mainly airflow noise; when a whisper is forcefully produced, a high-frequency hissing sound will occur. The method provided in this application does not use pure white noise (spectrally flat) or fixed aperiodic parameters. Instead, it mixes white noise with the aperiodic parameters of a normal speech frame based on vocal effort. Vocal effort is positively correlated with the structured mixed aperiodic parameters, meaning that the higher the vocal effort, the greater the difference between the mixed aperiodic parameters and the original aperiodic parameters, resulting in more prominent high-frequency noise in the speech frame and weaker original aperiodic components, making it sound more indistinct. Conversely, the lower the vocal effort, the less the difference between the mixed aperiodic parameters and the original aperiodic parameters, resulting in more prominent original aperiodic components and clearer sound. In other words, high-frequency noise increases with increasing vocal effort. This constructs a "hybrid excitation source" to finely model the noise characteristics of whispers, thus simulating the whisper excitation source and simulating the effect of enhanced high-frequency hissing when forcefully blowing air.

[0066] In one example, the mixing of white noise and aperiodic parameters based on the vocal effort, such that the vocal effort is positively correlated with the mixed aperiodic parameters, includes: determining a first weight and a second weight based on the vocal effort, wherein the vocal effort is positively correlated with the first weight and negatively correlated with the second weight; and performing a weighted summation of the white noise and aperiodic parameters based on the first weight and the second weight. This processing method allows for the mixing of white noise and the original aperiodic parameters of the normal speech frame through a weighted approach. Since the vocal effort is positively correlated with the first weight of the white noise, the higher the vocal effort, the greater the weight of the white noise in the high-frequency components, thus making the white noise in the high-frequency components more prominent. Since the vocal effort is negatively correlated with the second weight of the aperiodic parameters of the normal speech frame, the lower the vocal effort, the smaller the weight of the white noise in the high-frequency components, and the more prominent the original aperiodic components in the high-frequency components. By employing this weighted summation method of non-periodic parameter mixing, the mixing weights are dynamically controlled by the vocal effort. By constructing a "structured hybrid excitation source" to finely model the noise characteristics of whispers, the performance of non-periodic parameter fusion can be effectively improved.

[0067] In one example, adjusting the acoustic parameters of the excitation domain according to the vocal effort may further include the following steps: smoothing the aperiodic parameters; the mixing of white noise and the aperiodic parameters of the normal speech frame according to the vocal effort can be achieved as follows: mixing white noise and the smoothed aperiodic parameters according to the vocal effort. This processing method first smooths the original aperiodic parameters of the normal speech frame, which reduces the harshness and sharpness of the high-frequency bands of the normal speech frame, resulting in a softer sound corresponding to the smoothed aperiodic parameters; then, mixing the white noise and the smoothed aperiodic parameters according to the vocal effort effectively improves the realism of the high-frequency hissing sound in the mixed whisper.

[0068] In one example, determining the first weight based on vocal effort can be achieved as follows: determine the cutoff frequency of the high-pass filter based on vocal effort, such that vocal effort is negatively correlated with the cutoff frequency; obtain the first weight through the high-pass filter.

[0069] A high-pass filter (HPF) assigns frequency-varying weights to signals of different frequencies. Through frequency-selective attenuation / retention, it grants high weights to high-frequency signals (allowing them to pass through smoothly) and low weights to low-frequency signals (suppressing them), effectively assigning a dynamic weight to each frequency. The HPF's weights are determined by its frequency response function: for each frequency component in the input signal, the HPF calculates a gain coefficient (weight value), and the output signal = input signal * gain coefficient. The gain coefficient for high-frequency components is approximately 1 (high weight, almost no attenuation). The gain coefficient for low-frequency components is less than 1 (low weight, attenuated, the lower the frequency, the smaller the weight).

[0070] The method provided in this application embodiment uses a first weight that is dynamically controlled by vocal effort and is frequency-dependent. The cutoff frequency (fc) of the high-pass filter is controlled by the vocal effort; the cutoff frequency determines the weight boundary, and the lower the cutoff frequency, the greater the weight of high-frequency frequencies. Since vocal effort is negatively correlated with the cutoff frequency, the greater the vocal effort, the lower the cutoff frequency, and the greater the weight of high frequencies. The first weight, being frequency-dependent, is obtained by using the high-pass filter to acquire the weight of the frequency band containing the non-periodic component, thus simulating the physiological phenomenon that "the greater the effort, the stronger the high-frequency hissing sound."

[0071] In practice, other methods can also be used to determine the frequency-related weighting function based on the vocal effort.

[0072] In one example, the processing method of the above-mentioned structured hybrid excitation source can be expressed by the following formula: AP_hybrid(f, t) = w(f, Effort(t)) · 1 + (1 - w(f, Effort(t))) · AP_smooth(f, t) This formula is used to mix pure white noise (represented by AP=1) with the smoothed original aperiodic parameter AP_smooth. The mixing weight w(f, Effort(t)) (i.e., the first weight) is determined by the following function: w(f,Effort(t))=1-exp(-(f / (κ-Δκ·Effort(t)))^δ) .

[0073] This weighting function is a high-pass filter controlled by the vocal effort Effort(t). As Effort(t) increases, the denominator κ - Δκ · Effort(t) decreases, causing the "cutoff frequency" of the high-pass filter to shift to lower frequencies. This means that the higher the vocal effort, the greater the weight of pure white noise in the high-frequency range, thus simulating the physical effect of "increased high-frequency hissing when blowing hard."

[0074] Step S109: Based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, synthesize a pseudo-whisper speech frame corresponding to the normal speech frame, and form pseudo-whisper speech data with multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data.

[0075] This step synthesizes a pseudo-whisper waveform based on the cross-domain parameters dynamically modulated by the vocal effort system. Specifically, a vocoder can synthesize a pseudo-whisper speech frame corresponding to the normal speech frame using adjusted acoustic parameters derived from the multi-domain representation of the normal speech frame. A vocoder is a speech analysis and synthesis system. It can decompose a normal speech signal into a set of fundamental parameters describing its acoustic characteristics (i.e., cross-domain acoustic parameters adjusted based on effort, such as fundamental frequency, spectral envelope, and aperiodic parameters), and can also resynthesize the pseudo-whisper speech based on parameters adjusted from these fundamental parameters (such as adjusted loudness, adaptively smoothed spectral envelope, and aperiodic parameters after structured mixing).

[0076] In one example, cross-domain parameter joint modulation, centered on vocal effort, performs coordinated dynamic modulation on the loudness, spectral envelope, and excitation source of a normal speech frame. Using a vocoder, for each normal speech frame, a pseudo-whisper speech frame is synthesized based on the adjusted loudness, adaptively smoothed spectral envelope, and structured mixed aperiodic parameters corresponding to the normal speech frame. This pseudo-whisper speech data is then formed by combining multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data.

[0077] In one example, step S109 can be implemented as follows: the fundamental frequency of the normal speech frame is set as the initial fundamental frequency; the pseudo-whisper speech frame is synthesized based on the initial fundamental frequency and the adjusted acoustic parameters. For example, the original fundamental frequency (F0) sequence is set to zero; the spectral envelope (SP_mod) after dynamic loudness adjustment and adaptive smoothing, and the aperiodic parameters (AP_mod) after structured mixing are input into the synthesizer part of the vocoder to generate a highly natural pseudo-whisper speech file.

[0078] As can be seen from the above embodiments, the pseudo-whisper generation method provided in this application obtains normal speech data; for normal speech frames in the normal speech data, it obtains the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames; based on the force-related acoustic features and the intelligibility-related acoustic features, it obtains the vocal effort of the normal speech frames, wherein the vocal effort is positively correlated with force and intelligibility; based on the vocal effort, it adjusts the acoustic parameters of the multi-domain representation of the normal speech frames; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, it synthesizes pseudo-whisper speech frames corresponding to the normal speech frames, and forms pseudo-whisper speech data with multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data. This processing method enables the determination of vocal effort that represents the degree of vocal "force" and "intelligence" based on the multi-scale acoustic features of normal speech, and performs cross-domain parameter joint modulation based on the vocal effort to model the physiological mechanism of whisper vocalization. By adjusting the acoustic parameters in the time domain based on vocal effort, the synthesized pseudo-whisper's time-domain acoustic parameters dynamically change, reflecting the dynamic fluctuations in real whisper caused by semantic emphasis, emotional stress, or breathing rhythm. Adjusting the acoustic parameters in the frequency domain based on vocal effort simulates the phenomenon of clearer pronunciation with effort. Adjusting the acoustic parameters in the excitation domain based on vocal effort simulates the effect of enhanced high-frequency hissing when forcefully exhaling. Therefore, the naturalness of the pseudo-whisper can be effectively improved. Furthermore, this approach allows for the construction of a highly interpretable "white-box" generation model, thus effectively enhancing the interpretability of the pseudo-whisper.

[0079] Second Embodiment In the above embodiments, a method for generating pseudo-whispers is provided. Correspondingly, this application also provides a device for generating pseudo-whispers. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0080] This application also provides a pseudo-whisper generation device, comprising: a normal speech acquisition unit for acquiring normal speech data; an acoustic feature extraction unit for acquiring force-related acoustic features and intelligibility-related acoustic features of normal speech frames in the normal speech data; a vocal effort calculation unit for acquiring the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort is positively correlated with force and intelligibility; a cross-domain parameter joint adjustment unit for adjusting the acoustic parameters of the multi-domain representation of the normal speech frame based on the vocal effort; wherein the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; and a pseudo-whisper synthesis unit for synthesizing a pseudo-whisper speech frame corresponding to the normal speech frame based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, thereby forming pseudo-whisper speech data with multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data.

[0081] In one example, obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the clarity-related acoustic features includes: performing a weighted calculation based on the force weight and the force-related acoustic features, and the clarity weight and the clarity-related acoustic features, and using the weighted calculation value as the vocal effort.

[0082] In one example, obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features includes: obtaining simulated perception data of the human ear on the force-related acoustic features; and obtaining the vocal effort based on the simulated perception data and the intelligibility-related acoustic features.

[0083] In one example, obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features includes: normalizing the force-related acoustic features and the intelligibility-related acoustic features; obtaining the vocal effort based on the normalized force-related acoustic features and the normalized intelligibility-related acoustic features; and normalizing the vocal effort.

[0084] In one example, the acoustic parameters in the time domain include loudness-related parameters as the first parameter; adjusting the acoustic parameters in the time domain according to the vocal effort includes: obtaining the peak-to-peak value of the first parameter of the normal speech frame, obtaining the first parameter corresponding to low sound or high sound; obtaining the fluctuation value of the first parameter according to the peak-to-peak value of the first parameter and the vocal effort; obtaining the adjusted first parameter according to the first parameter corresponding to low sound or high sound and the fluctuation value.

[0085] In one example, adjusting the acoustic parameters in the time domain according to the vocal effort further includes: obtaining a first parameter contrast; obtaining a vocal effort affected by the first parameter contrast based on the vocal effort and the first parameter contrast; obtaining the fluctuation value of the first parameter based on the peak-to-peak value of the first parameter and the vocal effort includes: obtaining the fluctuation value based on the peak-to-peak value of the first parameter and the vocal effort affected by the first parameter contrast.

[0086] In one example, the acoustic parameters in the frequency domain include the smoothing bandwidth of the spectral envelope; adjusting the acoustic parameters in the frequency domain according to the phonation effort includes: obtaining a target reduction in the smoothing bandwidth according to the phonation effort; and adjusting the smoothing bandwidth according to the target reduction.

[0087] In one example, adjusting the acoustic parameters in the frequency domain according to the vocal effort further includes: obtaining the spectral centroid; obtaining the target increment of the smoothing bandwidth according to the spectral centroid; the target increment corresponding to voiceless consonants is less than the target increment corresponding to vowels; and adjusting the smoothing bandwidth according to the target increment.

[0088] In one example, the acoustic parameters of the excitation domain include aperiodic parameters; adjusting the acoustic parameters of the excitation domain according to the vocal effort includes: mixing white noise and aperiodic parameters according to the vocal effort, such that the vocal effort is positively correlated with the mixed aperiodic parameters.

[0089] In one example, the step of mixing white noise and aperiodic parameters according to the vocal effort, such that the vocal effort is positively correlated with the mixed aperiodic parameters, includes: determining a first weight and a second weight according to the vocal effort, wherein the vocal effort is positively correlated with the first weight and negatively correlated with the second weight; and performing a weighted summation of the white noise and aperiodic parameters according to the first weight and the second weight.

[0090] Third Embodiment In the above embodiments, a method for generating pseudo-whispers is provided. Correspondingly, this application also provides a method for constructing a whisper recognition model. This method corresponds to the embodiments of the above methods, so it is described simply. For relevant details, please refer to the description of the method embodiment one. The method embodiments described below are merely illustrative.

[0091] Please refer to Figure 1 This is a flowchart of the whisper recognition model construction method provided in this embodiment. The whisper recognition model construction method of this embodiment includes the following steps: Step S201: Acquire multiple normal speech data.

[0092] Step S203: For the normal speech frames in the normal speech data, obtain the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames; Step S205: Based on the force-related acoustic features and the clarity-related acoustic features, obtain the vocal effort of the normal speech frame, wherein the vocal effort is positively correlated with force and clarity; Step S207: Adjust the acoustic parameters of the multi-domain representation of the normal speech frame according to the vocal effort; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain. Step S209: Based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, synthesize a pseudo-whisper speech frame corresponding to the normal speech frame, and form pseudo-whisper speech data corresponding to the normal speech data with the multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data. Step S211: Construct a whisper recognition model based on multiple pseudo-whisper speech data corresponding to multiple normal speech data.

[0093] In this embodiment, multiple generated pseudo-whisper speech data are used as training samples, and a whisper recognition model is learned from the training samples.

[0094] As can be seen from the above embodiments, the whisper recognition model construction method provided in this application involves: acquiring multiple normal speech data; acquiring the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames in the normal speech data; acquiring the vocal effort of the normal speech frames based on the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort is positively correlated with force and intelligibility; adjusting the acoustic parameters of the multi-domain representation of the normal speech frames based on the vocal effort; wherein the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; synthesizing pseudo-whisper speech frames corresponding to the normal speech frames based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, and forming pseudo-whisper speech data corresponding to the normal speech data with the multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data; and constructing a whisper recognition model based on the multiple pseudo-whisper speech data corresponding to the multiple normal speech data. This processing method generates pseudo-whispers with high naturalness and strong interpretability. The whisper recognition model is trained based on the pseudo-whispers with high naturalness, thus effectively improving the accuracy of the whisper recognition model.

[0095] Fourth embodiment In the above embodiments, a method for constructing a whisper recognition model is provided. Correspondingly, this application also provides a device for constructing a whisper recognition model. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0096] This application also provides a whispering recognition model construction device, comprising: a normal speech acquisition unit for acquiring multiple normal speech data; an acoustic feature extraction unit for acquiring force-related acoustic features and intelligibility-related acoustic features of normal speech frames in the normal speech data; a vocal effort calculation unit for acquiring the vocal effort of the normal speech frames based on the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort is positively correlated with force and intelligibility; a cross-domain parameter joint adjustment unit for adjusting the acoustic parameters of the multi-domain representation of the normal speech frames based on the vocal effort; wherein the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; a pseudo-whispering synthesis unit for synthesizing pseudo-whispering speech frames corresponding to the normal speech frames based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, and forming pseudo-whispering speech data with the multiple pseudo-whispering speech frames corresponding to the multiple normal speech frames in the normal speech data; and a model construction unit for constructing a whispering recognition model based on the multiple pseudo-whispering speech data corresponding to the multiple normal speech data.

[0097] Fourth embodiment In the above embodiments, a method for generating pseudo-whispers is provided. Correspondingly, this application also provides a method for whisper recognition. This method corresponds to the embodiments of the above methods, so it is described simply. For relevant details, please refer to the description of the first method embodiment. The method embodiments described below are merely illustrative.

[0098] The whisper recognition method of this embodiment includes the following steps: Step S301: Collect whispered signals.

[0099] Step S303: Obtain the whisper content based on the whisper signal using the whisper recognition model; The whisper recognition model is constructed as follows: Multiple normal speech data are acquired; for normal speech frames in the normal speech data, the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames are acquired; based on the force-related acoustic features and the intelligibility-related acoustic features, the vocal effort of the normal speech frames is acquired, where the vocal effort is positively correlated with force and intelligibility; based on the vocal effort, the acoustic parameters of the multi-domain representation of the normal speech frames are adjusted; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, pseudo-whisper speech frames corresponding to the normal speech frames are synthesized, and multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data are combined to form pseudo-whisper speech data corresponding to the normal speech data; based on the multiple pseudo-whisper speech data corresponding to the multiple normal speech data, a whisper recognition model is constructed.

[0100] As can be seen from the above embodiments, the whisper recognition method provided in this application collects whisper signals; and obtains whisper content based on the whisper signals using a whisper recognition model. The whisper recognition model is constructed as follows: acquiring multiple normal speech data; acquiring force-related acoustic features and intelligibility-related acoustic features of normal speech frames in the normal speech data; acquiring the vocal effort of the normal speech frames based on the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort is positively correlated with force and intelligibility; adjusting the acoustic parameters of the multi-domain representation of the normal speech frames based on the vocal effort; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; synthesizing pseudo-whisper speech frames corresponding to the normal speech frames based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, and forming pseudo-whisper speech data corresponding to the normal speech data with the multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data; and constructing a whisper recognition model based on the multiple pseudo-whisper speech data corresponding to the multiple normal speech data. This processing method generates pseudo-whispers with high naturalness and strong interpretability. A high-accuracy whisper recognition model is trained based on the high-naturalness pseudo-whispers corpus, and whisper recognition is performed based on the high-accuracy whisper recognition model. Therefore, the whisper recognition accuracy can be effectively improved, thereby improving the user experience.

[0101] Sixth Embodiment In the above embodiments, a whisper recognition method is provided. Correspondingly, this application also provides a whisper recognition device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0102] This application also provides a whisper recognition device, comprising: a whisper acquisition unit for acquiring whisper signals; and a whisper recognition unit for obtaining whisper content based on the whisper signals using a whisper recognition model.

[0103] The whisper recognition model is constructed as follows: Multiple normal speech data are acquired; for normal speech frames in the normal speech data, the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames are acquired; based on the force-related acoustic features and the intelligibility-related acoustic features, the vocal effort of the normal speech frames is acquired, where the vocal effort is positively correlated with force and intelligibility; based on the vocal effort, the acoustic parameters of the multi-domain representation of the normal speech frames are adjusted; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, pseudo-whisper speech frames corresponding to the normal speech frames are synthesized, and multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data are combined to form pseudo-whisper speech data corresponding to the normal speech data; based on the multiple pseudo-whisper speech data corresponding to the multiple normal speech data, a whisper recognition model is constructed.

[0104] Seventh Embodiment In the above embodiments, a method for generating pseudo-whispers is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0105] The electronic device of this embodiment includes: a memory and a processor; the memory is used to store a program that implements any of the above methods, and the device is powered on and runs the program of any of the above methods through the processor.

[0106] Memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0107] In specific implementations, the electronic device may further include one or more of the following components: a power supply component, an input / output (I / O) interface, and a communication component. The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device. The I / O interface provides an interface between the processor 503 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and user devices (such as smartphones, tablets, etc.).

[0108] Eighth embodiment This application also provides a computer-readable storage medium. Since the embodiments of the computer-readable storage medium are substantially similar to the method embodiments, the description is relatively simple; relevant details can be found in the description of the method embodiments. The computer-readable storage medium embodiments described below are merely illustrative.

[0109] In this embodiment, a non-transitory computer-readable storage medium including instructions is provided, such as a memory including instructions, which can be executed by a processor of an electronic device to perform any of the methods provided in this disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0110] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0111] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0112] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

[0113] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0114] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0115] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0116] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A method for generating pseudo-whispers, characterized in that, include: Acquire normal voice data; For normal speech frames in normal speech data, obtain the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames; Based on the force-related acoustic features and the clarity-related acoustic features, the vocal effort of the normal speech frame is obtained, and the vocal effort is positively correlated with force and clarity. Based on the vocal effort, the acoustic parameters of the multi-domain representation of the normal speech frame are adjusted; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain. Based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, a pseudo-whisper speech frame corresponding to the normal speech frame is synthesized, and the pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data form pseudo-whisper speech data.

2. The method according to claim 1, characterized in that, The step of obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features includes: Based on the force weight and the force-related acoustic features, and the clarity weight and the clarity-related acoustic features, a weighted calculation is performed, and the weighted calculation value is used as the vocal effort.

3. The method according to claim 1, characterized in that, The step of obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features includes: Acquire simulated perception data of the force-related acoustic features of the human ear; The vocal effort is obtained based on the simulated perception data and the clarity-related acoustic features.

4. The method according to claim 1, characterized in that, The step of obtaining the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features includes: The force-related acoustic features and the sharpness-related acoustic features are normalized. The vocal effort is obtained based on normalized force-related acoustic features and normalized intelligibility-related acoustic features. The vocal effort was normalized.

5. The method according to claim 1, characterized in that, The acoustic parameters in the time domain include loudness-related parameters, which serve as the first parameter; Based on the vocal effort, the acoustic parameters in the time domain are adjusted, including: Obtain the peak-to-peak value of the first parameter of the normal speech frame, and obtain the first parameter corresponding to the low sound or the first parameter corresponding to the high sound; The fluctuation value of the first parameter is obtained based on the peak-to-peak value of the first parameter and the vocal effort. The adjusted first parameter is obtained based on the first parameter corresponding to the low sound or the first parameter corresponding to the high sound, and the fluctuation value.

6. The method according to claim 5, characterized in that, Adjusting the acoustic parameters in the time domain according to the vocal effort also includes: Obtain the first parameter, contrast. Based on the vocal effort and the first parameter contrast, the vocal effort affected by the first parameter contrast is obtained; The step of obtaining the fluctuation value of the first parameter based on the peak-to-peak value of the first parameter and the vocal effort includes: The fluctuation value is obtained based on the peak-to-peak value of the first parameter and the vocal effort affected by the contrast of the first parameter.

7. The method according to claim 1, characterized in that, The acoustic parameters in the frequency domain include the smooth bandwidth of the spectral envelope; Based on the vocal effort, the acoustic parameters in the frequency domain are adjusted, including: Based on the vocal effort, obtain the target reduction in smooth bandwidth; The smoothing bandwidth is adjusted according to the target reduction amount.

8. The method according to claim 7, characterized in that, The step of adjusting the acoustic parameters in the frequency domain according to the vocal effort further includes: Obtain the spectral centroid; Based on the spectral centroid, the target increment of the smoothing bandwidth is obtained; the target increment corresponding to voiceless consonants is less than the target increment corresponding to vowels; The smoothing bandwidth is adjusted according to the target increment.

9. The method according to claim 1, characterized in that, The acoustic parameters of the excitation domain include non-periodic parameters; Based on the vocal effort, the acoustic parameters of the excitation domain are adjusted, including: Based on the vocal effort, white noise and aperiodic parameters are mixed such that the vocal effort is positively correlated with the mixed aperiodic parameters.

10. The method according to claim 9, characterized in that, The step of mixing white noise and aperiodic parameters based on the vocal effort, such that the vocal effort is positively correlated with the mixed aperiodic parameters, includes: Based on vocal effort, a first weight and a second weight are determined. Vocal effort is positively correlated with the first weight and negatively correlated with the second weight. Based on the first weight and the second weight, the white noise and aperiodic parameters are weighted and summed.

11. A method for constructing a whisper recognition model, characterized in that, include: Acquire multiple normal voice data; For the normal speech frames in the normal speech data, obtain the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames; Based on the force-related acoustic features and the clarity-related acoustic features, the vocal effort of the normal speech frame is obtained, and the vocal effort is positively correlated with force and clarity. Based on the vocal effort, the acoustic parameters of the multi-domain representation of the normal speech frame are adjusted; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain. Based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, a pseudo-whisper speech frame corresponding to the normal speech frame is synthesized, and multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data form pseudo-whisper speech data corresponding to the normal speech data. A whisper recognition model is constructed based on multiple pseudo-whisper speech data corresponding to multiple normal speech data.

12. A whisper recognition method, characterized in that, include: Collect whispered signals; The whisper recognition model is used to obtain the whisper content based on the whisper signal. The whisper recognition model is constructed as follows: Multiple normal speech data are acquired; for normal speech frames in the normal speech data, the force-related acoustic features and intelligibility-related acoustic features of the normal speech frames are acquired; based on the force-related acoustic features and the intelligibility-related acoustic features, the vocal effort of the normal speech frames is acquired, where the vocal effort is positively correlated with force and intelligibility; based on the vocal effort, the acoustic parameters of the multi-domain representation of the normal speech frames are adjusted; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain; based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frames, pseudo-whisper speech frames corresponding to the normal speech frames are synthesized, and multiple pseudo-whisper speech frames corresponding to the multiple normal speech frames in the normal speech data are combined to form pseudo-whisper speech data corresponding to the normal speech data; based on the multiple pseudo-whisper speech data corresponding to the multiple normal speech data, a whisper recognition model is constructed.

13. A pseudo-whisper generation device, characterized in that, include: Normal speech acquisition unit, used to acquire normal speech data; An acoustic feature extraction unit is used to obtain force-related acoustic features and intelligibility-related acoustic features of normal speech frames in normal speech data. The vocal effort calculation unit is used to obtain the vocal effort of the normal speech frame based on the force-related acoustic features and the intelligibility-related acoustic features, wherein the vocal effort is positively correlated with force and intelligibility. A cross-domain parameter joint adjustment unit is used to adjust the acoustic parameters of the multi-domain representation of the normal speech frame according to the vocal effort; the multi-domain representation includes at least two of the following: time domain, frequency domain, and excitation domain. The pseudo-whisper synthesis unit is used to synthesize a pseudo-whisper speech frame corresponding to the normal speech frame based on the adjusted acoustic parameters of the multi-domain representation of the normal speech frame, and form pseudo-whisper speech data with multiple pseudo-whisper speech frames corresponding to multiple normal speech frames in the normal speech data.

14. An electronic device, characterized in that, include: processor; as well as A memory for storing a program for implementing the method according to any one of claims 1 to 12, wherein the device is powered on and the program for running the method is executed by the processor.