Voice changing processing methods, storage media, chips and electronic devices

By normalizing and correcting the frequency domain information of the voice changer, the problem of inaccurate simulation of formant frequency differences in existing voice changer algorithms is solved, and more natural and accurate voice changer speech signal generation is achieved.

CN115641858BActive Publication Date: 2026-04-21SHENZHEN BLUETRUM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN BLUETRUM TECH CO LTD
Filing Date
2022-10-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing voice-changing algorithms cannot effectively simulate the formant frequency differences between male and female voices, resulting in voice-changing speech signals that are not natural enough.

Method used

By acquiring the frequency domain information of the altered voice, normalization and correction processes are performed, including determining the normalized envelope coefficients and compensation vectors of the formants, and combining them with phase information to generate a time-domain speech signal.

Benefits of technology

It improves the naturalness and coherence of the voice-changing speech signal and enhances the accuracy of the voice-changing speech signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641858B_ABST
    Figure CN115641858B_ABST
Patent Text Reader

Abstract

This invention relates to the field of voice-changing processing technology, and discloses a voice-changing processing method, storage medium, chip, and electronic device. The method includes: acquiring voice-changing frequency domain information; normalizing the voice-changing frequency domain information to obtain normalized frequency domain information, the normalized frequency domain information including formant-normalized frequency domain information; correcting the normalized frequency domain information to obtain corrected frequency domain information, the corrected frequency domain information including formant-corrected frequency domain information; and generating time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information. This embodiment can not only perform voice-changing processing on the entire speech signal, but also normalize the formants and correct them, thereby obtaining new formants, making the voice-changing speech signal more accurate and natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice changing technology, specifically to a voice changing method, storage medium, chip, and electronic device. Background Technology

[0002] Existing voice-changing algorithms typically alter the fundamental frequency of the speech signal to achieve the desired voice change. For example, increasing the fundamental frequency can change a male voice to a female voice, or decreasing it can change a female voice to a male voice. However, due to the different structures of the male and female vocal tracts, in addition to the difference in fundamental frequency, the frequencies of formants also differ. Therefore, the voice-changing signals obtained by existing algorithms are not natural enough. Summary of the Invention

[0003] One objective of this invention is to provide a voice-changing processing method, storage medium, chip, and electronic device, aiming to improve the problem that existing voice-changing processing methods do not process natural-sounding speech signals.

[0004] In a first aspect, embodiments of the present invention provide a voice-changing processing method, comprising:

[0005] Acquire voice-changing frequency domain information, which is information obtained by performing a Fourier transform on the voice-changing speech signal, the voice-changing speech signal is a speech signal processed by a voice-changing algorithm, and the voice-changing frequency domain information includes frequency domain information corresponding to formants.

[0006] The frequency domain information of the altered sound is normalized to obtain normalized frequency domain information, which includes the frequency domain information of the formant after normalization.

[0007] The normalized frequency domain information is corrected to obtain corrected frequency domain information, which includes the frequency domain information of the formant after correction.

[0008] Time-domain speech information is generated based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information.

[0009] Optionally, the normalization process for the variable frequency domain information to obtain normalized frequency domain information includes:

[0010] Determine the normalized envelope coefficients of the voice-changing speech signal in the frequency domain;

[0011] The normalized frequency domain information is normalized based on the normalized envelope coefficients to obtain normalized frequency domain information.

[0012] Optionally, determining the normalized envelope coefficients of the voice-changing speech signal in the frequency domain includes:

[0013] According to the linear prediction algorithm, the p-th order linear prediction coefficients of the voice-changing speech signal in the frequency domain are calculated, where p is a positive integer;

[0014] The normalized envelope coefficients are determined based on the p-order linear prediction coefficients.

[0015] Optionally, determining the normalized envelope coefficients based on the p-order linear prediction coefficients includes:

[0016] The length of the p-order linear prediction coefficients is extended to the length of the target window function to obtain extended linear prediction coefficients, wherein the target window function is a window function that participates in the calculation of the variable sound frequency domain information;

[0017] Perform a Fourier transform on the extended linear prediction coefficients to obtain the Fourier information of the coefficients;

[0018] The normalized envelope coefficients are obtained by taking the modulus of the Fourier information of the coefficients.

[0019] Optionally, the step of normalizing the voice-changing frequency domain information based on the normalized envelope coefficients to obtain normalized frequency domain information includes:

[0020] Calculate the normalization factor based on the normalized envelope coefficients;

[0021] The normalized frequency domain information is determined based on the normalization factor and the variable sound frequency domain information.

[0022] Optionally, the step of correcting the normalized frequency domain information to obtain corrected frequency domain information includes:

[0023] The compensation vector of the resonance peak is calculated based on the normalized envelope coefficient.

[0024] Based on the normalized frequency domain information and the compensation vector, the corrected frequency domain information is determined.

[0025] Optionally, calculating the compensation vector of the resonance peak based on the normalized envelope coefficient includes:

[0026] According to the preset resonance peak correction ratio, the normalized envelope coefficient is corrected to obtain the corrected envelope coefficient, wherein the number of coefficients of the corrected envelope coefficient is equal to the number of coefficients of the normalized envelope coefficient.

[0027] The compensation vector of the resonance peak is determined based on the modified envelope coefficient.

[0028] Optionally, the step of correcting the normalized envelope coefficient according to a preset resonance correction ratio to obtain the corrected envelope coefficient includes:

[0029] Based on the preset resonance correction ratio, the number of coefficients in the normalized envelope coefficient is reduced or increased to obtain the scaled envelope coefficient.

[0030] The scaled envelope coefficients are linearly interpolated using a linear interpolation algorithm to obtain the corrected envelope coefficients.

[0031] Optionally, determining the compensation vector of the resonance peak based on the modified envelope coefficient includes: filtering the modified envelope coefficient to obtain the compensation vector of the resonance peak.

[0032] Optionally, determining the corrected frequency domain information based on the normalized frequency domain information and the compensation vector includes:

[0033] The normalized frequency domain information and the compensation vector are multiplied together to obtain the corrected frequency domain information.

[0034] Optionally, generating time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information includes:

[0035] Perform an inverse Fourier transform on the corrected frequency domain information and the phase information to obtain preliminary time domain information;

[0036] The preliminary time-domain information is smoothed to obtain time-domain speech information.

[0037] In a second aspect, embodiments of the present invention provide a storage medium characterized in that it stores computer-executable instructions, which are used to cause an electronic device to perform the above-described voice-changing processing method.

[0038] In a third aspect, embodiments of the present invention provide a chip, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein...

[0039] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the voice-changing processing method described above.

[0040] In a fourth aspect, embodiments of the present invention provide an electronic device, comprising:

[0041] At least one processor; and,

[0042] A memory communicatively connected to the at least one processor; wherein,

[0043] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the voice-changing processing method described above.

[0044] In the voice-changing processing method provided in this embodiment of the invention, voice-changing frequency domain information is obtained. This information is obtained by performing a Fourier transform on the voice-changing speech signal. The voice-changing speech signal is a speech signal processed by a voice-changing algorithm. The voice-changing frequency domain information includes the frequency domain information corresponding to the formants. Normalization processing is performed on the voice-changing frequency domain information to obtain normalized frequency domain information, which includes the normalized frequency domain information of the formants. Correction processing is then performed on the normalized frequency domain information to obtain corrected frequency domain information, which includes the corrected frequency domain information of the formants. Based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information, time-domain speech information is generated. This embodiment not only performs voice-changing processing on the speech signal as a whole but also normalizes the formants and corrects them, thereby obtaining new formants and making the voice-changing speech signal more accurate and natural. Attached Figure Description

[0045] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0046] Figure 1 A schematic flowchart of a voice-changing processing method provided in an embodiment of the present invention;

[0047] Figure 2 for Figure 1 The flowchart of S12 is shown below;

[0048] Figure 3 for Figure 1 The flowchart of S13 is shown below;

[0049] Figure 4 This is a schematic diagram of the structure of a voice-changing processing device provided in an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of the circuit structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0052] It should be noted that, unless otherwise specified, the various features in the embodiments of this invention can be combined with each other, all of which are within the protection scope of this invention. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this invention do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0053] This invention provides a voice-changing processing method. Please refer to [link / reference]. Figure 1 The voice changing process includes the following steps:

[0054] S11: Obtain the voice-changing frequency domain information. The voice-changing frequency domain information is the information obtained by performing a Fourier transform on the voice-changing speech signal. The voice-changing speech signal is the speech signal processed by the voice-changing algorithm. The voice-changing frequency domain information includes the frequency domain information corresponding to the formants.

[0055] In this step, the speech signal is acquired and processed using a voice-changing algorithm to obtain voice-changing speech data. Based on the length of the filter, the voice-changing speech signal x(n) of the nth frame is extracted from the voice-changing speech data. The voice-changing algorithm includes the WSOLA algorithm (Waveform Similarity Overlap-Add) or a vocoder and other pitch-changing constant-speed algorithms. x(n) = [x(i), x(i-1), ..., x(iM-1)], where M is the length of the filter and i represents the i-th sampling point.

[0056] This embodiment applies a windowing function to the voice-changing speech signal to obtain windowed speech information. Then, a Fourier transform is performed on the windowed speech information to obtain the voice-changing frequency domain information. For example, this embodiment performs a Fourier transform on the voice-changing speech signal according to Equation 1 to obtain the voice-changing frequency domain information:

[0057] X(n) = fft([x(n-1); x(n)·win]) Equation 1

[0058] Where X(n) represents the frequency domain information of the nth frame, win represents the target window function, which can be the Hanning window in this paper. The length of the Hanning window is 2*M, and fft represents the Fourier transform. In some embodiments, the target window function is the Hanning window function.

[0059] S12: Normalize the frequency domain information of the voice changer to obtain normalized frequency domain information, which includes the frequency domain information after normalization of the formant.

[0060] In this step, the normalized frequency domain information is the frequency domain information obtained after normalizing the voice-changing frequency domain information. Formants are regions in the speech signal spectrum where energy is relatively concentrated.

[0061] S13: Correct the normalized frequency domain information to obtain corrected frequency domain information, which includes the corrected frequency domain information of the formants.

[0062] In this step, the corrected frequency domain information is the frequency domain information after the normalized frequency domain information has been corrected.

[0063] S14: Generate time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information.

[0064] In this step, the time-domain speech information is the target frequency-domain information converted from the frequency domain to the time domain. The target frequency-domain information is obtained from the phase information of the jointly corrected frequency-domain information and the voice-changing frequency-domain information. For example, according to Equations 2 and 3, the time-domain speech information can be obtained in this embodiment as follows:

[0065]

[0066]

[0067] Where Y1(n) represents the corrected frequency domain information of the nth frame, and Y2(n) represents the target frequency domain information of the nth frame. X(n) represents the phase information of the frequency domain information of the nth frame's altered audio, and X(n) represents the frequency domain information of the nth frame's altered audio.

[0068] In some embodiments, this embodiment performs an inverse Fourier transform on the target frequency domain information to obtain time-domain speech information. For example, according to Equation 4, this embodiment can obtain the time-domain speech information as follows:

[0069] y(n) = ifft(Y2(n)) * win (Formula 4)

[0070] Where y(n) represents time-domain speech information, and ifft represents the inverse Fourier transform.

[0071] In some embodiments, generating time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information includes the following steps:

[0072] S141: Perform an inverse Fourier transform on the corrected frequency domain information and phase information to obtain preliminary time domain information.

[0073] S142: Smooth the initial time-domain information to obtain time-domain speech information.

[0074] In S141, as mentioned above, this embodiment can use Equation 4 to perform an inverse Fourier transform on the corrected frequency domain information and phase information in order to obtain preliminary time domain information.

[0075] In S142, in order to improve the naturalness and smoothness of the voice-changing speech signal, this embodiment can perform smoothing processing on the preliminary time-domain information to obtain time-domain speech information. For example, according to Equations 5 and 6, this embodiment can obtain the time-domain speech information as follows:

[0076] out(n) = y(1:M) + out_last(n) (Formula 5)

[0077] out_last(n) = y(M+1:2*M) (Formula 6)

[0078] Where out(n) is the time-domain speech information, and out_last(n) is the second half of the previous frame data.

[0079] Through Equations 5 and 6, this embodiment can effectively smooth and recover the initial time-domain information, thereby enabling the natural and smooth output of time-domain speech information.

[0080] In some embodiments, please refer to Figure 2 The normalization process for the voice-changing frequency domain information to obtain normalized frequency domain information includes the following steps:

[0081] S121: Determine the normalized envelope coefficients of the voice-changing speech signal in the frequency domain.

[0082] S122: Based on the normalized envelope coefficients, the frequency domain information of the voice changer is normalized to obtain the normalized frequency domain information.

[0083] In S121, the normalized envelope coefficients are the set of amplitudes of each resonance peak in the frequency domain.

[0084] In S122, in some embodiments, this embodiment calculates the p-th order linear prediction coefficients of the voice-changing speech signal in the frequency domain according to the linear prediction algorithm, where p is a positive integer, and determines the normalized envelope coefficients based on the p-th order linear prediction coefficients.

[0085] In this embodiment, the Levinson-Dubin formula autocorrelation method is used to solve for the linear prediction coefficients. The linear prediction system can be represented by Equation 7:

[0086]

[0087] in, Let x(n) be an estimated value. a is obtained by linear combination of the past p values. iFor a linear prediction function, the transfer function of a p-order linear predictor is as follows:

[0088]

[0089] The autocorrelation function r(j) is shown in Equation 9:

[0090]

[0091] The minimum mean square error E can be written as:

[0092] The recursive derivation is performed step by step using a recursive solution:

[0093] ① When i = 0, E = r(0), a0 = 1

[0094] ② Corresponding to the i-th recursion (i = 1, 2, 3, ..., p):

[0095]

[0096] a j (i) =k i

[0097] a j (i) =a j (i-1) -k i a i-j (i-1)

[0098] E i =(1-k) i 2 E i-1

[0099] The superscript in the parentheses above indicates the order of the predictor. By recursively solving for i = 1, 2, ..., p, we obtain:

[0100] a i =a j (p) 1≤j≤p

[0101] In this embodiment, the linear prediction coefficient a of order p is retained. i .

[0102] In some embodiments, determining the normalized envelope coefficients based on the p-order linear prediction coefficients includes the following steps:

[0103] S1221: Extend the length of the p-order linear prediction coefficients to the length of the target window function to obtain the extended linear prediction coefficients, where the target window function is the window function used to calculate the variable sound frequency domain information.

[0104] S1222: Perform a Fourier transform on the extended linear prediction coefficients to obtain the Fourier information of the coefficients.

[0105] S1223: Calculate the normalized envelope coefficients by taking the modulus of the Fourier information of the coefficients.

[0106] In S1221, this embodiment pads the p-th order linear prediction coefficients with 0s to extend the length of the p-th order linear prediction coefficients to the length of the target window function, that is, to extend it to a frame of data of size 2*M, where 2*M is the length of the target window function. For example, in this embodiment, padding the p-th order linear prediction coefficients with 0s yields ar(n) = [a1, a2, a3, ..., a p ,0,0,0,......,0].

[0107] In S1222, this embodiment performs a Fourier transform on the extended linear prediction coefficients according to Equation 10 to obtain the Fourier information of the coefficients, as follows:

[0108] F = fft(ar(n)) (Equation 10)

[0109] Where F represents the Fourier information of the coefficients.

[0110] In S1223, this embodiment calculates the normalized envelope coefficients by taking the modulus of the Fourier information of the coefficients according to Equation 11, as follows:

[0111] AR(n) = abs(F, 2M) Equation 11

[0112] Where AR(n) is the normalized envelope coefficient of the nth frame, and abs represents the modulus.

[0113] In some embodiments, normalizing the frequency domain information of the voice changer according to the normalized envelope coefficients to obtain the normalized frequency domain information includes the following steps:

[0114] S1221: Calculate the normalization factor based on the normalized envelope coefficient.

[0115] S1222: Determine the normalized frequency domain information based on the normalization factor and the variable frequency domain information.

[0116] In S1221, in some embodiments, this embodiment calculates the reciprocal of the normalized envelope coefficient and uses the reciprocal of the normalized envelope coefficient as the normalization factor. In some embodiments, this embodiment calculates the normalization factor based on the normalized envelope coefficient and the division protection factor. For example, this embodiment calculates the normalization factor according to Equation Twelve, as follows:

[0117] α(n) = 1 / (AR(n) + δ) (Equation 12)

[0118] Where α(n) is the normalization factor and δ is the division protection factor. δ can be customized by the designer according to business needs, for example, δ is 0.00001.

[0119] In S1222, in some embodiments, this embodiment calculates the modulus of the variable voice frequency domain information, and then divides the modulus of the variable voice frequency domain information by a normalization factor to obtain the normalized frequency domain information. For example, this embodiment determines the normalized frequency domain information according to formula thirteen, as follows:

[0120] Y(n) = abs(X(n)) / α(n) Equation Thirteen

[0121] Where Y(n) represents the normalized frequency domain information.

[0122] In some embodiments, please refer to Figure 3 The process of correcting the normalized frequency domain information to obtain the corrected frequency domain information includes the following steps:

[0123] S131: Calculate the compensation vector of the resonance peak based on the normalized envelope coefficient.

[0124] S132: Determine the corrected frequency domain information based on the normalized frequency domain information and the compensation vector.

[0125] In S131, the compensation vector is used to correct the amplitude of the resonance peak. In some embodiments, calculating the compensation vector of the resonance peak based on the normalized envelope coefficient includes: correcting the normalized envelope coefficient to obtain a corrected envelope coefficient according to a preset resonance peak correction ratio, wherein the number of coefficients in the corrected envelope coefficient is equal to the number of coefficients in the normalized envelope coefficient; and determining the compensation vector of the resonance peak based on the corrected envelope coefficient. The preset resonance peak correction ratio can be customized by the designer based on engineering experience.

[0126] In some embodiments, correcting the normalized envelope coefficients according to a preset formant correction ratio to obtain corrected envelope coefficients includes: reducing or increasing the number of coefficients in the normalized envelope coefficients according to the preset formant correction ratio to obtain scaled envelope coefficients; and performing linear interpolation on the scaled envelope coefficients according to a linear interpolation algorithm to obtain corrected envelope coefficients.

[0127] For example, the normalized envelope coefficient AR(n) = [1,2,3,.....,2*M], with a preset formant correction ratio rate = 2. Therefore, in this embodiment, the number of coefficients in the normalized envelope coefficient is reduced by half, resulting in the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M]. In this embodiment, based on a linear interpolation algorithm, the uninterpolated normalized envelope coefficient AR(n) = [1,2,3,.....,2*M] is linearly inserted into the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M], thus expanding the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M] into a frame of data of size 2*M, resulting in the corrected envelope coefficient AR_interp1(n) = [1,3,5,.....,M,α(M+1),.....,α(2*M)].

[0128] It is understandable that when the preset resonant correction ratio rate = 0.5, this embodiment amplifies the number of coefficients in the normalized envelope coefficients, resulting in the scaled envelope coefficients AR_interp(n) = [1, 1.5, 2, 2.5, 3, 3.5, ..., 2*M].

[0129] In some embodiments, determining the compensation vector of the resonance peak based on the modified envelope coefficient includes: directly using the modified envelope coefficient as the compensation vector of the resonance peak.

[0130] In some embodiments, determining the compensation vector of the formant based on the modified envelope coefficient includes: filtering the modified envelope coefficient to obtain the compensation vector of the formant. Thus, this embodiment can obtain a compensation vector with a relatively high degree of smoothness and fewer spikes, which is beneficial for obtaining a relatively smooth modified frequency domain information based on the normalized frequency domain information and the compensation vector. Furthermore, based on the phase information of the modified frequency domain information and the voice-changing frequency domain information, a relatively smooth time-domain speech information can be generated, which makes the voice changing more natural and coherent.

[0131] In some embodiments, the filtering process includes an initial filtering process. Filtering the modified envelope coefficients to obtain the compensation vector of the resonance peak includes the following steps: performing an initial filtering process on the modified envelope coefficients to obtain an initial vector of the resonance peak, and determining the compensation vector of the resonance peak based on the initial vector of the resonance peak.

[0132] In some embodiments, the initial filtering of the modified envelope coefficients to obtain the initial vector of the formant includes the following steps: obtaining the modified envelope coefficients of the previous frame, the modified envelope coefficients of the current frame, and the modified envelope coefficients of the next frame; calculating the first coefficient weighting result based on the modified envelope coefficients of the previous frame and the first preset weight; calculating the second coefficient weighting result based on the modified envelope coefficients of the current frame and the second preset weight; calculating the third coefficient weighting result based on the modified envelope coefficients of the next frame and the third preset weight; and adding the first coefficient weighting result, the second coefficient weighting result, and the third coefficient weighting result to obtain the initial vector of the formant.

[0133] In some embodiments, the second preset weight is greater than the first preset weight and the third preset weight, thereby increasing the influence of the current frame's corrected envelope coefficients on determining the initial vector. In some embodiments, the first preset weight is greater than the third preset weight, thereby increasing the influence of the previous frame's corrected envelope coefficients on determining the initial vector.

[0134] For example, in this embodiment, the initial vector of the resonance peak is determined according to Equation Fourteen, as follows:

[0135]

[0136] Where AR_interp1(i-1) is the corrected envelope coefficient of the previous frame, η1 is the first preset weight, and ψ1 is the first weighted result. AR_interp1(i) is the corrected envelope coefficient of the current frame, η2 is the second preset weight, and ψ2 is the second weighted result. AR_interp1(i+1) is the corrected envelope coefficient of the next frame, η3 is the third preset weight, and ψ3 is the third weighted result. AR_interp2(i) is the initial vector. i = 2, 3, 4, ..., 2*M-1. In some embodiments, η1 is 0.2, η2 is 1, and η3 is 0.01.

[0137] In some embodiments, determining the compensation vector of a resonance peak based on its initial vector includes: directly using the initial vector of the resonance peak as the compensation vector of the resonance peak.

[0138] In some embodiments, the filtering process includes median filtering, and determining the compensation vector of the resonance peak based on the initial vector of the resonance peak includes: performing median filtering on the initial vector of the resonance peak to obtain the compensation vector of the resonance peak.

[0139] For example, in this embodiment, according to Equation 15, the initial vector of the resonance peak is subjected to median filtering to obtain the compensation vector of the resonance peak, as follows:

[0140] AR_int erp3(n) = medfilt(AR_int erp2(n)) (Equation 15)

[0141] This embodiment, through initial filtering and median filtering, can reliably and effectively filter out the spikes in the compensation vector, making the compensation vector smoother.

[0142] In some embodiments, determining the corrected frequency domain information based on the normalized frequency domain information and the compensation vector includes multiplying the normalized frequency domain information and the compensation vector to obtain the corrected frequency domain information. For example, in this embodiment, the corrected frequency domain information is obtained according to Equation Sixteen, thereby enabling the refitting and generation of formants, as follows:

[0143] Y1(n) = Y(n) * AR_int erp3(n) (Equation 16)

[0144] After obtaining the corrected frequency domain information, this embodiment can combine Equations 2, 3, 4, 5, and 6 mentioned above to output time-domain speech information. As mentioned earlier, the voice-changing time-domain speech information provided in this embodiment is relatively natural and smooth.

[0145] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.

[0146] As another aspect of the embodiments of the present invention, the present invention provides a voice-changing processing device. The voice-changing processing device can be a software module, which includes several instructions stored in a memory. A processor can access the memory, call the instructions, and execute them to complete the voice-changing processing methods described in the above embodiments.

[0147] In some embodiments, the voice-changing processing device can also be constructed from hardware components. For example, the voice-changing processing device can be constructed from one or more chips, which can work in coordination to complete the voice-changing processing methods described in the various embodiments above. As another example, the voice-changing processing device can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0148] Please see Figure 4 The voice-changing processing device 400 includes a voice-changing conversion module 41, a normalization processing module 42, a correction processing module 43, and a time-domain speech generation module 44.

[0149] The voice-changing conversion module 41 is used to acquire voice-changing frequency domain information, which is information obtained by performing a Fourier transform on the voice-changing speech signal. The voice-changing speech signal is a speech signal processed by a voice-changing algorithm, and the voice-changing frequency domain information includes the frequency domain information corresponding to the formants. The normalization processing module 42 is used to normalize the voice-changing frequency domain information to obtain normalized frequency domain information, which includes the normalized frequency domain information of the formants. The correction processing module 43 is used to correct the normalized frequency domain information to obtain corrected frequency domain information, which includes the corrected frequency domain information of the formants. The time-domain speech generation module 44 is used to generate time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information.

[0150] This embodiment can not only perform voice-changing processing on the entire speech signal, but also normalize the formants and make corrections on the normalized formants, thereby obtaining new formants, which helps to make the voice-changing speech signal more accurate and natural.

[0151] In some embodiments, the normalization processing module 42 is specifically used to: determine the normalized envelope coefficients of the voice-changing speech signal in the frequency domain, and perform normalization processing on the voice-changing frequency domain information according to the normalized envelope coefficients to obtain normalized frequency domain information.

[0152] In some embodiments, the normalization processing module 42 is further specifically used to: calculate the p-th order linear prediction coefficients of the voice-changing speech signal in the frequency domain according to the linear prediction algorithm, where p is a positive integer, and determine the normalized envelope coefficients based on the p-th order linear prediction coefficients.

[0153] In some embodiments, the normalization processing module 42 is further specifically used to: extend the length of the p-order linear prediction coefficients to the length of the target window function to obtain extended linear prediction coefficients, wherein the target window function is a window function that participates in the calculation of the variable sound frequency domain information; perform Fourier transform on the extended linear prediction coefficients to obtain coefficient Fourier information; and perform modulus calculation on the coefficient Fourier information to obtain normalized envelope coefficients.

[0154] In some embodiments, the normalization processing module 42 is further configured to: calculate a normalization factor based on the normalized envelope coefficients, and determine normalized frequency domain information based on the normalization factor and the variable frequency domain information.

[0155] In some embodiments, the correction processing module 43 is specifically used to: calculate the compensation vector of the resonance peak according to the normalized envelope coefficient, and determine the correction frequency domain information according to the normalized frequency domain information and the compensation vector.

[0156] In some embodiments, the correction processing module 43 is further specifically used to: correct the normalized envelope coefficient according to a preset resonance peak correction ratio to obtain a corrected envelope coefficient, wherein the number of coefficients of the corrected envelope coefficient is equal to the number of coefficients of the normalized envelope coefficient, and determine the compensation vector of the resonance peak according to the corrected envelope coefficient.

[0157] In some embodiments, the correction processing module 43 is further specifically used to: reduce or increase the number of coefficients of the normalized envelope coefficient according to a preset formant correction ratio to obtain a scaled envelope coefficient, and perform linear interpolation on the scaled envelope coefficient according to a linear interpolation algorithm to obtain a corrected envelope coefficient.

[0158] In some embodiments, the correction processing module 43 is further specifically used to: filter the correction envelope coefficients to obtain the compensation vector of the resonance peak.

[0159] In some embodiments, the correction processing module 43 is further specifically used to: multiply the normalized frequency domain information and the compensation vector to obtain the corrected frequency domain information.

[0160] In some embodiments, the time-domain speech generation module 44 is specifically used to perform inverse Fourier transform on the corrected frequency domain information and the phase information to obtain preliminary time-domain information, and to smooth the preliminary time-domain information to obtain time-domain speech information.

[0161] It should be noted that the above-described voice-changing processing device can execute the voice-changing processing method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in the embodiments of the voice-changing processing device can be found in the voice-changing processing method provided in the embodiments of the present invention.

[0162] Please see Figure 5 , Figure 5 This is a circuit structure diagram of an electronic device provided in an embodiment of the present invention, wherein the electronic device may be a voice changer, a mobile phone, a computer, etc.

[0163] like Figure 5 As shown, the electronic device 500 includes one or more processors 51 and a memory 52. ​​Wherein, Figure 5 Take the 51 processor as an example.

[0164] The processor 51 and the memory 52 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0165] The memory 52, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice-changing processing method in the embodiments of the present invention. The processor 51 executes various functional applications and data processing of the voice-changing processing device by running the non-volatile software programs, instructions, and modules stored in the memory 52, thereby realizing the functions of the voice-changing processing method provided in the above method embodiments and the various modules or units in the above device embodiments.

[0166] Memory 52 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 52 may optionally include memory remotely located relative to processor 51, which can be connected to processor 51 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0167] The program instructions / modules are stored in the memory 52 and, when executed by one or more processors 51, perform the voice-changing processing method in any of the above method embodiments.

[0168] This invention also provides a storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 5 One of the processors 51 can enable the one or more processors to execute the voice-changing processing method in any of the above method embodiments.

[0169] This invention also provides a chip, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice-changing processing method in any of the above method embodiments.

[0170] This invention also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by an electronic device, cause the electronic device to perform the voice-changing processing method in any of the above method embodiments.

[0171] The device or equipment embodiments described above are merely illustrative. The unit modules described as separate components may or may not be physically separate. The components shown as module units may or may not be physical units; that is, they may be located in one place or distributed across multiple network module units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0172] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice-changing processing method, characterized in that, include: Acquire voice-changing frequency domain information, which is information obtained by performing a Fourier transform on the voice-changing speech signal, the voice-changing speech signal is a speech signal processed by a voice-changing algorithm, and the voice-changing frequency domain information includes frequency domain information corresponding to formants. The frequency domain information of the altered sound is normalized to obtain normalized frequency domain information, which includes the frequency domain information of the formant after normalization. The normalized frequency domain information is corrected to obtain corrected frequency domain information, which includes the frequency domain information of the formant after correction. Time-domain speech information is generated based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information.

2. The method according to claim 1, characterized in that, The normalization process for the voice-changing frequency domain information to obtain normalized frequency domain information includes: Determine the normalized envelope coefficients of the voice-changing speech signal in the frequency domain; The normalized frequency domain information is normalized based on the normalized envelope coefficients to obtain normalized frequency domain information.

3. The method according to claim 2, characterized in that, Determining the normalized envelope coefficients of the voice-changing speech signal in the frequency domain includes: According to the linear prediction algorithm, the p-th order linear prediction coefficients of the voice-changing speech signal in the frequency domain are calculated, where p is a positive integer; The normalized envelope coefficients are determined based on the p-order linear prediction coefficients.

4. The method according to claim 3, characterized in that, The step of determining the normalized envelope coefficient based on the p-order linear prediction coefficient includes: The length of the p-order linear prediction coefficients is extended to the length of the target window function to obtain extended linear prediction coefficients, wherein the target window function is a window function that participates in the calculation of the variable sound frequency domain information; Perform a Fourier transform on the extended linear prediction coefficients to obtain the Fourier information of the coefficients; The normalized envelope coefficients are obtained by taking the modulus of the Fourier information of the coefficients.

5. The method according to claim 2, characterized in that, The step of normalizing the voice-changing frequency domain information based on the normalized envelope coefficients to obtain normalized frequency domain information includes: Calculate the normalization factor based on the normalized envelope coefficients; The normalized frequency domain information is determined based on the normalization factor and the variable sound frequency domain information.

6. The method according to claim 2, characterized in that, The step of correcting the normalized frequency domain information to obtain corrected frequency domain information includes: The compensation vector of the resonance peak is calculated based on the normalized envelope coefficient. Based on the normalized frequency domain information and the compensation vector, the corrected frequency domain information is determined.

7. The method according to claim 6, characterized in that, The step of calculating the compensation vector of the resonance peak based on the normalized envelope coefficient includes: According to the preset resonance peak correction ratio, the normalized envelope coefficient is corrected to obtain the corrected envelope coefficient, wherein the number of coefficients of the corrected envelope coefficient is equal to the number of coefficients of the normalized envelope coefficient. The compensation vector of the resonance peak is determined based on the modified envelope coefficient.

8. The method according to claim 7, characterized in that, The step of correcting the normalized envelope coefficient according to a preset resonance peak correction ratio to obtain the corrected envelope coefficient includes: Based on the preset resonance correction ratio, the number of coefficients in the normalized envelope coefficient is reduced or increased to obtain the scaled envelope coefficient. The scaled envelope coefficients are linearly interpolated using a linear interpolation algorithm to obtain the corrected envelope coefficients.

9. The method according to claim 7, characterized in that, The step of determining the compensation vector of the resonance peak based on the modified envelope coefficient includes: filtering the modified envelope coefficient to obtain the compensation vector of the resonance peak.

10. The method according to claim 6, characterized in that, The step of determining the corrected frequency domain information based on the normalized frequency domain information and the compensation vector includes: The normalized frequency domain information and the compensation vector are multiplied together to obtain the corrected frequency domain information.

11. The method according to any one of claims 1 to 10, characterized in that, The step of generating time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information includes: Perform an inverse Fourier transform on the corrected frequency domain information and the phase information to obtain preliminary time domain information; The preliminary time-domain information is smoothed to obtain time-domain speech information.

12. A storage medium, characterized in that, The device stores computer-executable instructions for causing the electronic device to perform the voice-changing processing method as described in any one of claims 1 to 11.

13. A chip, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the voice-changing processing method as described in any one of claims 1 to 11.

14. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the voice-changing processing method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and system using a long-term correlation difference between left and right channels for time domain down mixing a stereo sound signal into primary and secondary channels

    CN108352164A

  • Voice change processing method, device thereof and computer-readable storage medium

    CN109410973A