Speech pitch shifting methods, storage media and electronic devices
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]本发明实施例的一个目的旨在提供一种语音变调方法、存储介质及电子设备,旨在改善相关技术中语音变调不够自然的技术问题
[0013]在本发明实施例提供的语音变调方法中,获取语音信号,每帧语音信号包括多个语音采样点,确定目标语音采样点的至少一类目标相位信息,目标语音采样点为多个语音采样点中的一个语音采样点,根据目标相位信息平滑调整目标语音采样点的幅值,以得到变调语音信号。本实施例能够根据目标语音采样点的目标相位信息平滑调整目标语音采样点的幅值,如此可避免帧间不连续的现象出现,会使得变调语音信号变得更为自然。
Smart Images

Figure CN115985332B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech tone shifting technology, specifically to a speech tone shifting method, storage medium, and electronic device. Background Technology
[0002] Related technologies often use relatively simple pitch-shifting time-domain algorithms to process speech signals. For example, pitch-shifting time-domain algorithms include Synchronous Overlap Add (SOLA) joint resampling algorithm and Waveform Similarity Overlap-Add (WSOLA) joint resampling algorithm, which are pitch-shifting algorithms that do not change the speed. Although these pitch-shifting time-domain algorithms have low complexity, they can produce inter-frame discontinuity during the pitch-shifting process, which can cause the sound to click or stutter, making the pitch shifting less natural. Summary of the Invention
[0003] One objective of this invention is to provide a speech pitch shifting method, storage medium, and electronic device, aiming to improve the technical problem of unnatural speech pitch shifting in related technologies.
[0004] In a first aspect, embodiments of the present invention provide a speech pitch-changing method, comprising:
[0005] Acquire speech signals, each frame of which includes multiple speech sampling points;
[0006] Determine at least one type of target phase information for the target speech sampling point, wherein the target speech sampling point is one of the plurality of speech sampling points;
[0007] The amplitude of the target speech sampling point is smoothly adjusted according to the target phase information to obtain the pitch-shifted speech signal.
[0008] In a second aspect, embodiments of the present invention provide a storage medium storing computer-executable instructions for causing an electronic device to perform the aforementioned speech modulation method.
[0009] In a third aspect, embodiments of the present invention provide an electronic device, comprising:
[0010] At least one processor; and,
[0011] A memory communicatively connected to the at least one processor; wherein,
[0012] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described speech modulation method.
[0013] In the speech modulation method provided in this embodiment of the invention, a speech signal is acquired, each frame of the speech signal includes multiple speech sampling points, at least one type of target phase information of the target speech sampling point is determined, the target speech sampling point is one of the multiple speech sampling points, and the amplitude of the target speech sampling point is smoothly adjusted according to the target phase information to obtain the pitch-modified speech signal. This embodiment can smoothly adjust the amplitude of the target speech sampling point according to the target phase information of the target speech sampling point, thus avoiding the phenomenon of discontinuity between frames and making the pitch-modified speech signal more natural. Attached Figure Description
[0014] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0015] Figure 1 A schematic flowchart of a speech pitch shifting method provided in an embodiment of the present invention;
[0016] Figure 2 This is a schematic diagram of the circuit structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0018] It should be noted that, unless otherwise specified, the various features in the embodiments of this invention can be combined with each other, all of which are within the protection scope of this invention. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this invention do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0019] This invention provides a method for speech pitch shifting. Please refer to [link / reference]. Figure 1 The method of speech tone shifting includes the following steps:
[0020] S11: Acquire the speech signal. Each frame of the speech signal includes multiple speech sampling points.
[0021] In this step, the speech signal is the signal collected by the speech sensor, which includes sensors such as microphones. The speech signal can be represented as x(n)=[x(i),x(i-1),…,x(iM-1)], where M is the length of the filter and i represents the i-th speech sampling point.
[0022] S12: Determine at least one type of target phase information for the target speech sampling point, where the target speech sampling point is one of multiple speech sampling points.
[0023] In this step, phase information is used to represent the phase of the speech sampling point in the speech signal. In this embodiment, one type of phase information can be used to represent the phase of each speech sampling point, or two or more types of phase information can be used to represent the phase of each speech sampling point from different dimensions. The target phase information is the phase information corresponding to the target speech sampling point, which can be the i-th speech sampling point, i = 0, 1, 2, ...
[0024] S13: Smoothly adjust the amplitude of the target speech sampling points according to the target phase information to obtain the pitch-shifted speech signal.
[0025] This embodiment can smoothly adjust the amplitude of the target speech sampling point according to the target phase information of the target speech sampling point, thus avoiding the phenomenon of inter-frame discontinuity and making the pitch-shifted speech signal more natural.
[0026] In some embodiments, determining at least one type of target phase information for a target speech sampling point includes the following steps:
[0027] S121: Determine the phase step value and at least one type of phase information of the previous sampling point, wherein the previous sampling point is the speech sampling point that is located in time before the target speech sampling point.
[0028] S122: Determine at least one type of target phase information for the target speech sampling point based on at least one type of phase information and phase step value of the previous sampling points.
[0029] In S121, the phase step value is the difference in phase between two adjacent speech sampling points. In some embodiments, determining the phase step value includes the following steps: obtaining the pitch period, period modification rate, and sampling rate of the speech signal; and calculating the phase step value based on the pitch period, period modification rate, and sampling rate. Thus, this embodiment can incorporate the pitch period element and combine it with a user-defined period modification rate, thereby enabling the pitch period of the speech signal to be stretched or compressed in a custom way, so that the amplitude of each speech sampling point can be smoothly adjusted according to the phase information in the future.
[0030] In some embodiments, this embodiment uses Equation 1 to calculate the phase step value, as shown below:
[0031]
[0032] in, denoted as phase step, r as period modification rate, maxdelay as maximum pitch period length (which can be any period length between 20ms and 40ms), and Fs as sampling rate (where the period modification rate r can be positive or negative). By adjusting the period modification rate r, the pitch period of the stretched or compressed speech signal can be obtained.
[0033] The previous sampling point can be a preset number of semantic sampling points preceding the target speech sampling point, such as 1 or 2. When the preset number is 1, the previous sampling point is the speech sampling point preceding the target speech sampling point. That is, if the target speech sampling point is the i-th speech sampling point, then the previous sampling point is the (i-1)-th speech sampling point.
[0034] In S122, in some embodiments, determining at least one type of target phase information of the target speech sampling point based on at least one type of phase information of the previous sampling points and the phase step value includes:
[0035] S1221: Add at least one type of phase information and phase step value from the previous sampling points to obtain the summation result.
[0036] S1222: Perform a modulo operation on the addition result to obtain the modulo result, and use the modulo result as at least one type of target phase information of the target speech sampling point.
[0037] In S1221, this embodiment calculates the summation result according to formula two, as shown below:
[0038]
[0039] in, For the j-th phase information of the (i-1)-th speech sampling point, The summation result of the j-th phase information at the i-th speech sampling point. The phase step value is i, which is less than or equal to M, and both i and j are positive integers.
[0040] In S1222, this embodiment calculates the remainder result according to formula three, as shown below:
[0041]
[0042] in, Let be the j-th phase information (i.e., the target phase information) of the i-th speech sampling point, and mod represents the modulo operation.
[0043] It is understood that j can be 1, 2, or 3, etc. When j is 1, this embodiment can smoothly adjust the amplitude based on one type of phase information. When j is 2 or 3, this embodiment can smoothly adjust the amplitude based on multiple types of phase information.
[0044] In some embodiments, the phase information includes first phase information and second phase information. In this embodiment, the first phase information is calculated according to formula four, and the second phase information is calculated according to formula five, as shown below:
[0045]
[0046]
[0047] In some embodiments, smoothly adjusting the amplitude of the target speech sampling point based on the target phase information includes the following steps:
[0048] S131: Select the speech smoothing model that meets the phase conditions of the target phase information as the target speech smoothing model.
[0049] S132: Determine the number of smoothing points and the smoothing gain based on the target speech smoothing model.
[0050] S133: Smoothly adjust the amplitude of the target speech sampling points based on the number of smoothing points and the smoothing gain.
[0051] In S131, the phase condition is a condition used to match the target phase information to select a speech smoothing model. The phase condition corresponds to the speech smoothing model. Different speech smoothing models have different phase conditions. Multiple speech smoothing models are associated sequentially according to a preset order. In this embodiment, the corresponding speech smoothing model can be selected sequentially from multiple speech smoothing models as the target speech smoothing model according to the target phase information.
[0052] In some embodiments, the phase conditions include a first phase condition, a second phase condition, a third phase condition, and a fourth phase condition. The first phase condition corresponds to a first speech smoothing model ①, the second phase condition corresponds to a second speech smoothing model ②, the third phase condition corresponds to a third speech smoothing model ③, and the fourth phase condition corresponds to a fourth speech smoothing model ④. The association order of the speech smoothing models is any of the following: ①→②→③→④, ②→③→④→①, ③→④→①→②, ④→①→②→③.
[0053] When the target phase information satisfies the first phase condition of the first speech smoothing model ①, this embodiment selects the first speech smoothing model ① as the target speech smoothing model.
[0054] When the target phase information changes, the changed target phase information does not satisfy the first phase condition, but it will satisfy the second phase condition of the second speech smoothing model ②. Therefore, in this embodiment, the second speech smoothing model ② is selected as the target speech smoothing model.
[0055] When the target phase information changes, the changed target phase information does not satisfy the second phase condition, but it will satisfy the third phase condition of the third speech smoothing model ③. Therefore, in this embodiment, the third speech smoothing model ③ is selected as the target speech smoothing model.
[0056] When the target phase information changes, the changed target phase information does not satisfy the third phase condition, but it will satisfy the fourth phase condition of the fourth speech smoothing model ④. Therefore, in this embodiment, the fourth speech smoothing model ④ is selected as the target speech smoothing model.
[0057] When the target phase information changes, the changed target phase information does not satisfy the fourth phase condition, but it will satisfy the first phase condition of the first speech smoothing model ①. Therefore, in this embodiment, the first speech smoothing model ① is selected as the target speech smoothing model, and so on.
[0058] Understandably, the number of speech smoothing models can be 2, 3, or 5, etc.
[0059] In some embodiments, the speech smoothing model includes a normal speech smoothing model and an activated speech smoothing model, wherein the activated speech smoothing model is used to modify target phase information in order to activate the normal speech smoothing model located after the activated speech smoothing model.
[0060] The normal speech smoothing model and the active speech smoothing model are alternately arranged in a preset order. For example, as mentioned above, the first speech smoothing model ① and the third speech smoothing model ③ are both normal speech smoothing models, and the second speech smoothing model ② and the fourth speech smoothing model ④ are both active speech smoothing models. The two normal speech smoothing models and the two active speech smoothing models are alternately arranged in a preset order.
[0061] In S132, the normal speech smoothing model includes a normal smoothing point model and a normal gain smoothing model. The normal smoothing point search model is used to determine the number of smoothing points in the normal smoothing mode based on the target phase information, and the normal gain smoothing model is used to determine the smoothing gain in the normal smoothing mode based on the target phase information.
[0062] In some embodiments, determining the number of smoothing points and the smoothing gain based on the target speech smoothing model includes the following steps: when the target speech smoothing model is a normal speech smoothing model, determining the number of smoothing points based on the normal smoothing point model and determining the smoothing gain based on the normal gain smoothing model.
[0063] In some embodiments, the normal gain smoothing model is represented using trigonometric functions, such as cosine or sine functions. The independent variable of the normal gain smoothing model follows the change in target phase information. The smoothing gain of the normal gain smoothing model is adjusted by changing the normal gain smoothing model. In this embodiment, a trigonometric function is used as the normal gain smoothing model to control the smoothing gain to change in [0,1] in accordance with the changes in the target phase information, so as to smoothly and effectively smooth the amplitude of the target speech sampling points.
[0064] In some embodiments, the target phase information is normalized phase information. For example, as described in Equation 3 above, the normalized target phase information can be obtained by performing a modulo operation between the sum and the natural number 1. This ensures that the target phase information can be mapped and associated with the smoothing gain in the range [0,1] using trigonometric functions, which is beneficial for ensuring continuous and uninterrupted inter-frame operation.
[0065] When the target phase information is within the first numerical range [0, a), the target phase information is positively correlated with the smoothing gain. When the target phase information is within the second numerical range [a, 1], the target phase information is negatively correlated with the smoothing gain. This embodiment utilizes this relationship to ensure that the smoothing gain changes regularly within [0, 1] as the target phase information changes. For example, it first increases within the first numerical range [0, a), and then decreases within the second numerical range [a, 1]. The value of 'a' can be customized by the designer based on engineering experience, for example, 'a' can be 0.5.
[0066] For example:
[0067] When the target phase information is 0, the independent variable of the normal gain smoothing model is 90 degrees. Therefore, the normal gain smoothing model cos90°=0, that is, the smoothing gain is 0.
[0068] When the target phase information is 0.1, the independent variable of the normal gain smoothing model is 72 degrees. Therefore, the normal gain smoothing model cos72°=0.309, that is, the smoothing gain is 0.309.
[0069] When the target phase information is 0.2, the independent variable of the normal gain smoothing model is 72 degrees. Therefore, the normal gain smoothing model cos54°=0.588, that is, the smoothing gain is 0.588.
[0070] When the target phase information is 0.5, the independent variable of the normal gain smoothing model is 0 degrees. Therefore, the normal gain smoothing model cos0°=1, that is, the smoothing gain is 1.
[0071] When the target phase information is 0.6, the independent variable of the normal gain smoothing model is 18 degrees. Therefore, the normal gain smoothing model cos18°=0.951, that is, the smoothing gain is 0.951.
[0072] When the target phase information is 0.8, the independent variable of the normal gain smoothing model is 54 degrees. Therefore, the normal gain smoothing model cos54°=0.588, that is, the smoothing gain is 0.588.
[0073] In some embodiments, the normal speech smoothing model includes a first normal speech smoothing model and a second normal speech smoothing model, wherein the first normal speech smoothing model and the second normal speech smoothing model adopt the same normal smoothing point model and normal gain smoothing model.
[0074] In some embodiments, the expression for the normal gain smoothing model is as shown in Equation 6:
[0075]
[0076] in, The smoothing gain is the j-th phase information of the i-th speech sampling point, and k is an adjustment factor, such as k = 1, 2, or 3.
[0077] In some embodiments, the target phase information includes normalized first phase information and second phase information. In this embodiment, the first smoothing gain is calculated according to Equation 7, and the second smoothing gain is calculated according to Equation 8, as shown below:
[0078]
[0079]
[0080] Where a = 1 / k. When k is 2, a is 0.5.
[0081] In some embodiments, the normal smoothing point model is constrained by the pitch period, sampling rate, and target phase information. Please refer to Equation Nine; the expression for the normal smoothing point model is as follows:
[0082]
[0083] in, The number of smoothing points corresponding to the j-th phase information of the i-th speech sampling point.
[0084] In some embodiments, the phase information includes first phase information and second phase information. In this embodiment, the number of first smoothing points is calculated according to formula ten, and the number of second smoothing points is calculated according to formula eleven, as shown below:
[0085]
[0086]
[0087] In general, normal speech smoothing models include normal smoothing point models and normal gain smoothing models. The expression for the normal smoothing point model is: The expression for the normal gain smoothing model is:
[0088] In some embodiments, the target phase information includes normalized first phase information and second phase information, as described above, the normalized first phase information and second phase information are respectively and
[0089] The speech smoothing model includes an active speech smoothing model. Determining the number of smoothing points and smoothing gain based on the target speech smoothing model includes: when the target speech smoothing model is an active speech smoothing model, determining the number of smoothing points and smoothing gain based on the active speech smoothing model.
[0090] In some embodiments, the target phase information includes normalized first phase information and second phase information, the activated speech smoothing model includes a first activated speech smoothing model and a second activated speech smoothing model, and determining the number of smoothing points and smoothing gain according to the activated speech smoothing model includes: when the target speech smoothing model is the first activated speech smoothing model, determining the number of third smoothing points and the third smoothing gain according to the first activated speech smoothing model; when the target speech smoothing model is the second activated speech smoothing model, determining the number of fourth smoothing points and the fourth smoothing gain according to the second activated speech smoothing model.
[0091] In some embodiments, each active speech smoothing model includes an active smoothing point model and an active gain smoothing model. Determining the number of third smoothing points and the third smoothing gain based on the first active speech smoothing model includes the following steps: when the target speech smoothing model is the first active speech smoothing model, determining the number of third smoothing points based on the first active smoothing point model, and determining the third smoothing gain based on the first active gain smoothing model.
[0092] In some embodiments, determining the number of second smoothing points and the second smoothing gain based on the second active speech smoothing model includes the following steps: when the target speech smoothing model is the second active speech smoothing model, determining the number of fourth smoothing points based on the second active smoothing point model, and determining the fourth smoothing gain based on the second active gain smoothing model.
[0093] In some embodiments, the expression for the first activated smoothing point model is shown in Equations XII and XIII:
[0094]
[0095]
[0096] In some embodiments, the expressions for the first activation gain smoothing model are shown in Equations 14 and 15:
[0097]
[0098]
[0099] In some embodiments, the expression for the second activated smoothing point model is shown in Equations 16 and 17:
[0100]
[0101]
[0102] In some embodiments, the expressions for the second activation gain smoothing model are shown in Equations 18 and 19:
[0103]
[0104]
[0105] In some embodiments, the target phase information includes normalized first phase information and second phase information. The first phase condition is: 0 ≤ first phase information ≤ 0.5, 0.5 ≤ second phase information ≤ 1. The second phase condition is: 0.5 < first phase information ≤ 1, 0 ≤ second phase information ≤ 0.5. The third phase condition is: 0 ≤ first phase information < 0.5, 0 ≤ second phase information < 0.5. The fourth phase condition is: 0.5 ≤ first phase information ≤ 1, 0.5 ≤ second phase information ≤ 1. Wherein, the first phase condition corresponds to a first normal speech smoothing model, the second phase condition corresponds to a first activated speech smoothing model, the third phase condition corresponds to a second normal speech smoothing model, and the fourth phase condition corresponds to a second activated speech smoothing model.
[0106] when Within the range [0, 0.5] and When the target phase information matches the first phase condition within the range of [0.5,1], the first normal speech smoothing model is selected as the target speech smoothing model. The first normal speech smoothing model includes a normal smoothing point model and a normal gain smoothing model, expressed as follows:
[0107]
[0108]
[0109]
[0110]
[0111] when Within the range (0.5,1) and When the target phase information matches the second phase condition within the range [0, 0.5], the first activation speech smoothing model is selected as the target speech smoothing model. The first activation speech smoothing model includes a first activation smoothing point model and a first activation gain smoothing model, expressed as follows:
[0112]
[0113]
[0114]
[0115]
[0116] when Within the range [0, 0.5) and When the target phase information matches the third phase condition within the range [0, 0.5), the second normal speech smoothing model is selected as the target speech smoothing model. The second normal speech smoothing model includes a normal smoothing point model and a normal gain smoothing model, expressed as follows:
[0117]
[0118]
[0119]
[0120]
[0121] when Within the range [0.5,1] and When the target phase information matches the fourth phase condition within the range of [0.5,1], the second activation speech smoothing model is selected as the target speech smoothing model. The second activation speech smoothing model includes a second activation smoothing point model and a second activation gain smoothing model, expressed as follows:
[0122]
[0123]
[0124]
[0125]
[0126] As mentioned above, when there are multiple speech smoothing models, the multiple speech smoothing models are associated sequentially in a preset order. The speech pitch shifting method also includes: when the target speech smoothing model is an active speech smoothing model, modifying at least one type of target phase information so that the modified target phase information satisfies the phase condition corresponding to the candidate speech smoothing model, and the candidate speech smoothing model is the speech smoothing model located after the target speech smoothing model.
[0127] When the target phase information includes normalized first phase information and second phase information, modifying at least one type of target phase information such that the modified target phase information satisfies the phase condition corresponding to the candidate speech smoothing model includes: modifying the first phase information and keeping the second phase information so that the modified target phase information satisfies the phase condition corresponding to the candidate speech smoothing model.
[0128] For example, since the target phase information satisfies the second phase condition, the target speech smoothing model is the first activated speech smoothing model. Since the second normal speech smoothing model follows the first activated speech smoothing model, the second normal speech smoothing model is a candidate speech smoothing model. Next, this embodiment will use the first phase information... Revised to: Maintain second phase information Unchanged, thus, the modified first phase information (Right now and second phase information (Right now The third phase condition is met; therefore, subsequent steps are based on the target phase information. and It can then automatically enter the second normal speech smoothing model to perform smoothing operations.
[0129] For another example, since the target phase information satisfies the fourth phase condition, the target speech smoothing model is the second activated speech smoothing model. Because the first normal speech smoothing model follows the second activated speech smoothing model, the first normal speech smoothing model is a candidate speech smoothing model. Next, this embodiment will use the first phase information... Revised to: Maintain the second phase information p i 2 Unchanged, thus, the modified first phase information (Right now and second phase information (Right now The first phase condition is met; therefore, subsequent steps are based on the target phase information. and It can then automatically enter the first normal speech smoothing model for smoothing operation.
[0130] This embodiment modifies the target phase information so that the smoothing gain changes sequentially and cyclically in multiple speech smoothing models, thereby improving the continuity and smoothness between frames.
[0131] In some embodiments, smoothly adjusting the amplitude of the target speech sampling point according to the number of smoothing points and the smoothing gain includes the following steps:
[0132] S1331: Determine the target smoothing point corresponding to the number of smoothing points for each type. The target smoothing point is located before the target speech sampling point in time. The target speech sampling point is spaced apart from the target smoothing point by the target number of speech sampling points. The target number is equal to the number of smoothing points minus the natural number 1.
[0133] S1332: Based on the smoothing gain and the amplitude of the target smoothing point, smoothly adjust the amplitude of the target speech sampling point.
[0134] In S1331, the number of smoothed points can be of one or more categories. In this embodiment, the target smoothed points are obtained according to the following formula: Let j be the target smoothing point corresponding to the j-th phase information of the i-th speech sampling point.
[0135] In some embodiments, the target phase information includes normalized first phase information and second phase information, and the first target smoothing point corresponding to the first phase information is: Let be the first target smoothing point corresponding to the first phase information of the i-th speech sampling point. The first target smoothing point is the point that is the nth time relative to the i-th speech sampling point in the past. One voice sampling point.
[0136] The second target smoothing point corresponding to the second phase information is: The second target smoothing point is the second phase information corresponding to the i-th speech sampling point. The second target smoothing point is the point that is the nth time relative to the i-th speech sampling point in the past. One voice sampling point.
[0137] For example, a frame of signal is {x1, x2, x3, x4, x5, x6, x7, x8}. When the loop reaches the 7th speech sampling point (i.e., the target speech sampling point) at i=7, when... At that time, then The fourth speech sampling point is x4, which is the target smoothing point. When the second frame {x9, x10, x11, x12} comes in, and the loop reaches the third point, the third point is x11. hour, x8 is the target smoothing point.
[0138] In S1332, when the target phase information is a type of phase information, the amplitude of the target speech sampling point is smoothly adjusted according to the smoothing gain and the amplitude of the target smoothing point, which includes: multiplying the smoothing gain by the amplitude of the target smoothing point to obtain the smoothed amplitude, and using the smoothed amplitude as the amplitude of the target speech sampling point.
[0139] When the target phase information includes at least two types of phase information, the amplitude of the target speech sampling point is smoothly adjusted according to the smoothing gain and the amplitude of the target smoothing point. This includes: calculating the weighted amplitude based on the amplitude of each target smoothing point and the smoothing gain corresponding to each target smoothing point, and combining the weighted amplitudes of multiple target smoothing points to obtain the amplitude of the target speech sampling point.
[0140] This embodiment uses the following formula to calculate the weighted magnitude: in, It is the weighted amplitude corresponding to the j-th phase information of the i-th speech sampling point.
[0141] In some embodiments, the target phase information includes normalized first phase information and second phase information, and the amplitude of the target speech sampling point is: The amplitude of the target speech sampling point.
[0142] This embodiment comprehensively considers multiple types of phase information of the target speech sampling point, calculates the amplitude corresponding to each type of phase information through a weighted method, and then combines the amplitudes of multiple types of phase information to obtain the amplitude of the target speech sampling point. This can smooth the speech signal more, improve the continuity between frames, and thus improve the naturalness of the speech signal.
[0143] As mentioned earlier, this embodiment utilizes trigonometric functions to stretch or compress the pitch period, achieving a time-domain method for modifying the pitch effect. Furthermore, this embodiment can reconstruct the formants of the pitch-shifted speech signal, further enhancing the voice-changing effect, as detailed below:
[0144] In this embodiment, the voice-changing frequency domain information is obtained, and the voice-changing frequency domain information is normalized to obtain normalized frequency domain information. The normalized frequency domain information includes the frequency domain information after the formants are normalized. The normalized frequency domain information is then corrected to obtain corrected frequency domain information. The corrected frequency domain information includes the frequency domain information after the formants are corrected. Based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information, time-domain speech information is generated.
[0145] The voice-changing frequency domain information is the information obtained by performing a Fourier transform on the pitch-shifted speech signal, and includes the frequency domain information corresponding to the formants. In this embodiment, the pitch-shifted speech signal is windowed according to a target window function to obtain windowed speech information. A Fourier transform is then performed on the windowed speech information to obtain the voice-changing frequency domain information. For example, this embodiment performs a Fourier transform on the pitch-shifted speech signal according to the following formula to obtain the voice-changing frequency domain information:
[0146] X(n)=fft([x(n-1);x(n)·win])
[0147] Where X(n) represents the frequency domain information of the nth frame, win represents the target window function, which can be the Hanning window in this paper. The length of the Hanning window is 2*M, and fft represents the Fourier transform. In some embodiments, the target window function is the Hanning window function.
[0148] Normalized frequency domain information is the frequency domain information obtained after normalizing the voice-changing frequency domain information. Formants are regions in the speech signal spectrum where energy is relatively concentrated.
[0149] Time-domain speech information is information converted from the frequency domain to the time domain of the target frequency domain information. The target frequency domain information is obtained from the phase information of the jointly corrected frequency domain information and the voice-changing frequency domain information. For example, in this embodiment, the time-domain speech information can be obtained according to the following formula:
[0150]
[0151]
[0152] Where Y1(n) represents the corrected frequency domain information of the nth frame, and Y2(n) represents the target frequency domain information of the nth frame. X(n) represents the phase information of the frequency domain information of the nth frame's altered audio, and X(n) represents the frequency domain information of the nth frame's altered audio.
[0153] In some embodiments, this embodiment performs an inverse Fourier transform on the target frequency domain information to obtain time-domain speech information. For example, this embodiment can obtain time-domain speech information according to the following formula:
[0154] y(n) = ifft(Y2(n)) * win
[0155] Where y(n) represents time-domain speech information, and ifft represents the inverse Fourier transform.
[0156] In some embodiments, generating time-domain speech information based on the phase information of the corrected frequency domain information and the voice-changing frequency domain information includes the following steps:
[0157] S141: Perform an inverse Fourier transform on the corrected frequency domain information and phase information to obtain preliminary time domain information.
[0158] S142: Smooth the initial time-domain information to obtain time-domain speech information.
[0159] In S141, as mentioned above, this embodiment can use Equation 4 to perform an inverse Fourier transform on the corrected frequency domain information and phase information in order to obtain preliminary time domain information.
[0160] In S142, in order to improve the naturalness and smoothness of the voice-changing speech signal, this embodiment can perform smoothing processing on the preliminary time-domain information to obtain time-domain speech information. For example, this embodiment can obtain the time-domain speech information according to the following formula, as follows:
[0161] out(n) = y(1:M) + out_last(n)
[0162] out_last(n) = y(M+1:2*M)
[0163] Where out(n) is the time-domain speech information, and out_last(n) is the second half of the previous frame data.
[0164] This embodiment can effectively smooth and recover the initial time-domain information, thereby enabling the natural and smooth output of time-domain speech information.
[0165] In some embodiments, normalizing the voice-changing frequency domain information to obtain normalized frequency domain information includes the following steps: determining the normalized envelope coefficients of the pitch-changing speech signal in the frequency domain, and normalizing the voice-changing frequency domain information according to the normalized envelope coefficients to obtain normalized frequency domain information.
[0166] The normalized envelope coefficient is the set of amplitudes of each resonance peak in the frequency domain.
[0167] In some embodiments, this embodiment calculates the p-th order linear prediction coefficients of the pitch-shifted speech signal in the frequency domain according to a linear prediction algorithm, where p is a positive integer, and determines the normalized envelope coefficients based on the p-th order linear prediction coefficients.
[0168] In this embodiment, the Levinson-Dubin formula autocorrelation method is used to solve for the linear prediction coefficients. The linear prediction system can be represented by the following formula:
[0169]
[0170] in, Let x(n) be an estimated value. a is obtained by linear combination of the past p values. iFor a linear prediction function, the transfer function of a p-order linear predictor is as follows:
[0171]
[0172] The autocorrelation function r(j) is shown below:
[0173]
[0174] The minimum mean square error E can be written as:
[0175] The recursive derivation is performed step by step using a recursive solution:
[0176] ① When i = 0, E = r(0), a0 = 1
[0177] ② Corresponding to the i-th recursion (i = 1, 2, 3, ..., p):
[0178]
[0179] a j (i) =k i
[0180] a j (i) =a j (i-1) -k i a i-j (i-1)
[0181] E i =(1-k) i 2 E i-1
[0182] The superscript in the parentheses above indicates the order of the predictor. By recursively solving for i = 1, 2, ..., p, we obtain:
[0183] a i =a j (p) 1≤j≤p
[0184] In this embodiment, the linear prediction coefficient a of order p is retained. i .
[0185] In some embodiments, determining the normalized envelope coefficients based on the p-order linear prediction coefficients includes the following steps: extending the length of the p-order linear prediction coefficients to the length of the target window function to obtain extended linear prediction coefficients, wherein the target window function is a window function that participates in the calculation of the variable sound frequency domain information; performing a Fourier transform on the extended linear prediction coefficients to obtain coefficient Fourier information; and performing modulus calculation on the coefficient Fourier information to obtain the normalized envelope coefficients.
[0186] In this embodiment, zeros are padded after the p-order linear prediction coefficients to extend the length of the p-order linear prediction coefficients to the length of the target window function, that is, to expand it to a frame of data of size 2*M, where 2*M is the length of the target window function. For example, in this embodiment, zeros are padded after the p-order linear prediction coefficients to obtain ar(n) = [a1, a2, a3, ..., a p ,0,0,0,......,0].
[0187] In this embodiment, the extended linear prediction coefficients are subjected to Fourier transform according to the following formula to obtain the Fourier information of the coefficients, as follows:
[0188] F = fft(ar(n))
[0189] Where F represents the Fourier information of the coefficients.
[0190] In this embodiment, the normalized envelope coefficients are obtained by modulo the Fourier information of the coefficients according to the following formula:
[0191] AR(n) = abs(F, 2M)
[0192] Where AR(n) is the normalized envelope coefficient of the nth frame, and abs represents the modulus.
[0193] In some embodiments, normalizing the voice-changing frequency domain information according to the normalized envelope coefficients to obtain normalized frequency domain information includes the following steps: calculating the normalization factor according to the normalized envelope coefficients, and determining the normalized frequency domain information according to the normalization factor and the voice-changing frequency domain information.
[0194] In some embodiments, this embodiment calculates the reciprocal of the normalized envelope coefficient and uses it as the normalization factor. In some embodiments, this embodiment calculates the normalization factor based on the normalized envelope coefficient and the division protection factor. For example, this embodiment calculates the normalization factor according to the following formula:
[0195] α(n) = 1(AR(n) + δ)
[0196] Where α(n) is the normalization factor and δ is the division protection factor. δ can be customized by the designer according to business needs, for example, δ is 0.00001.
[0197] In some embodiments, this embodiment calculates the modulus of the variable frequency domain information, and then divides the modulus of the variable frequency domain information by a normalization factor to obtain the normalized frequency domain information. For example, this embodiment determines the normalized frequency domain information according to the following formula:
[0198] Y(n)=abs(X(n))α(n)
[0199] Where Y(n) represents the normalized frequency domain information.
[0200] In some embodiments, the process of correcting the normalized frequency domain information to obtain the corrected frequency domain information includes the following steps: calculating the compensation vector of the formant based on the normalized envelope coefficients, and determining the corrected frequency domain information based on the normalized frequency domain information and the compensation vector.
[0201] The compensation vector is used to correct the amplitude of the resonance peak. In some embodiments, calculating the compensation vector of the resonance peak based on the normalized envelope coefficient includes: correcting the normalized envelope coefficient to obtain a corrected envelope coefficient according to a preset resonance peak correction ratio, wherein the number of coefficients in the corrected envelope coefficient is equal to the number of coefficients in the normalized envelope coefficient; and determining the compensation vector of the resonance peak based on the corrected envelope coefficient. The preset resonance peak correction ratio can be customized by the designer based on engineering experience.
[0202] In some embodiments, correcting the normalized envelope coefficients according to a preset formant correction ratio to obtain corrected envelope coefficients includes: reducing or increasing the number of coefficients in the normalized envelope coefficients according to the preset formant correction ratio to obtain scaled envelope coefficients; and performing linear interpolation on the scaled envelope coefficients according to a linear interpolation algorithm to obtain corrected envelope coefficients.
[0203] For example, the normalized envelope coefficient AR(n) = [1,2,3,.....,2*M], with a preset formant correction ratio rate = 2. Therefore, in this embodiment, the number of coefficients in the normalized envelope coefficient is reduced by half, resulting in the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M]. In this embodiment, based on a linear interpolation algorithm, the uninterpolated normalized envelope coefficient AR(n) = [1,2,3,.....,2*M] is linearly inserted into the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M], thus expanding the scaled envelope coefficient AR_interp(n) = [1,3,5,.....,M] into a frame of data of size 2*M, resulting in the corrected envelope coefficient AR_interp1(n) = [1,3,5,.....,M,α(M+1),.....,α(2*M)].
[0204] It is understandable that when the preset resonant correction ratio rate = 0.5, this embodiment amplifies the number of coefficients in the normalized envelope coefficients, resulting in the scaled envelope coefficients AR_interp(n) = [1, 1.5, 2, 2.5, 3, 3.5, ..., 2*M].
[0205] In some embodiments, determining the compensation vector of the resonance peak based on the modified envelope coefficient includes: directly using the modified envelope coefficient as the compensation vector of the resonance peak.
[0206] In some embodiments, determining the compensation vector of the formant based on the modified envelope coefficient includes: filtering the modified envelope coefficient to obtain the compensation vector of the formant. Thus, this embodiment can obtain a compensation vector with a relatively high degree of smoothness and fewer spikes, which is beneficial for obtaining a relatively smooth modified frequency domain information based on the normalized frequency domain information and the compensation vector. Furthermore, based on the phase information of the modified frequency domain information and the voice-changing frequency domain information, a relatively smooth time-domain speech information can be generated, which makes the voice changing more natural and coherent.
[0207] In some embodiments, the filtering process includes an initial filtering process. Filtering the modified envelope coefficients to obtain the compensation vector of the resonance peak includes the following steps: performing an initial filtering process on the modified envelope coefficients to obtain an initial vector of the resonance peak, and determining the compensation vector of the resonance peak based on the initial vector of the resonance peak.
[0208] In some embodiments, the initial filtering of the modified envelope coefficients to obtain the initial vector of the formant includes the following steps: obtaining the modified envelope coefficients of the previous frame, the modified envelope coefficients of the current frame, and the modified envelope coefficients of the next frame; calculating the first coefficient weighting result based on the modified envelope coefficients of the previous frame and the first preset weight; calculating the second coefficient weighting result based on the modified envelope coefficients of the current frame and the second preset weight; calculating the third coefficient weighting result based on the modified envelope coefficients of the next frame and the third preset weight; and adding the first coefficient weighting result, the second coefficient weighting result, and the third coefficient weighting result to obtain the initial vector of the formant.
[0209] In some embodiments, the second preset weight is greater than the first preset weight and the third preset weight, thereby increasing the influence of the current frame's corrected envelope coefficients on determining the initial vector. In some embodiments, the first preset weight is greater than the third preset weight, thereby increasing the influence of the previous frame's corrected envelope coefficients on determining the initial vector.
[0210] For example, in this embodiment, the initial vector of the resonance peak is determined according to the following formula:
[0211] ψ1=AR_interp1(i-1)*η1
[0212] ψ2=AR_interp1(i)*η2
[0213] ψ3=AR_interp1(i+1)*η3
[0214] AR_interp2(i)=ψ1+ψ2+ψ3
[0215] Where AR_interp1(i-1) is the corrected envelope coefficient of the previous frame, η1 is the first preset weight, and ψ1 is the first weighted result. AR_interp1(i) is the corrected envelope coefficient of the current frame, η2 is the second preset weight, and ψ2 is the second weighted result. AR_interp1(i+1) is the corrected envelope coefficient of the next frame, η3 is the third preset weight, and ψ3 is the third weighted result. AR_interp2(i) is the initial vector. i = 2, 3, 4, ..., 2*M-1. In some embodiments, η1 is 0.2, η2 is 1, and η3 is 0.01.
[0216] In some embodiments, determining the compensation vector of a resonance peak based on its initial vector includes: directly using the initial vector of the resonance peak as the compensation vector of the resonance peak.
[0217] In some embodiments, the filtering process includes median filtering, and determining the compensation vector of the resonance peak based on the initial vector of the resonance peak includes: performing median filtering on the initial vector of the resonance peak to obtain the compensation vector of the resonance peak.
[0218] For example, in this embodiment, the initial vector of the resonance peak is subjected to median filtering according to the following formula to obtain the compensation vector of the resonance peak, as follows:
[0219] AR_interp3(n)=medfilt(AR_interp2(n))
[0220] This embodiment, through initial filtering and median filtering, can reliably and effectively filter out the spikes in the compensation vector, making the compensation vector smoother.
[0221] In some embodiments, determining the corrected frequency domain information based on the normalized frequency domain information and the compensation vector includes multiplying the normalized frequency domain information and the compensation vector to obtain the corrected frequency domain information. For example, this embodiment obtains the corrected frequency domain information according to the following formula, thereby enabling the refitting and generation of formants, as follows:
[0222] Y1(n) = Y(n) * AR_interp3(n)
[0223] After obtaining the corrected frequency domain information, this embodiment can combine it with the formula mentioned above to output time-domain speech information. As mentioned earlier, the voice-changing time-domain speech information provided in this embodiment is relatively natural and smooth.
[0224] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0225] Please see Figure 2 , Figure 2 This is a schematic diagram of the circuit structure of an electronic device provided in an embodiment of the present invention. Figure 2 As shown, the electronic device 200 includes one or more processors 21 and a memory 22. Wherein, Figure 2 Take a processor 21 as an example.
[0226] Processor 21 and memory 22 can be connected via a bus or other means. Figure 2 Taking the example of a connection between China and Israel via a bus.
[0227] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the speech pitch shifting method in the embodiments of the present invention. The processor 21 executes the function of the speech pitch shifting method provided in the above method embodiments by running the non-volatile software programs, instructions, and modules stored in the memory 22.
[0228] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may optionally include memory remotely located relative to processor 21, which can be connected to processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0229] The program instructions / modules are stored in the memory 22 and, when executed by one or more processors 21, perform the speech modulation method in any of the above method embodiments.
[0230] This invention also provides a storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 2One of the processors 21 can enable the one or more processors to execute the speech modulation method in any of the above method embodiments.
[0231] This invention also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by an electronic device, cause the electronic device to perform any of the speech modulation methods described above.
[0232] The device or equipment embodiments described above are merely illustrative. The unit modules described as separate components may or may not be physically separate. The components shown as module units may or may not be physical units; that is, they may be located in one place or distributed across multiple network module units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0233] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for speech tone redialing, characterized in that, include: Acquire a speech signal, each frame of the speech signal including multiple speech sampling points, each speech sampling point having at least one type of phase information, the phase information being used to characterize the phase of the speech sampling point in the speech signal; Determine at least one type of target phase information for the target speech sampling point, wherein the target speech sampling point is one of the plurality of speech sampling points; Select the speech smoothing model that satisfies the phase condition of the target phase information as the target speech smoothing model; The number of smoothing points and the smoothing gain are determined based on the target phase information. The amplitude of the target speech sampling point is smoothly adjusted according to the number of smoothing points and the smoothing gain to obtain the pitch-shifted speech signal.
2. The method according to claim 1, characterized in that, The at least one type of target phase information for determining the target speech sampling point includes: Determine at least one type of phase information, including the phase step value and previous sampling points, wherein the previous sampling points are speech sampling points that are located in time before the target speech sampling point; Based on at least one type of phase information of the previous sampling points and the phase step value, at least one type of target phase information of the target speech sampling point is determined.
3. The method according to claim 2, characterized in that, Determining the phase step value includes: Obtain the pitch period, period modification rate, and sampling rate of the speech signal; The phase step value is calculated based on the pitch period, period modification rate, and sampling rate.
4. The method according to claim 2, characterized in that, Determining at least one type of target phase information for the target speech sampling point based on at least one type of phase information from the previous sampling points and the phase step value includes: Add at least one type of phase information and the phase step value of the previously sampled points to obtain the addition result; The addition result is subjected to a remainder operation to obtain a remainder result, and the remainder result is used as at least one type of target phase information of the target speech sampling point.
5. The method according to claim 1, characterized in that, The speech smoothing model includes a normal speech smoothing model, which includes a normal smoothing point model and a normal gain smoothing model. Determining the number of smoothing points and the smoothing gain based on the target speech smoothing model includes: When the target speech smoothing model is a normal speech smoothing model, the number of smoothing points is determined according to the normal smoothing point model; The smoothing gain is determined based on the normal gain smoothing model.
6. The method according to claim 5, characterized in that, The normal gain smoothing model is represented using trigonometric functions; The independent variable of the normal gain smoothing model follows the change of the target phase information. The changes are made to adjust the smoothing gain of the normal gain smoothing model output.
7. The method according to claim 6, characterized in that, The trigonometric function is the cosine function.
8. The method according to claim 5, characterized in that, The target phase information is normalized phase information; When the target phase information is within the first numerical range [0, a), the target phase information is positively correlated with the smoothing gain; When the target phase information is within the second numerical range [a,1], the target phase information is negatively correlated with the smoothing gain, where a is a positive number.
9. The method according to claim 5, characterized in that, The smoothing point model is constrained by the pitch period, sampling rate, and target phase information.
10. The method according to claim 1, characterized in that, The speech smoothing model includes an active speech smoothing model. The step of determining the number of smoothing points and the smoothing gain based on the target speech smoothing model includes: when the target speech smoothing model is an active speech smoothing model, determining the number of smoothing points and the smoothing gain based on the active speech smoothing model.
11. The method according to claim 1, characterized in that, The method further includes associating multiple speech smoothing models sequentially in a preset order. When the target speech smoothing model is an active speech smoothing model, at least one type of target phase information is modified so that the modified target phase information satisfies the phase condition corresponding to the candidate speech smoothing model, wherein the candidate speech smoothing model is a speech smoothing model located after the target speech smoothing model.
12. The method according to claim 1, characterized in that, The step of smoothly adjusting the amplitude of the target speech sampling point according to the number of smoothing points and the smoothing gain includes: Determine a target smoothing point corresponding to the number of smoothing points in each category. The target smoothing point is located before the target speech sampling point in time. The target speech sampling point is spaced apart from the target smoothing point by a target number of speech sampling points. The target number is equal to the number of smoothing points minus the natural number 1. Based on the smoothing gain and the amplitude of the target smoothing point, the amplitude of the target speech sampling point is smoothly adjusted.
13. The method according to claim 12, characterized in that, Smoothing adjustment of the amplitude of the target speech sampling point based on the smoothing gain and the amplitude of the target smoothing point includes: Calculate the weighted amplitude based on the amplitude of each target smoothing point and the smoothing gain corresponding to each target smoothing point; The amplitude of the target speech sampling point is obtained by combining the weighted amplitudes of multiple target smoothing points.
14. The method according to claim 4, characterized in that, The target phase information includes normalized first phase information and second phase information, and the phase conditions include the following four sets of conditions: First phase condition: 0 ≤ first phase information ≤ 0.5, 0.5 ≤ second phase information ≤ 1; Second phase condition: 0.5 < first phase information ≤ 1, 0 ≤ second phase information ≤ 0.5; Third phase condition: 0 ≤ first phase information < 0.5, 0 ≤ second phase information < 0.5; Fourth phase condition: 0.5 ≤ first phase information ≤ 1, 0.5 ≤ second phase information ≤ 1.
15. A storage medium, characterized in that, The device stores computer-executable instructions for causing the electronic device to perform the speech modulation method as described in any one of claims 1 to 14.
16. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the speech modulation method as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Voice changing processing method, storage medium, chip and electronic equipment
CN115641858A