Audio signal processing method, device, computer equipment and storage medium
Through frame processing and audio signal processing methods of transforming spectrum envelopes, the problems of unsatisfactory sound quality and high delay in the existing voice converter algorithm are solved, real-time adjustable audio signal processing is realized, and audio quality and user experience are improved.
Patent Information
- Application Number
- CN202310095321.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-03
AI Technical Summary
The existing voice converter algorithms have problems such as poor sound quality, high delay or fixed tone in virtual anchors and metaverse applications, making it difficult to achieve real-time adjustable audio signal processing.
By obtaining the tone and timbre parameters of the original audio stream, moving the fundamental frequency and changing the spectrum envelope after frame processing, combining the real-time fundamental tone synchronization superposition algorithm and phase vocoder algorithm, real-time adjustment of tone and timbre is achieved.
Real-time adjustable processing of audio signals is realized, reducing hearing distortion, improving sound quality, improving user experience, and having less calculations.
Smart Images

Figure CN116092509B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technologies, and in particular, to an audio signal processing method, apparatus, computer device, and computer-readable storage medium. Background Art
[0002] In application scenarios such as virtual anchors and the metaverse, users have the need to beautify their voices and conceal their identities using voice changers. Currently, the voice changer algorithms on the market mainly fall into two categories: signal processing algorithms and deep learning algorithms. The former has a fast calculation speed, low latency, and adjustable parameters, but the sound quality is often not ideal and there is distortion in the listening experience; the latter has a listening experience close to natural human voices, but there are problems such as being unable to process audio signals in real time or having a high latency, and the timbre of the output voice signal is also often fixed. Summary of the Invention
[0003] The main objective of this application is to propose an audio signal processing method, apparatus, computer device, and computer-readable storage medium, aiming to solve the problem of how to perform real-time adjustable audio signal processing on human voices and ensure a better listening experience of the output sound.
[0004] To achieve the above objective, an embodiment of this application provides an audio signal processing method, and the method includes:
[0005] Obtain an original audio stream and predetermined parameters for the original audio stream, where the predetermined parameters include a pitch change parameter and a timbre change parameter;
[0006] Sample the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division on the sampling points to obtain a plurality of first speech frames;
[0007] Determine the fundamental frequency of each frame of the first speech frames, and move the fundamental frequency of each frame of the first speech frames according to the pitch change parameter;
[0008] Overlay and splice each frame of the first speech frames after the movement to obtain an input audio stream;
[0009] Divide the input audio stream into a plurality of second speech frames of the same length, and transform the spectral envelope of each frame of the second speech frames, where two adjacent frames of the second speech frames have an overlapping part;
[0010] Splice each frame of the second speech frames after the spectral envelope transformation to obtain an output audio stream, and resample the output audio stream according to the timbre change parameter to obtain a target audio stream.
[0011] Optionally, moving the fundamental frequency of each frame of the first speech frames according to the pitch change parameter includes:
[0012] According to the pitch-shifting parameter, shift the fundamental frequency of each of the input speech frames through a real-time pitch synchronous overlap-add algorithm.
[0013] Optionally, the shifting of the fundamental frequency of each of the first speech frames according to the pitch-shifting parameter through a real-time pitch synchronous overlap-add algorithm includes:
[0014] Obtain the fundamental frequency of each of the first speech frames;
[0015] Decompose each of the first speech frames according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments;
[0016] Extract speech segments from each of the pitch period segments by using a window function;
[0017] Copy and shift the speech segments according to the pitch-shifting parameter to obtain processed speech segments, and perform an overlap-add process on the extracted speech segments and the processed speech segments to obtain each of the first speech frames after shifting.
[0018] Optionally, the decomposing of the first speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments includes:
[0019] According to the period corresponding to the fundamental frequency, identify the maximum value in each period, and form a pitch period segment with the current maximum value as the center and the two adjacent maximum values before and after as the boundaries.
[0020] Optionally, the copying and shifting of the speech segments according to the pitch-shifting parameter to obtain processed speech segments, and performing an overlap-add process on the extracted speech segments and the processed speech segments to obtain each of the first speech frames after shifting includes:
[0021] Convert the fundamental frequency to a musical scale;
[0022] Perform discretization processing on the musical scale with the pitch-shifting parameter as the granularity of the discretization processing to obtain a discretized musical scale;
[0023] Convert the discretized musical scale back to the corresponding frequency value, and obtain a fundamental frequency shifting ratio according to the frequency value and the fundamental frequency;
[0024] Copy and shift the speech segments according to the fundamental frequency shifting ratio to obtain processed speech segments, and perform an overlap-add process on the extracted speech segments and the processed speech segments to obtain each of the first speech frames after shifting.
[0025] Optionally, the transforming of the spectral envelope of each of the second speech frames includes:
[0026] Transform the spectral envelope of each frame of the second speech frame through the phase vocoder algorithm.
[0027] Optionally, the transforming the spectral envelope of each frame of the second speech frame through the phase vocoder algorithm includes:
[0028] Translate the position of each frame of the second speech frame on the time axis according to the timbre change parameter;
[0029] Perform phase reconstruction on each frame of the second speech frame after translation, and use each frame of the second speech frame after phase reconstruction as each frame of the second speech frame after being transformed by the spectral envelope.
[0030] Optionally, the performing phase reconstruction on each frame of the second speech frame after translation includes:
[0031] Obtain each frame of the second speech frame;
[0032] Calculate the spectra of the m-th frame and the (m + 1)-th frame of the second speech frame respectively through short-time Fourier transform, and determine the instantaneous frequency and phase of the m-th frame and the (m + 1)-th frame of the second speech frame according to the spectra, where m is an integer greater than or equal to 1;
[0033] Calculate the phase of the (m + 1)-th frame of the second speech frame after translation according to the instantaneous frequency and phase of the m-th frame of the second speech frame and the time interval before and after the movement of each frame of the second speech frame;
[0034] Replace the phase of the spectrum of the (m + 1)-th frame of the second speech frame with the calculated phase;
[0035] Convert the spectrum of the (m + 1)-th frame of the second speech frame after phase replacement into a time-domain signal through inverse fast Fourier transform, and use the time-domain signal as the (m + 1)-th frame of the second speech frame after being transformed by the spectral envelope.
[0036] Optionally, the translating the position of each frame of the second speech frame on the time axis according to the timbre change parameter includes:
[0037] Obtain the first number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame;
[0038] Determine the second number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame after translation according to the timbre change parameter and the first number;
[0039] Translate the position of each frame of the second speech frame on the time axis according to the second number.
[0040] In addition, to achieve the above object, an embodiment of the present application further provides an audio signal processing device, which includes:
[0041] An acquisition module, configured to acquire an original audio stream and predetermined parameters of the original audio stream, where the predetermined parameters include a pitch change parameter and a timbre change parameter;
[0042] A framing module, configured to sample the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform framing processing on the sampling points to obtain a plurality of first speech frames;
[0043] A shifting module, configured to determine the fundamental frequency of each frame of the first speech frames, and shift the fundamental frequency of each frame of the first speech frames according to the pitch change parameter;
[0044] A splicing module, configured to splice each frame of the shifted first speech frames to obtain an input audio stream;
[0045] A transformation module, configured to divide the input audio stream into multiple second speech frames of the same length, and transform the spectral envelope of each frame of the second speech frames, where two adjacent frames of the second speech frames have an overlapping part;
[0046] A resampling module, configured to stack and splice each frame of the second speech frames after the spectral envelope transformation to obtain an output audio stream, and resample the output audio stream according to the timbre change parameter to obtain a target audio stream.
[0047] To achieve the above object, an embodiment of the present application further provides a computer device, which includes: a memory, a processor, and an audio signal processing program stored on the memory and executable on the processor. When the audio signal processing program is executed by the processor, the audio signal processing method as described above is implemented.
[0048] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, on which an audio signal processing program is stored. When the audio signal processing program is executed by a processor, the audio signal processing method as described above is implemented.
[0049] The audio signal processing method, apparatus, computer device, and computer-readable storage medium proposed in the embodiments of the present application analyze the signal characteristics of the original audio stream input by the user frame by frame according to the pitch change parameter, so as to realize the translation of the fundamental frequency of each frame of speech frame, thereby changing the pitch of the speech and reducing the distortion in the sense of hearing. At the same time, by transforming the spectral envelope of the speech after pitch change, the timbre of the speech is changed, thereby improving the sound quality of the output audio, enabling the user to have a better sense of hearing of the output audio, and further improving the user experience during the voice conversion process. In addition, the audio signal processing method in the present application can complete the changes in pitch and timbre with less computational effort and has a fast feedback speed, so that the audio signal processing method in the present application can also support the audio signal processing of real-time speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 An application environment architecture diagram for implementing various embodiments of the present application;
[0051] Figure 2 A flowchart of an audio signal processing method proposed in an embodiment of the present application;
[0052] Figure 3 A detailed process schematic diagram for moving the fundamental frequency of each frame of the first speech frame through the real-time pitch synchronous overlap algorithm according to the pitch change parameter in an embodiment of the present application;
[0053] Figure 4 A schematic diagram of a pitch period segment in the first embodiment of the present application;
[0054] Figure 5 According to Figure 4 The extracted speech segment and the fundamental frequency movement schematic diagram;
[0055] Figure 6 A detailed process schematic diagram for copying and moving the speech segment according to the pitch change parameter in an embodiment of the present application to obtain a processed speech segment, and performing superposition processing on the extracted speech segment and the processed speech segment to obtain each frame of the first speech frame after movement;
[0056] Figure 7 A detailed process schematic diagram for transforming the spectral envelope of each frame of the second speech frame through the phase vocoder algorithm in an embodiment of the present application;
[0057] Figure 8 A detailed process schematic diagram for translating the position of each frame of the second speech frame on the time axis according to the timbre change parameter in an embodiment of the present application;
[0058] Figure 9Schematic diagram of the refinement process for phase reconstruction of each second voice frame after translation
[0059] Figure 10 Schematic diagram of the processing process for processing the input audio stream to obtain the output audio stream
[0060] Figure 11 Schematic diagram of the program module of an audio signal processing device proposed in an embodiment of the present application
[0061] Figure 12 Schematic diagram of the hardware architecture of an electronic device proposed in an embodiment of the present application Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0063] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0064] Please refer to Figure 1 , Figure 1 Application environment architecture diagram for implementing various embodiments of the present application
[0065] In an exemplary embodiment, the system of the application environment may include a computer device 10 and a server 20. Among them, the computer device 10 and the server 20 are connected through a wireless or wired network. The computer device 10 may be, for example, a smart phone, a tablet device, a laptop computer, a smart TV, a vehicle-mounted terminal, etc. The server 20 may be a rack server, a blade server, a tower server or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. The network may include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls and / or proxy devices, etc. The network may also include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links and their combinations and / or analogs.
[0066] Next, several embodiments will be provided in the above exemplary application environment to illustrate the audio signal processing solution in the present application.
[0067] As Figure 2 shown, it is a flowchart of an audio signal processing method proposed in the first embodiment of the present application. It can be understood that the flowchart in the embodiment of the present method is not used to limit the order of execution steps. According to needs, some steps in the flowchart can also be added or deleted.
[0068] The method includes the following steps:
[0069] S20, obtain the original audio stream and the predetermined parameters of the original audio stream, where the predetermined parameters include a pitch change parameter and a timbre change parameter.
[0070] Specifically, the original audio stream may be an audio stream collected in real time, or an audio stream imported by the user through an audio import interface.
[0071] The pitch change parameter is a parameter used to represent the strength of the voice pitch, which can be represented by a and is used to change the pitch of the voice subsequently.
[0072] The timbre change parameter is a parameter used to represent the degree of adjustment of the spectral envelope of the voice, which can be represented by b and is used to change the timbre of the voice subsequently.
[0073] It should be noted that the voice change parameter may be a default parameter or a parameter input by the user.
[0074] S21, sample the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform frame division processing on the sampling points to obtain a plurality of first speech frames.
[0075] Specifically, after obtaining the original audio stream, the voice signal in the original audio stream needs to be sampled first. The sampling rate, also known as the sampling speed or sampling frequency, refers to the number of samples extracted from a continuous signal (analog signal) and composed into a discrete signal per unit time, which is expressed in Hertz (Hz). In this embodiment, the original audio stream can be sampled at any feasible sampling rate to obtain a series of discrete sampling points. Generally speaking, the commonly used sampling rates are mainly 44100 Hz (Hertz) and 48000 Hz. After sampling the original audio stream, frame division processing can be performed on these sampling points, that is, a predetermined number of sampling points are used as a first voice frame. For example, 1024 sampling points are used as a first voice frame. After performing frame division processing on the sampling points, multiple first voice frames can be obtained.
[0076] S22, determine the fundamental frequency of each frame of the first voice frame, and shift the fundamental frequency of each frame of the first voice frame according to the pitch change parameter.
[0077] Specifically, the fundamental frequency refers to the most basic vibration frequency in the voice signal, which represents the degree to which the human ear can distinguish the pitch of a sound in terms of auditory perception. In this embodiment, the methods for measuring the fundamental frequency of a voice frame can be roughly divided into the time domain method and the frequency domain method. Among them, the time domain method takes the waveform of the sound as the input, and its basic principle is to find the smallest positive period of the waveform. Of course, the periodicity of the actual signal can only be approximate. The frequency domain method will first perform a Fourier transform on the signal to obtain the frequency spectrum (only taking the magnitude spectrum and discarding the phase spectrum). There will be spikes at integer multiples of the fundamental frequency on the frequency spectrum, and the basic principle of the frequency domain method is to find the greatest common divisor of these spike frequencies.
[0078] In one embodiment, after obtaining multiple first voice frames, for each frame of the first voice frame, the fundamental frequency of the first voice frame can be measured by means such as the time domain autocorrelation algorithm or the frequency domain harmonic-to-noise ratio algorithm. In other embodiments, any other feasible existing method can also be used to measure the fundamental frequency, which is not limited here.
[0079] In one embodiment, according to the pitch change parameter, the fundamental frequency of each frame of the first voice frame can be shifted by the real-time pitch synchronous overlap-add algorithm. During the shifting process, the pitch change parameter can be directly used as the fundamental frequency shifting ratio to shift the fundamental frequency of each frame of the first voice frame.
[0080] In the existing Pitch Synchronous Overlap and Add (PSOLA) algorithm, generally the fundamental frequency of the entire audio segment is shifted, so real-time processing cannot be achieved. In this embodiment, through the improved real-time PSOLA algorithm, it is possible to decompose the fundamental frequency period into segments based on each of the input speech frames, and then shift the fundamental frequency of each segment, rather than shifting the entire audio segment, thereby achieving real-time processing of the audio signal.
[0081] In this embodiment, the basic idea of the real-time PSOLA algorithm is to decompose the speech signal into segments corresponding to the fundamental frequency period, and then perform operations such as copying and translation on these segments. Since the short-term correlation in the signal is not changed, that is, the signal within the segment does not change, the spectral envelope of the signal is maintained, and only the fundamental frequency of the signal is changed, so the quality of the speech signal is retained to the greatest extent.
[0082] It should be noted that the shifting of the fundamental frequency in this embodiment refers to a processing method of changing the frequency of the fundamental frequency part in the signal through an algorithm without changing the frequency of other harmonic parts.
[0083] In an exemplary embodiment, referring to Figure 3 , the shifting of the fundamental frequency of each of the first speech frames through the real-time pitch synchronous overlap and add algorithm according to the pitch change parameter includes:
[0084] Step S30, obtaining the fundamental frequency of each of the first speech frames.
[0085] Step S31, decomposing each of the first speech frames according to the period corresponding to the fundamental frequency to obtain a plurality of fundamental period segments.
[0086] Specifically, in order to improve the processing speed of the original audio stream, after performing frame segmentation on the original audio stream, it is possible to continue to segment each of the first speech frames according to the fundamental period to obtain a plurality of fundamental period segments. In this way, when shifting each of the first speech frames, it becomes shifting for each fundamental period segment.
[0087] In an exemplary embodiment, when segmenting each of the first speech frames, it is possible to identify the maximum value in each period according to the period corresponding to the fundamental frequency, and then use the current maximum value as the center, and the two adjacent maximum values before and after as the boundaries to form a fundamental period segment.
[0088] As Figure 4 shown, it is a schematic diagram of the fundamental period segment in this embodiment. In Figure 4 , the maximum value identified in each period is represented by an x point, and then using each x point as the center and the adjacent x points on the left and right as the boundaries is a fundamental period segment.
[0089] Step S32: Use a window function to extract voice segments from each frame of the pitch period segments.
[0090] Specifically, the window function can be a rectangular window, a triangular window, a Hanning window, a Hamming window, etc., which is not limited in this embodiment. Among them, the window length of the window function can be set according to the actual situation or can be the default window length. For example, the window length is half of the length of the pitch period segment.
[0091] In this embodiment, by applying the window function to the pitch period segments, the voice segments corresponding to each pitch period segment can be extracted. As Figure 5 shown, it is a schematic diagram of the voice segments and fundamental frequency shift extracted according to the input pitch period segments. In Figure 5 , from the topmost pitch period segment (original voice signal), several voice segments can be extracted, as Figure 5 shown in the left row in the middle.
[0092] Step S33: Copy and move the voice segments according to the pitch shift parameter to obtain the processed voice segments, and perform superposition processing on the extracted voice segments and the processed voice segments to obtain each frame of the first voice frame after movement.
[0093] Specifically, using the pitch shift parameter as the fundamental frequency shift multiple, copy and perform fundamental frequency shift on each voice segment, and then the processed voice segments of each voice segment can be obtained. As an example, when copying and moving multiple voice segments as Figure 5 shown in the middle, the moved segments corresponding to these voice segments can be obtained, as Figure 5 shown in the right row in the middle.
[0094] In this embodiment, after obtaining the processed voice segments of all the voice segments extracted from each pitch period segment of the current first voice frame, perform superposition processing on the voice segments extracted from all the pitch period segments and their processed voice segments, and then the movement result of the current first voice frame can be obtained. As an example, when superposing multiple voice segments as Figure 5 shown in the middle and the corresponding voice segments after copying and moving processing, the movement result of the input pitch period segments can be obtained, as Figure 5 shown in the lower part. Concatenate the movement results of all the pitch period segments, and then the movement result of the fundamental frequency of the first voice frame can be obtained.
[0095] In an exemplary embodiment, in order to perform better pitch shift on the voice, refer to Figure 6, copying and moving the speech segment according to the pitch change parameter to obtain a processed speech segment, and performing superposition processing on the extracted speech segment and the processed speech segment, so that each frame of the first speech frame after movement includes:
[0096] Step S60: Convert the fundamental frequency into a musical scale.
[0097] Specifically, the musical scale refers to a series of tones from low to high, increasing according to an exponential relationship. In this embodiment, the musical scale can be represented by MIDI (Musical Instrument Digital Interface) values. In the currently common twelve-tone system in the music field, there is an exponential function relationship between the musical scale and the corresponding frequency. Specifically, for every twelve musical scales raised, the pitch is raised by one octave, and the corresponding frequency becomes twice the original, that is, the frequency ratio of two adjacent musical scales is 2^(1 / 12) = 1.05946. At the same time, in music theory, it is stipulated that the musical scale number corresponding to middle C (C4) is 60, and the corresponding frequency is 261.6 Hz. Therefore, the conversion formula from the frequency value to the musical scale value is: MIDI = 60 + 12 * log2(F / 261.6), where F represents the frequency value, that is, the fundamental frequency of each frame of the first speech frame, and the unit is Hz. According to the above formula, the fundamental frequency of each frame of the first speech frame can be converted into a MIDI value, that is, converted into a musical scale.
[0098] Step S61: Use the pitch change parameter as the granularity of discretization processing to perform discretization processing on the musical scale to obtain a discretized musical scale.
[0099] Specifically, in the traditional tuning scheme, vocal electroacoustic music generally realizes the discretization processing of the musical scale by converting the frequency of the input speech signal to the nearest integer MIDI value. For example, assume that the fundamental frequency of the input speech frame is 250 Hz, and the corresponding MIDI value is 59.215, which can be converted to the nearest integer MIDI value 59, that is, at 246.9 Hz.
[0100] In this embodiment, in order to improve the granularity of scale discretization, when performing the discretization process on the scale, the variable pitch parameter can be used as the granularity of the discretization process to discretize the scale, and the discretized scale is obtained. Specifically, the variable pitch parameter a can be used as the granularity of the discretization process. After converting the fundamental frequency to the scale MIDI value, the formula for scale discretization is a * round(MIDI / a), where the round function represents rounding, that is, taking an integer. For example, when a = 1, it represents converting the frequency of the human voice speech signal to the nearest integer MIDI value in the traditional tuning scheme. Assuming that the fundamental frequency of the first speech frame is 250 Hz, the corresponding MIDI value is 59.215. Discretizing it with a = 1 as the granularity and converting it to the nearest integer MIDI value is 59, that is, at 246.9 Hz. When a = 2.5, discretizing it with 2.5 as the granularity and converting it to an integer MIDI value is 2.5 * round(59.215 / 2.5) = 60, that is, at 261.6 Hz. The larger the value of the variable pitch parameter a, the stronger the electro - music effect in the resulting voice - changing result.
[0101] Step S62: Convert the discretized scale back to the corresponding frequency value, and obtain the fundamental - frequency shift ratio according to the frequency value and the fundamental frequency.
[0102] Specifically, after discretizing the scale with the predetermined parameter a as the granularity, according to the above conversion formula in reverse, the scale value can be converted back to the corresponding frequency value. Then, taking the ratio of the frequency value and the fundamental frequency of the first speech frame, the corresponding fundamental - frequency shift ratio can be obtained. That is, the fundamental - frequency shift ratio of each frame of the first speech frame = frequency value / fundamental frequency.
[0103] Step S63: Copy and move the speech segment according to the fundamental - frequency shift ratio to obtain the processed speech segment, and perform superposition processing on the extracted speech segment and the processed speech segment to obtain each frame of the first speech frame after movement.
[0104] In this embodiment, by using the ratio of frequency value / fundamental frequency as the fundamental - frequency shift ratio instead of directly using the variable pitch parameter as the fundamental - frequency shift ratio, each frame of the first speech frame after the final movement process has a better pitch - changing result, and the distortion in the listening experience can be further reduced.
[0105] It should be noted that the fundamental - frequency shift ratio is the ratio for moving the speech segment. To facilitate understanding of how to copy and move the speech segment according to the fundamental - frequency shift ratio, the following will be described in detail with a specific example.
[0106] For example, the fundamental frequency shift multiple is 2, and there are a total of 6 voice segments, namely voice segment 1 containing sampling points 1 - 200, voice segment 2 containing sampling points 101 - 300, voice segment 3 containing sampling points 201 - 400, voice segment 4 containing sampling points 301 - 500, voice segment 5 containing sampling points 401 - 600, and voice segment 6 containing sampling points 501 - 700. After copying and moving the voice segments according to the fundamental frequency shift multiple, the 6 obtained voice segments are respectively voice segment 7 containing sampling points 1 - 200, voice segment 8 containing sampling points 201 - 400, voice segment 9 containing sampling points 401 - 600, voice segment 10 containing sampling points 601 - 800, voice segment 11 containing sampling points 801 - 1000, and voice segment 12 containing sampling points 1001 - 1200.
[0107] S23, stack and splice each frame of the moved first voice frame to obtain an input audio stream.
[0108] Specifically, after obtaining each frame of the moved first voice frame, re - splice all the voice frames frame by frame and output, then the input audio stream can be obtained.
[0109] S24, divide the input audio stream into multiple second voice frames of the same length, and transform the spectral envelope of each frame of the second voice frame, where two adjacent second voice frames have an overlapping part.
[0110] Specifically, after obtaining the input audio stream, first sample the voice signal in the input audio stream to obtain a series of discrete sampling points. Then, divide this series of discrete sampling points into multiple second voice frames of the same length. Among them, in the process of dividing into multiple second voice frames, two adjacent second voice frames have an overlapping part, and the number of sampling points included in the overlapping part can be set according to the actual situation. For example, the number of sampling points included in the overlapping part is 10.
[0111] In one embodiment, the spectral envelope of each frame of the second voice frame can be transformed by the phase vocoder algorithm.
[0112] The phase vocoder algorithm is an algorithm for transforming the spectral envelope of a voice signal. Among them, the transformation of the spectral envelope refers to changing the energy distribution of the fundamental frequency and harmonics of the audio frequency signal in the frequency domain.
[0113] In an exemplary embodiment, refer to Figure 7 , the transformation of the spectral envelope of each frame of the second voice frame by the phase vocoder algorithm includes:
[0114] Step S70, translate the position of each frame of the second speech frame on the time axis according to the voice color change parameter.
[0115] Specifically, translate the position of each frame of the second speech frame on the time axis according to the voice color change parameter to change the size of the overlapping part, that is, change the number of sampling points included in the overlapping part.
[0116] In an exemplary embodiment, refer to Figure 8 , the translating the position of each frame of the second speech frame on the time axis according to the voice color change parameter includes:
[0117] Step S80, obtain the first number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame.
[0118] Step S81, determine the second number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame after translation according to the voice color change parameter and the first number.
[0119] Specifically, the product of the voice color change parameter and the first number can be used as the second number, that is, the second number = the voice color change parameter * the first number. In one embodiment, the first number divided by the voice color change parameter and then added to the first number can also be used as the second number, that is, the second number = the first number / the voice color change parameter + the first number.
[0120] Step S82, translate the position of each frame of the second speech frame on the time axis according to the second number.
[0121] Specifically, during the process of translating each frame of the second speech frame, the first speech frame is not translated.
[0122] As an example, there are a total of 5 second speech frames, namely speech frame 1 containing sampling points 0 - 20, speech frame 2 containing sampling points 11 - 30, speech frame 3 containing sampling points 21 - 40, speech frame 4 containing sampling points 31 - 50, and speech frame 5 containing sampling points 41 - 60. If the second number is 15, then the speech frames obtained after translating these 5 second speech frames are speech frame 6 containing sampling points 0 - 20, speech frame 7 containing sampling points 6 - 25, speech frame 8 containing sampling points 11 - 30, speech frame 9 containing sampling points 16 - 35, and speech frame 10 containing sampling points 21 - 40.
[0123] Step S71: Perform phase reconstruction on each of the translated second speech frames, and use each of the second speech frames after phase reconstruction as each of the second speech frames after being transformed by the spectral envelope.
[0124] Specifically, in order to ensure the continuity of the phase between each of the second speech frames after the translation process, so that the finally obtained target audio stream has better sound quality. In this embodiment, after obtaining each of the translated second speech frames, it is necessary to perform phase reconstruction on each of the second speech frames.
[0125] In an exemplary embodiment, refer to Figure 9 , the performing phase reconstruction on each of the translated second speech frames includes: Step S90: Obtain each of the second speech frames; Step S91: Calculate the spectra of the m-th second speech frame and the (m + 1)-th second speech frame respectively through short-time Fourier transform, and determine the instantaneous frequency and phase of the m-th second speech frame and the (m + 1)-th second speech frame according to the spectra, where m is an integer greater than or equal to 1; Step S92: Calculate the phase of the (m + 1)-th second speech frame after translation according to the instantaneous frequency and phase of the m-th second speech frame and the time interval before and after the movement of each of the second speech frames; Step S93: Replace the phase of the spectrum of the (m + 1)-th second speech frame with the calculated phase; Step S94: Convert the spectrum of the (m + 1)-th second speech frame after phase replacement into a time-domain signal through inverse fast Fourier transform, and use the time-domain signal as the (m + 1)-th second speech frame after being transformed by the spectral envelope.
[0126] Specifically, the time interval before and after the movement of each of the second speech frames can be determined according to the timestamps corresponding to the second speech frames before and after the movement. For example, if the timestamp corresponding to the second speech frame before the movement is T1, and the timestamp corresponding to the second speech frame after the movement is T2, then the time interval = T2 - T1.
[0127] The phase of the (m + 1)-th second speech frame after translation = 2πf(T2 - T1)+X, where f is the instantaneous frequency of the m-th second speech frame, and X is the phase of the m-th second speech frame.
[0128] S25: Concatenate each of the second speech frames after being transformed by the spectral envelope to obtain an output audio stream, and resample the output audio stream according to the variable timbre parameter to obtain a target audio stream.
[0129] Specifically, after obtaining each second speech frame transformed by the spectral envelope, all the speech frames are re - stitched frame by frame and then output, and thus the output audio stream can be obtained. Among them, the pitch and spectral envelope of the output audio stream are consistent with those of the input audio stream, but the total length of the audio has changed.
[0130] As Figure 10 shown, it is a schematic diagram of the processing process for obtaining the output audio stream by processing the input audio stream. Among them, Figure 10 the first figure is a schematic diagram of the input audio stream, the second figure is a schematic diagram of obtaining multiple second speech frames after dividing the input audio stream, the third figure is a schematic diagram of the speech frames obtained after transforming the spectral envelopes of each second speech frame, and the fourth figure is a schematic diagram of the output audio stream.
[0131] After obtaining the output audio stream, the output audio stream can be resampled according to the variable timbre parameter, so that the length of the target audio stream is the same as that of the original audio stream. At the same time, the spectral envelope of the target audio stream is compressed or stretched in the frequency domain, so that the timbre of the target audio stream changes compared with the original audio stream in terms of listening perception.
[0132] In this embodiment, based on the variable pitch parameter, the signal characteristics of the original audio stream input by the user are analyzed frame by frame to realize the translation of the fundamental frequency of each speech frame, thereby changing the pitch of the speech and reducing the distortion in listening perception. At the same time, by transforming the spectral envelope of the speech after pitch - shifting, the timbre of the speech is changed, so as to improve the sound quality of the output audio, enabling the user to have a better listening experience for the output audio, and further improving the user experience during the voice - changing process. In addition, the audio signal processing method in this application can complete the changes in pitch and timbre with less computational complexity and has a fast feedback speed, so that the audio signal processing method in this application can also support the audio signal processing of real - time speech.
[0133] Refer to Figure 11 shown, it is a program module diagram of an embodiment of the audio signal processing device 110 of this application.
[0134] In this embodiment, the audio signal processing device 110 includes a series of computer program instructions stored in the memory. When these computer program instructions are executed by the processor, the audio signal processing functions of each embodiment of this application can be realized. In some embodiments, based on the specific operations implemented by each part of these computer program instructions, the audio signal processing device 110 can be divided into one or more modules, and the specific modules that can be divided are as follows:
[0135] An acquisition module 111, configured to acquire an original audio stream and predetermined parameters of the original audio stream, where the predetermined parameters include a pitch-changing parameter and a timbre-changing parameter.
[0136] A framing module 112, configured to sample the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform framing processing on the sampling points to obtain a plurality of first speech frames.
[0137] A shifting module 113, configured to determine the fundamental frequency of each frame of the first speech frames, and shift the fundamental frequency of each frame of the first speech frames according to the pitch-changing parameter.
[0138] An overlay splicing module 114, configured to perform overlay splicing on each shifted frame of the first speech frames to obtain an input audio stream.
[0139] A transformation module 115, configured to divide the input audio stream into multiple second speech frames of the same length, and transform the spectral envelope of each frame of the second speech frames, where two adjacent frames of the second speech frames have an overlapping part.
[0140] A resampling module 116, configured to splice each frame of the second speech frames after the spectral envelope transformation to obtain an output audio stream, and resample the output audio stream according to the timbre-changing parameter to obtain a target audio stream.
[0141] In an exemplary embodiment, the shifting module 113 is further configured to shift the fundamental frequency of each input speech frame by a real-time pitch synchronous overlap-add algorithm according to the pitch-changing parameter.
[0142] In an exemplary embodiment, the shifting module 113 is further configured to obtain the fundamental frequency of each frame of the first speech frames; decompose each frame of the first speech frames according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments; extract speech segments from each frame of the pitch period segments by using a window function; copy and shift the speech segments according to the pitch-changing parameter to obtain processed speech segments, and perform an overlay process on the extracted speech segments and the processed speech segments to obtain each shifted frame of the first speech frames.
[0143] In an exemplary embodiment, the shifting module 113 is further configured to identify the maximum value in each frame period according to the period corresponding to the fundamental frequency, and form a pitch period segment with the current maximum value as the center and the two adjacent maximum values before and after as the boundaries.
[0144] In an exemplary embodiment, the mobile module 113 is further configured to convert the fundamental frequency into a musical scale; discretize the musical scale by using the pitch shift parameter as the granularity of the discretization process to obtain a discretized musical scale; convert the discretized musical scale back into a corresponding frequency value, obtain a fundamental frequency shift ratio based on the frequency value and the fundamental frequency; copy and move the speech segment according to the fundamental frequency shift ratio to obtain a processed speech segment, and perform a superimposition process on the extracted speech segment and the processed speech segment to obtain each frame of the first speech frame after movement.
[0145] In an exemplary embodiment, the transformation module 115 is further configured to transform the spectral envelope of each frame of the second speech frame through a phase vocoder algorithm.
[0146] In an exemplary embodiment, the transformation module 115 is further configured to translate the position of each frame of the second speech frame on the time axis according to the timbre change parameter; perform phase reconstruction on each frame of the second speech frame after translation, and use each frame of the second speech frame after phase reconstruction as each frame of the second speech frame after being transformed by the spectral envelope.
[0147] In an exemplary embodiment, the transformation module 115 is further configured to obtain each frame of the second speech frame; calculate the spectra of the m-th frame of the second speech frame and the (m + 1)-th frame of the second speech frame respectively through short-time Fourier transform, and determine the instantaneous frequency and phase of the m-th frame of the second speech frame and the (m + 1)-th frame of the second speech frame according to the spectra, where m is an integer greater than or equal to 1; calculate the phase of the (m + 1)-th frame of the second speech frame after translation according to the instantaneous frequency and phase of the m-th frame of the second speech frame and the time interval before and after the movement of each frame of the second speech frame; replace the phase of the spectrum of the (m + 1)-th frame of the second speech frame with the calculated phase; convert the spectrum of the (m + 1)-th frame of the second speech frame after phase replacement into a time-domain signal through inverse fast Fourier transform, and use the time-domain signal as the (m + 1)-th frame of the second speech frame after being transformed by the spectral envelope.
[0148] In an exemplary embodiment, the transformation module 115 is further configured to obtain the first number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame; determine the second number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame after translation according to the timbre change parameter and the first number; translate the position of each frame of the second speech frame on the time axis according to the second number.
[0149] In this embodiment, based on the pitch change parameter, the signal characteristics of the original audio stream input by the user are analyzed frame by frame to achieve the translation of the fundamental frequency of each frame of speech frame, thereby changing the pitch of the speech and reducing the distortion in the auditory sense. At the same time, by transforming the spectral envelope of the speech after pitch change, the timbre of the speech is changed, thereby improving the sound quality of the output audio, enabling the user to have a better auditory sense of the output audio, and further improving the user experience during the voice conversion process. In addition, the audio signal processing method in this application can complete the changes in pitch and timbre with less computational effort and has a fast feedback speed, so that the audio signal processing method in this application can also support the audio signal processing of real-time speech.
[0150] As Figure 12 shown, the figure is a schematic diagram of the hardware architecture of a computer device 12 proposed in an embodiment of this application. In this embodiment, the computer device 12 may include, but is not limited to, a memory 21, a processor 22, and a network interface 23 that are communicatively connected to each other through a system bus. It should be noted that Figure 12 only the computer device 12 with components 21-23 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. In this embodiment, the computer device 12 can be the client or the server.
[0151] The memory 21 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 21 can be an internal storage unit of the computer device 12, such as the hard disk or memory of the computer device 12. In other embodiments, the memory 21 can also be an external storage device of the computer device 12, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 12. Of course, the memory 21 can also include both the internal storage unit and the external storage device of the computer device 12. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed on the computer device 12, such as the program code of the audio signal processing device 110. In addition, the memory 21 can also be used to temporarily store various data that have been output or will be output.
[0152] In some embodiments, the processor 22 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 22 is generally used to control the overall operation of the computer device 12. In this embodiment, the processor 22 is used to run the program code stored in the memory 21 or process data, such as running the audio signal processing device 110, etc.
[0153] The network interface 23 may include a wireless network interface or a wired network interface, and this network interface 23 is generally used to establish a communication connection between the computer device 12 and other electronic devices.
[0154] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing an audio signal processing program, and the audio signal processing program can be executed by at least one processor, so that the at least one processor executes the steps of the audio signal processing method as described above.
[0155] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0156] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0157] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device, so that they can be stored in the storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0158] The above are only the preferred embodiments of the embodiments of the present application, and do not limit the patent scope of the embodiments of the present application. Any equivalent structural or equivalent process transformation made by using the description and drawings of the embodiments of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the embodiments of the present application.
Claims
1. An audio signal processing method, characterized in that The method includes: Obtaining an original audio stream and predetermined parameters of the original audio stream, where the predetermined parameters include a pitch-changing parameter and a timbre-changing parameter; Sampling the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and performing frame division on the sampling points to obtain a plurality of first speech frames; Determining the fundamental frequency of each frame of the first speech frame, and shifting the fundamental frequency of each frame of the first speech frame according to the pitch-changing parameter; Performing superposition and splicing on each shifted frame of the first speech frame to obtain an input audio stream; Dividing the input audio stream into a plurality of second speech frames of the same length, and transforming the spectral envelope of each frame of the second speech frame, where adjacent two frames of the second speech frames have an overlapping part; Performing splicing on each frame of the second speech frame after the spectral envelope transformation to obtain an output audio stream, and resampling the output audio stream according to the timbre-changing parameter to obtain a target audio stream.
2. The audio signal processing method according to claim 1, wherein Shifting the fundamental frequency of each frame of the first speech frame according to the pitch-changing parameter includes: Shifting the fundamental frequency of each frame of the first speech frame according to the pitch-changing parameter by a real-time pitch synchronous overlap-add algorithm.
3. The audio signal processing method according to claim 2, wherein Shifting the fundamental frequency of each frame of the first speech frame according to the pitch-changing parameter by a real-time pitch synchronous overlap-add algorithm includes: Obtaining the fundamental frequency of each frame of the first speech frame; Decomposing each frame of the first speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments; Extracting speech segments from each frame of the pitch period segments by using a window function; Copying and shifting the speech segments according to the pitch-changing parameter to obtain processed speech segments, and performing superposition processing on the extracted speech segments and the processed speech segments to obtain each shifted frame of the first speech frame.
4. The audio signal processing method according to claim 3, wherein Decomposing the first speech frame according to the period corresponding to the fundamental frequency to obtain a plurality of pitch period segments includes: Identifying the maximum value in each period according to the period corresponding to the fundamental frequency, and forming a pitch period segment with the current maximum value as the center and the two adjacent maximum values before and after as the boundaries.
5. The audio signal processing method according to claim 3, wherein Copying and shifting the speech segments according to the pitch-changing parameter to obtain processed speech segments, and performing superposition processing on the extracted speech segments and the processed speech segments to obtain each shifted frame of the first speech frame includes: Converting the fundamental frequency to a musical scale; Performing discretization processing on the musical scale with the pitch-changing parameter as the granularity of the discretization processing to obtain a discretized musical scale; Converting the discretized musical scale back to the corresponding frequency value, and obtaining a fundamental frequency shifting magnification according to the frequency value and the fundamental frequency; Copying and shifting the speech segments according to the fundamental frequency shifting magnification to obtain processed speech segments, and performing superposition processing on the extracted speech segments and the processed speech segments to obtain each shifted frame of the first speech frame.
6. The audio signal processing method according to any one of claims 1 to 5, characterized in that, Transforming the spectral envelope of each frame of the second speech frame includes: Transforming the spectral envelope of each frame of the second speech frame by a phase vocoder algorithm.
7. The audio signal processing method according to any one of claims 6, characterized in that, The transformation of the spectral envelope of each frame of the second speech frame by the phase vocoder algorithm includes: Translating the position of each frame of the second speech frame on the time axis according to the timbre change parameter; Performing phase reconstruction on each frame of the second speech frame after translation, and using each frame of the second speech frame after phase reconstruction as each frame of the second speech frame after the transformation of the spectral envelope.
8. The audio signal processing method according to claim 7, wherein The performing phase reconstruction on each frame of the second speech frame after translation includes: Obtaining each frame of the second speech frame; Calculating the spectra of the m-th frame of the second speech frame and the (m + 1)-th frame of the second speech frame respectively through short-time Fourier transform, and determining the instantaneous frequency and phase of the m-th frame of the second speech frame and the (m + 1)-th frame of the second speech frame according to the spectra, where m is an integer greater than or equal to 1; Calculating the phase of the (m + 1)-th frame of the second speech frame after translation according to the instantaneous frequency and phase of the m-th frame of the second speech frame and the time interval before and after the movement of each frame of the second speech frame; Replacing the phase of the spectrum of the (m + 1)-th frame of the second speech frame with the calculated phase; Converting the spectrum of the (m + 1)-th frame of the second speech frame after phase replacement into a time-domain signal through inverse fast Fourier transform, and using the time-domain signal as the (m + 1)-th frame of the second speech frame after the transformation of the spectral envelope.
9. The audio signal processing method according to claim 7, wherein The translating the position of each frame of the second speech frame on the time axis according to the timbre change parameter includes: Obtaining the first number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame; Determining the second number of discrete sampling points included in the overlapping part of two adjacent frames of the second speech frame after translation according to the timbre change parameter and the first number; Translating the position of each frame of the second speech frame on the time axis according to the second number.
10. An audio signal processing device, characterized in that, The apparatus includes: An acquisition module, configured to acquire an original audio stream and predetermined parameters of the original audio stream, where the predetermined parameters include a pitch change parameter and a timbre change parameter; A framing module, configured to sample the original audio stream at a predetermined sampling rate to obtain a series of discrete sampling points, and perform framing processing on the sampling points to obtain a plurality of first speech frames; A moving module, configured to determine the fundamental frequency of each frame of the first speech frame, and move the fundamental frequency of each frame of the first speech frame according to the pitch change parameter; A splicing module, configured to perform superimposing and splicing on each frame of the first speech frame after movement to obtain an input audio stream; A transformation module, configured to divide the input audio stream into multiple frames of second speech frames with the same length, and transform the spectral envelope of each frame of the second speech frame, where two adjacent frames of the second speech frames have an overlapping part; A resampling module, configured to splice each frame of the second speech frame after the transformation of the spectral envelope to obtain an output audio stream, and resample the output audio stream according to the timbre change parameter to obtain a target audio stream.
11. A computer device, characterized in that, The computer device includes: a memory, a processor, and an audio signal processing program stored on the memory and executable on the processor. When the audio signal processing program is executed by the processor, it implements the audio signal processing method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, An audio signal processing program is stored on the computer-readable storage medium. When the audio signal processing program is executed by a processor, it implements the audio signal processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Audio processing method and device and storage medium
CN109003621A
Voice changing method and system for changing voice tones and timbres
CN111816198A