An audio processing method and computer apparatus
By obtaining the duet phrase group time period template of the target song, the voice change to be synthesized is decomposed into the target object phrase group and the user phrase group, and mixed with the target song accompaniment, the problem of lack of target object voice change accompaniment in the karaoke system is solved, and the effect of conveniently generating the target object voice change accompaniment audio is achieved.
Patent Information
- Application Number
- CN202111529065.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing karaoke systems are unable to provide voice-changing accompaniments with different target objects in a timely manner, such as the Minions voice-changing accompaniment or the Crayon Shin-chan voice-changing accompaniment, resulting in the inability to meet users' needs in different entertainment scenarios.
By obtaining the duet phrase group time period template of the target song, the voice change to be synthesized is decomposed into the target object phrase group and the user phrase group, and the target object phrase group is mixed with the target song accompaniment to generate the accompaniment audio with the target object voice change.
It realizes the rapid generation of accompaniment audio with voice change of the target object, improving the convenience of obtaining accompaniment audio and the mixing effect.
Smart Images

Figure CN114220409B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the synthesis processing of audio, in particular to an audio processing method and a computer device. BACKGROUND
[0002] In recent years, the speed of music creation has shown an explosive growth, and hundreds of original songs are produced every day. In daily entertainment, people may need to have variable sound accompaniment with different target objects in different application scenarios, such as variable sound accompaniment of "Little Yellow Man" or "Crayon Shin-chan", and such variable sound accompaniment corresponding to the original songs cannot be provided in time for use in KTV systems.
[0003] Therefore, there is an urgent need to provide a method for generating a synthesized accompaniment with variable sound of a target object. SUMMARY
[0004] The embodiments of the present application provide an audio processing method and a computer device for generating a synthesized accompaniment with variable sound of a target object.
[0005] The first aspect of the embodiments of the present application provides an audio processing method, comprising:
[0006] obtaining a to-be-synthesized variable sound, the to-be-synthesized variable sound being an audio obtained by performing variable pitch and constant speed processing on original dry sound corresponding to a target song according to a pitch of a target object;
[0007] obtaining a duet sentence group time period template of the target song, and decomposing the to-be-synthesized variable sound into a target object sentence group and a user sentence group according to the duet sentence group time period template;
[0008] retaining the target object sentence group in the to-be-synthesized variable sound, and performing mute processing on the user sentence group in the to-be-synthesized variable sound;
[0009] mixing the retained target object sentence group with a target song accompaniment according to a preset energy ratio, so that the loudness of the volume of the mixed accompaniment is the same as the original loudness of the target song accompaniment, and is not lower than a preset loudness standard;
[0010] outputting the accompaniment audio with the variable sound of the target object after mixing.
[0011] Optionally, before the obtaining of the duet sentence group time period template of the target song, the method further comprises:
[0012] obtaining the to-be-synthesized variable sound;
[0013] performing first processing on the to-be-synthesized variable sound to generate a duet sentence group time period template of the target song, wherein the first processing is used to obtain timestamp information of an effective sound sentence group in the to-be-synthesized variable sound.
[0014] Optionally, the first processing is performed on the to-be-synthesized variable voice to generate a duet phrase period template of the target song, and the first processing comprises:
[0015] performing smoothing processing on a voice signal of the to-be-synthesized variable voice to obtain an amplitude envelope of the to-be-synthesized variable voice;
[0016] processing the amplitude envelope of the to-be-synthesized variable voice by using a threshold function to obtain timestamp information of an effective voice phrase group in the to-be-synthesized variable voice;
[0017] performing duet period marking on the to-be-synthesized variable voice according to the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice to generate the duet phrase period template of the target song.
[0018] Optionally, before performing smoothing processing on a voice signal of the to-be-synthesized variable voice to obtain an amplitude envelope of the to-be-synthesized variable voice, the method further comprises:
[0019] performing low-pass filtering on the voice signal of the to-be-synthesized variable voice to obtain a to-be-synthesized variable voice signal of a target frequency band.
[0020] Optionally, before performing low-pass filtering on the voice signal of the to-be-synthesized variable voice, the method further comprises:
[0021] performing normalization processing on an amplitude value of the voice signal of the to-be-synthesized variable voice to obtain a normalized to-be-synthesized variable voice.
[0022] Optionally, the duet period marking is performed on the to-be-synthesized variable voice according to the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice to generate the duet phrase period template of the target song, and the duet period marking comprises:
[0023] obtaining time lyric information of the to-be-synthesized variable voice;
[0024] performing duet period marking on the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice and the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice to generate the duet phrase period template of the target song.
[0025] Optionally, the duet period marking is performed on the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice and the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice to generate the duet phrase period template of the target song, and the duet period marking comprises:
[0026] performing alignment operation on the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice and timestamp information in the time lyric information of the to-be-synthesized variable voice to obtain corrected timestamp information of the effective voice phrase group;
[0027] According to the timestamp information of the corrected valid sound sentence group, a duet period mark is performed on the to-be-composed variable voice to generate a duet sentence group period template of the target song.
[0028] Optionally, before the mixing of the reserved target object sentence group with the target song accompaniment according to the preset energy ratio, the method further comprises:
[0029] The loudness of the target object sentence group is set to be less than the loudness of the target song accompaniment and lower than the loudness of the target song accompaniment by a preset value.
[0030] Optionally, the method further comprises:
[0031] Obtaining time lyric information of the to-be-composed variable voice;
[0032] After the target object sentence group in the to-be-composed variable voice is reserved and the user sentence group in the to-be-composed variable voice is muted, the method further comprises:
[0033] According to the time lyric information of the to-be-composed variable voice, the reserved target object sentence group, and the muted user sentence group, the time lyric information of the target object sentence group in the time lyric information of the to-be-composed variable voice is marked to generate time lyric information with target object marking.
[0034] Optionally, the marking of the time lyric information of the target object sentence group in the time lyric information of the to-be-composed variable voice according to the time lyric information of the to-be-composed variable voice, the reserved target object sentence group, and the muted user sentence group to generate time lyric information with target object marking comprises:
[0035] According to the time lyric information of the to-be-composed variable voice, the reserved target object sentence group, and the muted user sentence group, the timestamp position of the lyrics of the target object sentence group in the time lyric information of the to-be-composed variable voice is marked by using a note file or a midi file to generate time lyric information with target object marking.
[0036] Optionally, before the obtaining of the to-be-composed variable voice, the method further comprises:
[0037] Performing variable-pitch constant-speed processing on the original dry voice of the target song according to the pitch of the target object to obtain a target object variable voice set of the target song;
[0038] Selecting a to-be-composed variable voice that meets a preset standard from the target object variable voice set of the target song.
[0039] Optionally, before performing pitch conversion and constant speed processing on the original dry sound of the target song according to the pitch of the target object to obtain a target object pitch conversion set of the target song, the method further comprises:
[0040] filtering an original dry sound set meeting preset melody standards and preset timbre standards from the dry sound set of the target song;
[0041] filtering a to-be-synthesized pitch conversion meeting preset standards from the target object pitch conversion set of the target song, comprising:
[0042] filtering a to-be-synthesized pitch conversion meeting preset standards from the target object pitch conversion set of the target song according to naturalness and pleasantness standards of the listening.
[0043] Optionally, the preset melody standards comprise preset pitch standards and preset rhythm standards.
[0044] The naturalness and pleasantness standards of the listening comprise at least one of smoothness of audio breath, cosine similarity of a fundamental frequency sequence, and stability of audio beat distribution.
[0045] Optionally, the pitch conversion and constant speed processing comprises resampling pitch conversion and constant speed processing and phase frequency vocoder pitch conversion and constant speed processing.
[0046] Optionally, the pitch conversion and constant speed processing on the original dry sound of the target song according to the pitch of the target object comprises:
[0047] performing the resampling pitch conversion and constant speed processing on the speech signal of the original dry sound first, and then performing the phase frequency vocoder pitch conversion and constant speed processing.
[0048] or,
[0049] performing the phase frequency vocoder pitch conversion and constant speed processing on the speech signal of the original dry sound first, and then performing the resampling pitch conversion and constant speed processing.
[0050] Optionally, performing the resampling pitch conversion and constant speed processing on the speech signal of the original dry sound first, and then performing the phase frequency vocoder pitch conversion and constant speed processing, comprises:
[0051] performing resampling on the time domain speech signal of the original dry sound to obtain a second pitch conversion after pitch conversion and constant speed processing;
[0052] frame and windowing the second pitch conversion to decompose the second pitch conversion into a plurality of analysis frames;
[0053] performing time domain and frequency domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from a time domain function to a frequency domain function;
[0054] keeping the spectral amplitudes of the analysis frames unchanged, modifying phase information of the analysis frames to obtain a plurality of synthesis frames;
[0055] overlapping and adding the plurality of synthesis frames by a synthesis window function to obtain a third variable sound performing pitch-invariant speed processing on the original dry sound.
[0056] Optionally, the modifying the phase information of the analysis frames to obtain a plurality of synthesis frames comprises:
[0057] calculating a phase difference of adjacent analysis frames;
[0058] obtaining a pitch parameter of the second variable sound to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, an inverse of the pitch parameter of the second variable sound and the any synthesis frame;
[0059] constructing a frequency domain function of the target synthesis frame according to the spectral amplitudes of the analysis frames and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0060] Optionally, the performing the speed-invariant pitch strategy of the phase-frequency vocoder on the speech signal of the original dry sound and then performing the speed-variable pitch strategy of the resampling comprises:
[0061] framing and windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time domain functions;
[0062] performing time domain and frequency domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time domain functions to frequency domain functions;
[0063] keeping the spectral amplitudes of the analysis frames unchanged, modifying phase information of the analysis frames to obtain a plurality of synthesis frames;
[0064] overlapping and adding the plurality of synthesis frames by a synthesis window function to obtain a fourth variable sound performing speed-invariant pitch processing on the original dry sound;
[0065] performing resampling on a time domain signal of the fourth variable sound according to a target sampling rate to obtain a fifth variable sound performing pitch-invariant speed processing on the original dry sound, wherein a pitch parameter obtained according to the target sampling rate and a speed parameter of the fourth variable sound are inverses of each other.
[0066] Optionally, the modifying the phase information of the analysis frames to obtain a plurality of synthesis frames comprises:
[0067] calculating a phase difference of adjacent analysis frames;
[0068] acquire a variable speed parameter of the fourth variable voice, and construct a synthesis phase of a target synthesis frame adjacent to the any synthesis frame according to a phase difference of the adjacent analysis frames, the variable speed parameter of the fourth variable voice and the any synthesis frame;
[0069] construct a frequency domain function of the target synthesis frame according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame, so as to obtain a plurality of synthesis frames.
[0070] Optionally, if the audio in the target object variable voice set of the target song is a rising variable voice performed on the original dry voice according to a target object tone, the variable tone and constant speed processing performed on the original dry voice of the target song according to the target object tone comprises:
[0071] the original dry voice in the original dry voice set is first subjected to the variable speed and constant tone processing of the phase frequency vocoder, and then subjected to the variable speed and variable tone processing of the resampling.
[0072] Optionally, if the audio in the target object variable voice set of the target song is a rising variable voice performed on the original dry voice according to a target object tone, the variable tone and constant speed processing performed on the original dry voice of the target song according to the target object tone comprises:
[0073] the original dry voice in the original dry voice set is first subjected to the variable speed and variable tone processing of the resampling, and then subjected to the variable speed and constant tone processing of the phase frequency vocoder.
[0074] The second aspect of the embodiment of the present application provides an audio processing device, comprising:
[0075] an acquisition unit configured to acquire a to-be-synthesized variable voice, the to-be-synthesized variable voice being audio obtained by performing variable tone and constant speed processing on an original dry voice corresponding to a target song according to a target object tone;
[0076] a decomposition unit configured to acquire a duet phrase period template of the target song, and decompose the to-be-synthesized variable voice into a target object phrase and a user phrase according to the duet phrase period template;
[0077] a processing unit configured to retain the target object phrase in the to-be-synthesized variable voice, and perform mute processing on the user phrase in the to-be-synthesized variable voice;
[0078] a mixing unit configured to mix the retained target object phrase with a target song accompaniment according to a preset energy ratio, so that a loudness of a volume of the mixed accompaniment is the same as an original loudness of the target song accompaniment and is not lower than a preset loudness standard;
[0079] an output unit configured to output the mixed accompaniment audio with the target object variable voice.
[0080] Optionally, the audio processing apparatus further comprises:
[0081] an acquisition unit, configured to acquire the to-be-synthesized variable voice;
[0082] The processing unit is further configured to perform first processing on the to-be-synthesized variable voice to generate the duet phrase period template of the target song, wherein the first processing is configured to acquire timestamp information of an effective voice phrase in the to-be-synthesized variable voice.
[0083] Optionally, the processing unit is specifically configured to:
[0084] perform smoothing processing on a voice signal of the to-be-synthesized variable voice to acquire an amplitude envelope of the to-be-synthesized variable voice;
[0085] perform processing on the amplitude envelope of the to-be-synthesized variable voice by using a threshold function to acquire the timestamp information of the effective voice phrase in the to-be-synthesized variable voice;
[0086] perform duet period marking on the to-be-synthesized variable voice according to the timestamp information of the effective voice phrase in the to-be-synthesized variable voice to generate the duet phrase period template of the target song.
[0087] Optionally, the processing unit is further configured to:
[0088] perform low-pass filtering on a voice signal of the to-be-synthesized variable voice to acquire a to-be-synthesized variable voice signal of a target frequency band.
[0089] Optionally, the processing unit is further configured to:
[0090] perform normalization processing on an amplitude value of the voice signal of the to-be-synthesized variable voice to obtain a normalized to-be-synthesized variable voice.
[0091] Optionally, the acquisition unit is further configured to:
[0092] acquire time lyric information of the to-be-synthesized variable voice;
[0093] Optionally, the processing unit is further configured to:
[0094] perform duet period marking on the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice and the timestamp information of the effective voice phrase in the to-be-synthesized variable voice to generate the duet phrase period template of the target song.
[0095] Optionally, the processing unit is specifically configured to:
[0096] perform alignment operation on the timestamp information of the effective voice phrase in the to-be-synthesized variable voice and timestamp information in the time lyric information of the to-be-synthesized variable voice to obtain corrected timestamp information of the effective voice phrase;
[0097] According to the timestamp information of the corrected valid sound sentence group, a duet period mark is performed on the to-be-combined variable voice to generate a duet sentence group period template of the target song.
[0098] Optionally, the audio processing apparatus further comprises:
[0099] The setting unit is configured to set the loudness of the target object sentence group to be less than the loudness of the accompaniment of the target song and lower than the loudness of the accompaniment of the target song by a preset value.
[0100] Optionally, the acquisition unit is further configured to:
[0101] Acquire time lyric information of the to-be-combined variable voice;
[0102] The processing unit is further configured to:
[0103] According to the time lyric information of the to-be-combined variable voice, the retained target object sentence group and the muted user sentence group, mark the time lyric information of the target object sentence group in the time lyric information of the to-be-combined variable voice to generate time lyric information with target object marks.
[0104] Optionally, the processing unit is specifically configured to:
[0105] According to the time lyric information of the to-be-combined variable voice, the retained target object sentence group and the muted user sentence group, mark the time stamp position of the lyrics of the target object sentence group in the time lyric information of the to-be-combined variable voice by using a note file or a midi file to generate time lyric information with target object marks.
[0106] Optionally, the audio processing apparatus further comprises:
[0107] The variable voice unit is configured to, before acquiring the to-be-combined variable voice, perform variable pitch and constant speed processing on the original dry voice of the target song according to the pitch of the target object to obtain a target object variable voice set of the target song.
[0108] The screening unit is configured to screen the to-be-combined variable voice that meets the preset standard from the target object variable voice set of the target song.
[0109] Optionally, the screening unit is further configured to:
[0110] Screen an original dry voice set that meets the preset melody standard and the preset sound quality standard from a dry voice set of the target song.
[0111] The screening unit is specifically configured to:
[0112] According to the naturalness and pleasantness of the sound, the target voice of the target song is filtered from the target voice set of the target song.
[0113] Optionally, the preset melody standard includes a preset pitch standard and a preset rhythm standard.
[0114] The naturalness and pleasantness of the sound includes at least one of smoothness of audio breath, cosine similarity of a fundamental frequency sequence, and stability of audio beat distribution.
[0115] Optionally, the pitch change and speed invariant processing includes a resampling pitch change and speed change processing and a phase vocoder pitch change and speed invariant processing.
[0116] Optionally, the voice changing unit is specifically configured to:
[0117] The resampling pitch change and speed change processing is performed on the original dry voice, and then the phase vocoder pitch change and speed invariant processing is performed.
[0118] Or,
[0119] The phase vocoder pitch change and speed invariant processing is performed on the original dry voice, and then the resampling pitch change and speed change processing is performed.
[0120] Optionally, the voice changing unit is specifically configured to:
[0121] The resampling is performed on the time domain voice signal of the original dry voice to obtain a second voice after pitch change and speed change.
[0122] The second voice is framed and windowed to decompose the second voice into a plurality of analysis frames.
[0123] The plurality of analysis frames are converted from time domain functions to frequency domain functions.
[0124] The spectral amplitudes of the analysis frames are kept unchanged, and the phase information of the analysis frames is modified to obtain a plurality of synthesis frames.
[0125] The plurality of synthesis frames are superimposed and added through a synthesis window function to obtain a third voice after pitch change and speed invariant processing on the original dry voice.
[0126] Optionally, the voice changing unit is specifically configured to:
[0127] The phase difference of adjacent analysis frames is calculated.
[0128] The pitch change parameter of the second voice is obtained to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of adjacent analysis frames, the inverse of the pitch change parameter of the second voice, and the any synthesis frame.
[0129] constructing a frequency domain function of the target synthesis frame according to the spectral magnitude of the analysis frame and the synthesis phase of the target synthesis frame, to obtain a plurality of synthesis frames.
[0130] Optionally, the voice variation unit is specifically configured to:
[0131] windowing the original dry voice frame to decompose the original dry voice into a plurality of analysis frames, wherein the plurality of analysis frames are time domain functions;
[0132] performing time domain and frequency domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time domain functions to frequency domain functions;
[0133] keeping the spectral magnitude of the analysis frame unchanged, and modifying the phase information of the analysis frame to obtain a plurality of synthesis frames;
[0134] overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a fourth voice variation of the original dry voice subjected to the variable speed constant pitch processing;
[0135] performing resampling on a time domain signal of the fourth voice variation according to a target sampling rate to obtain a fifth voice variation of the original dry voice subjected to the variable pitch constant speed processing, wherein a variable pitch parameter obtained according to the target sampling rate and a variable speed parameter of the fourth voice variation are reciprocal to each other.
[0136] Optionally, the voice variation unit is specifically configured to:
[0137] calculating a phase difference of adjacent analysis frames;
[0138] obtaining a variable speed parameter of the fourth voice variation, to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the variable speed parameter of the fourth voice variation, and the any synthesis frame;
[0139] constructing a frequency domain function of the target synthesis frame according to the spectral magnitude of the analysis frame and the synthesis phase of the target synthesis frame, to obtain a plurality of synthesis frames.
[0140] Optionally, if the audio in the target object voice variation set of the target song is a rising pitch voice variation of the original dry voice according to the target object's voice tone, the voice variation unit is specifically configured to:
[0141] performing the variable speed constant pitch strategy of the phase frequency vocoder on the original dry voice in the original dry voice set first, and then performing the variable speed variable pitch strategy of the resampling.
[0142] Optionally, if the audio in the target object voice variation set of the target song is a falling pitch voice variation of the original dry voice according to the target object's voice tone, the voice variation unit is specifically configured to:
[0143] The variable-speed variable-pitch strategy is performed on the original dry sound in the original dry sound set first, and then the variable-speed constant-pitch strategy of the phase-frequency vocoder is performed.
[0144] The third aspect of the embodiments of the present application provides a computer device, comprising a processor, which is used to implement the audio processing method of the first aspect of the embodiments of the present application when executing a computer program stored on a memory.
[0145] The fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is used to implement the audio processing method of the first aspect of the embodiments of the present application when executed by a processor.
[0146] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0147] The target object variable sound is obtained, the target object variable sound is audio obtained by performing variable-pitch constant-speed processing on original dry sound corresponding to a target song according to a tone of a target object, a duet sentence group time period template of the target song is obtained, the target object variable sound is decomposed into a target object sentence group and a user sentence group according to the duet sentence group time period template, the target object sentence group in the target object variable sound is reserved, and the user sentence group in the target object variable sound is subjected to mute processing, the reserved target object sentence group is mixed with a target song accompaniment according to a preset energy ratio, so that the loudness of the volume of the mixed accompaniment is the same as the original loudness of the target song accompaniment, and is not lower than a preset loudness standard, and the accompaniment audio with the target object variable sound after mixing is output.
[0148] In the embodiments of the present application, the target object variable sound is obtained first, then the target object variable sound is decomposed into a target object sentence group and a user sentence group by using a duet sentence group time period template of a target song, and further the reserved target object sentence group is mixed with a target song accompaniment according to a preset energy ratio, so that the accompaniment audio with the target object variable sound is generated, and the convenience of obtaining the accompaniment audio with the target object variable sound is improved. BRIEF DESCRIPTION OF DRAWINGS
[0149] Figure 1 An embodiment of the audio processing method in the embodiments of the present application is shown;
[0150] Figure 2 Another embodiment of the audio processing method in the embodiments of the present application is shown;
[0151] Figure 3 is Figure 2 A detailed step of step 202 in the embodiments;
[0152] FIG. 4A is a waveform diagram of a voice signal before processing in an embodiment of the present application;
[0153] FIG. 4B is a diagram of an amplitude envelope of the voice signal after smoothing in an embodiment of the present application;
[0154] Figure 5 FIG. 4C is a diagram of the amplitude envelope of the voice signal with a threshold in an embodiment of the present application;
[0155] Figure 6 Figure 2 Another refinement of step 202 in an embodiment;
[0156] Figure 7 FIG. 5B is a diagram of another embodiment of an audio processing method in an embodiment of the present application;
[0157] Figure 8 FIG. 5B is a diagram of another embodiment of an audio processing method in an embodiment of the present application;
[0158] Figure 9 FIG. 5B is a diagram of another embodiment of an audio processing method in an embodiment of the present application;
[0159] Figure 10 FIG. 6A is a diagram of one embodiment of a pitch shifting process on the original dry voice according to the target object's pitch in an embodiment of the present application;
[0160] Figure 11 FIG. 6B is a diagram of another embodiment of a pitch shifting process on the original dry voice according to the target object's pitch in an embodiment of the present application; Figure 10 A refinement of step 1004 in an embodiment;
[0161] Figure 12 FIG. 6B is a diagram of another embodiment of a pitch shifting process on the original dry voice according to the target object's pitch in an embodiment of the present application;
[0162] Figure 13 FIG. 6B is a diagram of another embodiment of a pitch shifting process on the original dry voice according to the target object's pitch in an embodiment of the present application; Figure 12 A refinement of step 1204 in an embodiment;
[0163] Figure 14 FIG. 7 is a diagram of one embodiment of an audio processing device in an embodiment of the present application;
[0164] Figure 15 FIG. 8 is a diagram of a second pitch shifted voice and an analysis frame obtained by framing and windowing the second pitch shifted voice in an embodiment of the present application. DETAILED DESCRIPTION
[0165] Embodiments of the present application provide an audio processing method and a computer device for obtaining a backing track with a target object's pitch, to improve the convenience of obtaining a backing track with a target object's pitch.
[0166] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort should fall within the protection scope of the present application.
[0167] The terms "first", "second", "third", "fourth" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a particular order or sequence. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments described herein can be carried out in sequences other than those illustrated or described herein. Furthermore, the terms "comprise" and "have", and any variations thereof, are intended to cover non-exclusive inclusion, for example, processes, methods, systems, products, or devices that include a series of steps or units are not necessarily limited to those clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0168] Karaoke singing has become one of the lifestyles of people gathering and entertaining, and the demand for singing, practicing and listening to songs is gradually becoming a common demand of families. In the previous K song system, people will bring music accompaniment when singing to assist people to better sing and practice.
[0169] As a way of singing, antiphonal singing is a lively form of singing, which is different from duet, and is a single-voice song. In daily entertainment, people may need to bring variable accompaniment with different target objects, such as Minion variable accompaniment or Crayon Shin-chan variable accompaniment, in different application scenarios, and such variable accompaniment corresponding to the original song cannot be provided in time for use in the K song system.
[0170] The present application aims to provide an audio processing method and a computer device for obtaining accompaniment audio with variable voice of a target object, so as to improve the convenience of obtaining accompaniment audio with variable voice of a target object.
[0171] For the convenience of understanding, the audio processing method in the present application will be described below. Please refer to Figure 1 One embodiment of the audio processing method in the present application includes:
[0172] 101, obtaining a to-be-synthesized variable voice, the to-be-synthesized variable voice being audio obtained by performing variable pitch and constant speed processing on original dry voice corresponding to a target song according to a pitch of a target object;
[0173] In order to obtain a variable sound accompaniment with a different target object, such as a variable sound accompaniment with a Minion or a variable sound accompaniment with a Crayon Shin-chan, a to-be-synthesized variable sound needs to be obtained, wherein the to-be-synthesized variable sound is an audio obtained by performing pitch conversion without speed conversion on an original dry sound corresponding to a target song according to a pitch of a target object.
[0174] The target song is a song to be processed, such as a song of Blue and White Porcelain or a song of Little Donkey, and the original dry sound refers to pure human voice without music, that is, a singing audio of the target song without music accompaniment.
[0175] Specifically, a process of performing pitch conversion without speed conversion on the original dry sound of the target song according to the pitch of the target object will be described in detail in the following embodiments, and will not be described here again
[0176] 102. Obtain a duet sentence group time period template of the target song, and decompose the to-be-synthesized variable sound into a target object sentence group and a user sentence group according to the duet sentence group time period template;
[0177] After obtaining the to-be-synthesized variable sound, a duet sentence group time period template of the target song can be further obtained, and the to-be-synthesized variable sound is decomposed into a target object sentence group and a user sentence group according to the duet sentence group time period template.
[0178] The duet sentence group time period template of the target song refers to an identification template for identifying which sentence group in the target song belongs to which time period, and which sentence group belongs to the A song part in the duet and which sentence group belongs to the B song part in the duet, such as in the song of Little Donkey, “I have a little donkey, I never ride it” belongs to 0.00s to 0.04s, and belongs to the A song part, “one day I got the idea to ride it to the market” belongs to 0.05s to 0.09s, and belongs to the B song part, etc. The A song and the B song here refer to the part sung by the A song and the part sung by the B song in the duet.
[0179] After obtaining the duet sentence group time period template of the target song, the to-be-synthesized variable sound can be further decomposed into a target object sentence group and a user sentence group according to the template, wherein the target object sentence group refers to a part sung by the target object (such as Minion or Crayon Shin-chan), and the user sentence group refers to a part sung by the user.
[0180] 103. Retain the target object sentence group in the to-be-synthesized variable sound, and perform mute processing on the user sentence group in the to-be-synthesized variable sound;
[0181] After decomposing the to-be-synthesized variable sound into a target object sentence group and a user sentence group, in order to obtain an accompaniment audio with a target object variable sound, the user sentence group in the to-be-synthesized variable sound can be subjected to mute processing to obtain a variable sound set with only the target object variable sound.
[0182] 104. Mixing the retained target sentence group with the target song accompaniment according to a preset energy ratio, so that the loudness of the mixed accompaniment volume is the same as the original loudness of the target song accompaniment and is not lower than a preset loudness standard;
[0183] In order to further obtain an accompaniment with a voice change of the target object, the target object sentence group in the voice change to be synthesized can be mixed with the target song accompaniment according to a preset energy ratio, so that the loudness of the accompaniment volume after mixing is the same as the original loudness of the target song accompaniment and is not lower than a preset loudness standard, such as the loudness of the accompaniment volume after mixing is not lower than -14dB.
[0184] 105. Output the mixed accompaniment audio with the target object's voice changed.
[0185] After the target object and the target song accompaniment are mixed, the mixed accompaniment audio with the target object's voice change can be further output.
[0186] In an embodiment of the present application, the voice change to be synthesized is first obtained, and then the voice change to be synthesized is further decomposed into a target object sentence group and a user sentence group according to the duet sentence group time period template of the target song, the target object sentence group in the voice change to be synthesized is retained, and the user sentence group in the voice change to be synthesized is muted, and the retained target object sentence group is further mixed with the target song accompaniment according to a preset energy ratio, so as to obtain accompaniment audio with the target object voice change, thereby improving the convenience of obtaining accompaniment audio with the target object voice change.
[0187] based on Figure 1 In the embodiment described, before executing step 104, in order to further improve the mixing effect, the loudness of the target object sentence group can also be set to be smaller than the loudness of the target song accompaniment, and lower than the loudness of the target song accompaniment by a preset value. For example, the loudness of the target object sentence group is set to be 3dB lower than the loudness of the target song accompaniment, so as to prevent the dry sound energy after voice change from being too large to cover the volume of the accompaniment, thereby ensuring the mixing effect.
[0188] based on Figure 1 In the embodiment described above, before executing step 102 to obtain the duet phrase group time period template of the target song, the following steps may be performed to generate the duet phrase group time period template of the target song, see Figure 2 Another embodiment of the audio processing method in the embodiment of the present application includes:
[0189] 201. Perform a first processing on the voice change to be synthesized to generate a duet phrase group time period template of the target song, wherein the first processing is used to obtain timestamp information of valid voice phrase groups in the voice change to be synthesized.
[0190] After obtaining the to-be-synthesized variable voice, a first processing is performed on the to-be-synthesized variable voice to generate a duet sentence group period template of the target song, wherein the first processing is used to obtain timestamp information of an effective voice sentence group in the to-be-synthesized variable voice.
[0191] Specifically, the process of performing the first processing on the to-be-synthesized variable voice is as follows Figure 3 The:
[0192] 301. Perform smoothing processing on the voice signal of the to-be-synthesized variable voice to obtain the amplitude envelope of the to-be-synthesized variable voice.
[0193] The voice signal is a kind of sound signal, which will present a certain waveform. The original voice signal will present a certain burr due to various interference signals when it is not processed. In order to obtain the amplitude envelope feature of the to-be-synthesized variable voice, that is, to obtain the external contour information of the to-be-synthesized variable voice signal, smoothing processing can be performed on the voice signal of the to-be-synthesized variable voice to obtain the amplitude envelope feature of the to-be-synthesized variable voice.
[0194] Specifically, when performing smoothing processing on the voice signal of the to-be-synthesized variable voice, mean smoothing, exponential smoothing or median smoothing can be used, as long as the amplitude envelope feature of the to-be-synthesized variable voice can be obtained, which is not limited here.
[0195] Next, taking mean smoothing as an example, the process of performing smoothing processing on the voice signal of the to-be-synthesized variable voice is described.
[0196] Suppose each frame of voice signal is x(n), where n represents the sample index, that is, the sample point on each frame signal, then the process of mean smoothing can be represented as x m (n)=conv(x(n),w r (n)), where
[0197] For easy understanding, Fig. 4A shows a waveform diagram before performing smoothing processing on the voice signal, and Fig. 4B shows an amplitude envelope diagram of the to-be-synthesized variable voice obtained after performing smoothing processing on the voice signal.
[0198] 302. Use a threshold function to process the amplitude envelope of the to-be-synthesized variable voice to obtain timestamp information of an effective voice sentence group in the to-be-synthesized variable voice.
[0199] In order to distinguish the effective voice sentence group in the to-be-synthesized variable voice, a threshold function can be used to process the amplitude envelope of the to-be-synthesized variable voice to obtain the timestamp information of the effective voice sentence group in the to-be-synthesized variable voice.
[0200] Next, the process of using a threshold function to process the amplitude envelope of the to-be-synthesized variable voice is described.
[0201] Assume that the threshold function is v(n), Wherein, thr=w, w can be customized according to actual needs, so as to obtain the timestamp information of the effective voice sentence group in the synthesized variable voice through the threshold function.
[0202] For the convenience of understanding, Figure 5 The amplitude envelope of the synthesized variable voice with threshold mark is given.
[0203] 303, according to the timestamp information of the effective voice sentence group in the synthesized variable voice, the duet time period marking of the synthesized variable voice is performed to generate the duet sentence group time period template of the target song.
[0204] After obtaining the timestamp information of the effective voice sentence group in the synthesized variable voice, the duet time period marking of the synthesized variable voice is performed, such as the voice sentence group from 2s to 2.5s belongs to A song part, and the voice sentence group from 2.5s to 3s belongs to B song part, wherein the A song part belongs to A duet part, and the B song part belongs to B duet part.
[0205] After completing the duet time period marking of the synthesized variable voice, the duet sentence group time period template of the target song can be generated.
[0206] Based on the above first processing, before performing step 301 on the voice signal of the synthesized variable voice, step can also be performed to facilitate the quick processing of the voice signal, please refer to 6:
[0207] 601, the amplitude of the voice signal of the synthesized variable voice is normalized to obtain the normalized synthesized variable voice;
[0208] In order to summarize the statistical distribution of the sound signal, the amplitude of the voice signal of the synthesized variable voice can be normalized to obtain the normalized synthesized variable voice signal. Specifically, assume that |x(n)| represents the amplitude of each sample point, and max{|x(n)|} represents the maximum amplitude of the sample point.
[0209] Then the normalized amplitude feature is:
[0210] 602, the normalized synthesized variable voice signal is low-pass filtered to obtain the target frequency band synthesized variable voice.
[0211] After obtaining the normalized synthesized variable voice signal, because there are many sound signals in different frequency ranges in the synthesized variable voice, and the human voice signal is mainly concentrated within 1000Hz, it is necessary to low-pass filter the normalized synthesized variable voice signal to obtain the target frequency band synthesized variable voice, wherein the target frequency band is within 1000Hz.
[0212] Specifically, the low-pass filtering process is as follows: x r (n) = resample(x a (n)).
[0213] Based on Figures 3 to 6 The embodiment described above describes in detail the process of generating the duet sentence group period template of the target song. On the one hand, the normalization process improves the convenience of the synthesized voice processing, and on the other hand, the low-pass filtering obtains effective singing signals. The process of generating the duet sentence group period template of the target song is simple and convenient, and the convenience of generating the duet sentence group period template of the target song is improved.
[0214] Based on Figure 3 In step 303 in the embodiment, in order to improve the accuracy of the duet period marking, the following steps can also be performed when the synthesized voice is executed. Please refer to Figure 7 , Figure 7 Another embodiment of the audio processing method in the embodiment is as follows:
[0215] 701, obtaining the time lyrics information of the synthesized voice;
[0216] When the synthesized voice is executed, the time lyrics information of the synthesized voice can also be obtained. The specific time lyrics information refers to the file information of the lyrics corresponding to the time point, such as the QRC lyrics sentence information of each song.
[0217] 702, according to the time lyrics information of the synthesized voice and the timestamp information of the effective voice sentence group in the synthesized voice, performing the duet period marking on the synthesized voice to generate the duet sentence group period template of the target song.
[0218] According to the timestamp information of the effective voice sentence group in the synthesized voice and the time lyrics information, the timestamp information of the effective voice sentence group is aligned to correct the timestamp information of the effective voice sentence group in the synthesized voice. According to the corrected timestamp information of the effective voice sentence group, the duet period marking is performed on the effective voice sentence group in the synthesized voice to generate the duet sentence group period template of the target song corresponding to the synthesized voice segment.
[0219] In the embodiment, the QRC lyrics sentence information of the synthesized voice and the timestamp information of the effective voice sentence group are combined to perform the duet period marking on the synthesized voice, thereby improving the accuracy of obtaining the duet sentence group period template of the target song.
[0220] Based on Figures 1 to 7The embodiment can further perform the following steps to generate the time lyric information with target object mark, and improve the convenience of the user and the target object in singing, please refer to Figure 8 Another embodiment of the audio processing method in the present application comprises:
[0221] 801, obtaining time lyric information of the to-be-synthesized variable voice;
[0222] In order to output the time lyric information with target object mark in the accompaniment audio with target object variable voice, the time lyric information of the to-be-synthesized variable voice can be further obtained, such as the QRC lyric information of the to-be-synthesized variable voice.
[0223] 802, according to the time lyric information of the to-be-synthesized variable voice, the retained target object sentence group and the user sentence group after mute processing, marking the time lyric information of the target object sentence group in the time lyric information of the to-be-synthesized variable voice to generate the time lyric information with target object mark.
[0224] After obtaining the time lyric information of the to-be-synthesized variable voice, the time lyric information of the target object sentence group in the time lyric information of the to-be-synthesized variable voice can be further marked according to the retained target object sentence group and the user sentence group after mute processing to generate the time lyric information with target object mark, so as to guide the user to sing along.
[0225] Further, in order to improve the accuracy of generating the time lyric information with target object mark, the time stamp position of the lyrics of the target object sentence group in the time lyric information of the to-be-synthesized variable voice can be marked by using the note file or the midi file, so as to confirm the singing time stamp position to the word or note level, and improve the accuracy of generating the time lyric information with target object mark.
[0226] Based on Figures 1 to 8 The embodiment can further perform the following steps before step 101, please refer to Figure 9 , Figure 9 Another embodiment of the audio processing method comprises:
[0227] 901, screening an original dry sound set meeting preset melody standards and preset sound quality standards from the dry sound set of the target song;
[0228] Firstly, an original dry sound set meeting preset melody standards and preset sound quality standards is screened from the dry sound set of the target song, wherein the preset melody standards include preset pitch standards and preset rhythm standards.
[0229] 902、performing pitch conversion and constant speed processing on the original dry sounds of the target song according to the vocal tone of the target object to obtain a target object converted sound set of the target song;
[0230] After the original dry sound set meeting the preset melody standard and the preset timbre standard is screened out from the dry sound set of the target song, performing pitch conversion and constant speed processing on the original dry sound in the original dry sound set to obtain a target object converted sound set of the target song.
[0231] Specifically, the pitch conversion and constant speed processing in the embodiment includes the resampling pitch conversion and speed conversion processing and the phase frequency vocoder speed conversion and constant pitch processing, and the process of performing pitch conversion and constant speed processing on the original dry sound in the original dry sound set will be described in the following embodiments, which will not be described here.
[0232] 903、screening a to-be-synthesized converted sound meeting a preset standard from the target object converted sound set of the target song according to a naturalness and pleasantness standard of listening.
[0233] After obtaining the target object converted sound set of the target song, screening a to-be-synthesized converted sound meeting a preset standard from the target object converted sound set of the target song according to a naturalness and pleasantness standard of listening, wherein the naturalness and pleasantness standard of listening includes at least one of a smoothness of audio breath, a cosine similarity of a fundamental frequency sequence, and a stability of audio beat distribution.
[0234] The process of screening a to-be-synthesized converted sound meeting a preset standard from the target object converted sound set of the target song in the embodiment is described in detail, thereby providing a basis for generating a high-standard to-be-synthesized converted sound.
[0235] Based on the above-described embodiment, when the original dry sound in the original dry sound set is performed pitch conversion and constant speed processing according to the vocal tone of the target object, the following two ways can be used, which will be described as follows: Figure 9 I. performing the resampling speed conversion and pitch conversion strategy on the speech signal of the original dry sound first, and then performing the phase frequency vocoder speed conversion and constant pitch conversion strategy to obtain a target object converted sound set of the target song, for details, please refer to
[0236] : Figure 10
[0237] 1001、performing resampling on the time domain speech signal of the original dry sound to obtain a second converted sound after speed conversion and pitch conversion;
[0238] When performing resampling on the speech signal of the original dry sound, generally, the time domain speech signal of the original dry sound is performed resampling to obtain a second converted sound after speed conversion and pitch conversion.
[0239] The so-called resampling is to extract or interpolate the sample points in each frame of speech signal. If the pitch of the original dry speech signal is to be doubled, only one sample point is discarded every other sample point. If the pitch of the original dry speech signal is to be halved, only one sample point is inserted every other sample point.
[0240] In the process of implementing resampling, although the pitch is doubled, the length of the speech signal (i.e. the time of the speech signal) is correspondingly shortened or doubled. That is, in the process of extracting or interpolating the sample points in each frame of speech signal, the speed and pitch of the speech signal are changed.
[0241] Therefore, after performing resampling on the time-domain speech signal of the original dry speech, the second variable sound after changing the speed and pitch can be obtained.
[0242] 1002, frame and window the second variable sound to decompose the second variable sound into a plurality of analysis frames;
[0243] After obtaining the second variable sound, in order to implement the variable speed and constant pitch processing of the second variable sound, the frequency domain analysis needs to be performed on the second variable sound. However, when performing Fourier transform on the speech signal, it is required that the speech signal is as periodic as possible. Therefore, the embodiment of the present application performs frame and window processing on the second variable sound to decompose the second variable sound into a plurality of analysis frames.
[0244] Among them, frame the second variable sound is to divide the continuous speech signal into a plurality of discrete periodic signals. However, in the process of framing, in order to avoid the accuracy of signal analysis not being enough due to some jump signals in the signal, i.e. in order to improve the resolution of the signal, generally a certain overlap degree is set between two adjacent analysis frames in the framing process, i.e. there is a certain frame shift between adjacent analysis frames.
[0245] Suppose T u is the frame length, h a is the frame shift, and in order to ensure the continuity between the frames, generally h a ≤ T u / 4.
[0246] Further, windowing each analysis frame is to reduce the frequency domain leakage caused by the fact that the signal cannot be truncated by an integer period of times in the framing process. When windowing the analysis frame, the hann or hamming window has excellent performance in generating the time-frequency diagram of the speech signal. Therefore, the hann or hamming window is preferred in the embodiment of the present application.
[0247] For the convenience of understanding, Figure 15 a schematic diagram of the second variable sound and the analysis frame after framing and windowing the second variable sound is given.
[0248] 1003、performing time-domain to frequency-domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time-domain functions to frequency-domain functions;
[0249] After the second variable sound is decomposed into the plurality of analysis frames, time-domain to frequency-domain conversion is performed on each analysis frame to convert the plurality of analysis frames from time-domain functions to frequency-domain functions. As an embodiment, the plurality of analysis frames are converted from time-domain functions to frequency-domain functions by short-time Fourier transform in this application.
[0250] For convenience of description, it is assumed that each analysis frame is wherein is the time before the phase information is modified, u represents the frame number, Ω k represents the angular frequency of the kth frequency point in each analysis frame, φ a represents the phase of the analysis frame, and A represents the amplitude of each frame.
[0251] 1004、keeping the spectral amplitude of the analysis frame unchanged and modifying the phase information of the analysis frame to obtain a plurality of synthesis frames;
[0252] After the frequency-domain function of the analysis frame is obtained, the spectral amplitude of the analysis frame is kept unchanged, that is, |A| is unchanged, and the phase information of the analysis frame is modified, that is, φ is modified. Thus, variable speed and constant pitch of the second variable sound are realized.
[0253] Specifically, the process of modifying the phase information of the analysis frame will be described in the following embodiments, which will not be described herein.
[0254] 1005、the plurality of synthesis frames are superimposed by a synthesis window function to obtain a third variable sound performing variable pitch and constant speed on the original dry sound.
[0255] After the plurality of synthesis frames are obtained, the plurality of synthesis frames are superimposed by a synthesis window function, so that a third variable sound performing variable pitch and constant speed on the original dry sound is obtained.
[0256] It is assumed that is the frequency-domain signal of the synthesis frame of the uth frame, and y w (u, n) is the time-domain signal of the frequency-domain signal of the synthesis frame of the uth frame after inverse transformation, and the synthesis window function is f(n), so that After superimposed by the synthesis window function f(n), the third variable sound performing variable pitch and constant speed on the original dry sound can be obtained.
[0257] It is assumed that is the actual output signal of the uth frame, and in order to make y w (u, n) infinitely close to the actual output signal so that min, where u denotes the frame number, and n denotes the sample point index number within the frame.
[0258] Thus, the final synthesis signal is obtained as
[0259] Based on Figure 10 The process of modifying the analysis frame phase information is described in detail below with reference to step 1004 in the embodiment. Figure 11 , Figure 11 The detailed steps of step 1004 are as follows:
[0260] 1101. Calculate the phase difference between adjacent analysis frames.
[0261] In order to obtain the variable-speed constant-pitch synthesis frame, the phase information of the analysis frame can be modified, and the phase difference between adjacent analysis frames is calculated first: wherein φu denotes the phase of the u-th analysis frame, φu-1 denotes the phase of the (u-1)-th analysis frame, h a Δu denotes the frame shift between adjacent analysis frames, and Ω k ωk denotes the angular frequency of the k-th frequency point, and thus Δφu denotes the phase difference between the u-th and (u-1)-th analysis frames.
[0262] 1102. Obtain the variable-pitch parameter of the second variable sound, and construct the synthesis phase of the target synthesis frame adjacent to the any synthesis frame according to the phase difference between the adjacent analysis frames, the reciprocal of the variable-pitch parameter of the second variable sound, and the any synthesis frame.
[0263] Assuming that the variable-pitch parameter of the second variable sound is β, the synthesis phase of the u-th frame is constructed using the variable-pitch parameter β of the second variable sound. First, the principal value of the argument of the phase difference between adjacent analysis frames is extracted, and the error quantity falling within the interval [-π, π] is obtained: Further, the synthesis phase of the u-th frame is constructed using the variable-pitch parameter: wherein φu denotes the phase of the u-th synthesis frame, φu-1 denotes the phase of the (u-1)-th synthesis frame, h s Δu denotes the frame shift between adjacent synthesis frames, and Ω and β are reciprocals of each other. When constructing the synthesis phase of the u-th synthesis frame, it is generally considered that the phase of the initial synthesis frame is equal to the phase of the initial analysis frame, i.e.
[0264] Thus, it can be seen that the phase of the second synthesis frame can be constructed according to the initial synthesis frame, and the phase of the third synthesis frame can be constructed according to the phase of the second synthesis frame, so as to sequentially construct the phase of each synthesis frame.
[0265] That is, according to the phase difference of the adjacent analysis frames, the reciprocal of the pitch parameter of the second variable sound, and any synthesis frame, the synthesis phase of the target synthesis frame adjacent to the synthesis frame can be constructed.
[0266] 1103. Constructing the frequency domain function of the target synthesis frame according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0267] After obtaining the synthesis phase of the target synthesis frame, because the spectrum amplitude of the analysis frame is kept unchanged during the processing of the analysis frame, the frequency domain function of the target synthesis frame can be constructed according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0268] In the embodiments of the present application, the process of performing the variable speed and variable pitch strategy on the original dry sound speech signal first, and then performing the variable speed and constant pitch strategy of the phase frequency vocoder is described in detail, thereby improving the reliability of the process of obtaining the synthesis frame according to the analysis frame in the embodiments of the present application.
[0269] II. Performing the variable speed and constant pitch strategy of the phase frequency vocoder on the original dry sound speech signal first, and then performing the variable speed and variable pitch strategy of the resampling to obtain a target object variable sound set of a target song, please refer to Figure 12 :
[0270] 1201. Framing and windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time domain functions;
[0271] 1202. Converting the plurality of analysis frames from time domain functions to frequency domain functions by converting the time domain and the frequency domain;
[0272] 1203. Keeping the spectrum amplitude of the analysis frame unchanged, modifying the phase information of the analysis frame to obtain a plurality of synthesis frames;
[0273] 1204. Overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a frequency domain signal of a fourth variable sound performing variable speed and constant pitch processing on the original dry sound;
[0274] It should be noted that the steps 1201 to 1204 in the embodiments of the present application are similar to the description of the steps 1002 to 1005 in the embodiments of the present application, except that the object of framing and windowing is changed from the second variable sound to the original dry sound, which will not be described here. Figure 10
[0275] 1205. Converting the frequency domain signal of the fourth variable sound from the frequency domain to the time domain to obtain a time domain signal of the fourth variable sound;
[0276] In order to perform resampling on the fourth variant voice, the frequency domain signal of the fourth variant voice needs to be converted into a time domain signal to perform step 1206.
[0277] 1206, performing resampling on the time domain signal of the fourth variant voice according to a target sampling rate to obtain a fifth variant voice signal which is processed by pitch variation and constant speed on the original dry voice, wherein the pitch variation parameter obtained according to the target sampling rate and the speed variation parameter of the fourth variant voice are reciprocal to each other.
[0278] After obtaining the time domain signal of the fourth variant voice, resampling is performed on the time domain signal of the fourth variant voice according to a target sampling rate to obtain a fifth variant voice signal which is processed by pitch variation and constant speed on the original dry voice, wherein the pitch variation parameter obtained according to the target sampling rate and the speed variation parameter of the fourth variant voice are reciprocal to each other.
[0279] Wherein, the process of performing resampling on the time domain signal of the fourth variant voice is similar to the process of performing resampling on the time domain signal of the third variant voice. Figure 10 The description of step 1001 in the embodiment is similar, and will not be repeated here.
[0280] For the process of obtaining the fourth variant voice, the following detailed description is made with reference to Figure 12 The description of step 1204 in the embodiment is similar, and will not be repeated here. Figure 13 , Figure 13 The detailed steps of step 1204 are as follows:
[0281] 1301, calculating the phase difference of adjacent analysis frames;
[0282] 1302, obtaining the speed variation parameter of the fourth variant voice to construct the synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the speed variation parameter of the fourth variant voice and the any synthesis frame;
[0283] 1303, constructing the frequency domain function of the target synthesis frame according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0284] Specifically, steps 1301 to 1303 in the embodiment are similar to the description of steps 1101 to 1103 in the embodiment, and will not be repeated here. Figure 11 The description of step 1001 in the embodiment is similar, and will not be repeated here.
[0285] In the embodiment, the process of performing the speed variation and constant pitch strategy of the phase frequency vocoder on the speech signal of the original dry voice before performing the resampling process of the speed variation and pitch variation strategy is described in detail, thereby improving the reliability of the process of obtaining the synthesis frame according to the analysis frame in the embodiment.
[0286] Figures 10 to 13In the embodiment, the series connection process of the variable speed and variable pitch strategy using resampling and the variable speed and constant pitch strategy of the phase frequency vocoder is described in detail. The resampling of the original dry sound can be implemented by using the variable pitch parameter β to achieve the variable speed and variable pitch of the original dry sound, and then the parameter The variable speed and variable pitch of the speech signal is pulled back to the original length, and the expansion and contraction of the signal time domain can be implemented by using the variable speed parameter β', and then the parameter The resampling of the variable speed signal after expansion and contraction makes the length of the resampled signal the same as the length of the signal before expansion and contraction, and only the pitch changes.
[0287] For the Figures 10 to 13 In the embodiment, different strategies can be used to achieve different voice effects in the process of implementing the variable pitch and constant speed of the original dry sound. The following describes the different strategies:
[0288] I. If the audio in the target object voice set of the target song is the rising pitch voice of the original dry sound according to the target object's voice tone, the resampling variable speed and variable pitch strategy is first performed on the original dry sound in the original dry sound set, and then the variable speed and constant pitch strategy of the phase frequency vocoder is performed.
[0289] II. If the audio in the target object voice set of the target song is the falling pitch voice of the original dry sound according to the target object's voice tone, the resampling variable speed and variable pitch strategy is first performed on the original dry sound in the original dry sound set, and then the variable speed and constant pitch strategy of the phase frequency vocoder is performed.
[0290] Because according to multiple experiments, the effects of the voice processing when performing rising pitch and falling pitch voice are shown in Table 1:
[0291] Table 1
[0292]
[0293]
[0294] Therefore, in the embodiment, different variable pitch and constant speed processing is performed on the original dry sound according to the different voice tones of the target object, thereby improving the naturalness of different voices.
[0295] The audio processing method in the embodiment is described in detail above, and the audio processing device in the present application is described below. Please refer to Figure 14 One embodiment of the audio processing device in the embodiment includes:
[0296] The acquisition unit 1401 is configured to acquire a to-be-synthesized voice, and the to-be-synthesized voice is an audio obtained by performing variable pitch and constant speed processing on the original dry sound corresponding to a target song according to a voice tone of a target object.
[0297] a decomposition unit 1402, configured to acquire a duet phrase period template of the target song, and decompose the to-be-synthesized variable voice into a target object phrase and a user phrase according to the duet phrase period template;
[0298] a processing unit 1403, configured to retain the target object phrase in the to-be-synthesized variable voice, and perform mute processing on the user phrase in the to-be-synthesized variable voice;
[0299] a mixing unit 1404, configured to mix the retained target object phrase with a target song accompaniment according to a preset energy ratio, so that a loudness of a volume of the mixed accompaniment is the same as an original loudness of the target song accompaniment, and is not lower than a preset loudness standard;
[0300] an output unit 1405, configured to output the mixed accompaniment audio with the target object variable voice.
[0301] Optionally, the audio processing apparatus further includes:
[0302] The processing unit 1403 is further configured to perform first processing on the to-be-synthesized variable voice to generate the duet phrase period template of the target song, where the first processing is configured to acquire timestamp information of an effective voice phrase in the to-be-synthesized variable voice.
[0303] Optionally, the processing unit 1403 is specifically configured to:
[0304] perform smoothing processing on a speech signal of the to-be-synthesized variable voice to acquire an amplitude envelope of the to-be-synthesized variable voice;
[0305] perform processing on the amplitude envelope of the to-be-synthesized variable voice by using a threshold function to acquire the timestamp information of the effective voice phrase in the to-be-synthesized variable voice;
[0306] perform duet period marking on the to-be-synthesized variable voice according to the timestamp information of the effective voice phrase in the to-be-synthesized variable voice to generate the duet phrase period template of the target song.
[0307] Optionally, the processing unit 1403 is further configured to:
[0308] perform low-pass filtering on the speech signal of the to-be-synthesized variable voice to acquire a to-be-synthesized variable voice signal of a target frequency band.
[0309] Optionally, the processing unit 1403 is further configured to:
[0310] perform normalization processing on an amplitude value of the speech signal of the to-be-synthesized variable voice to obtain a normalized to-be-synthesized variable voice.
[0311] Optionally, the acquisition unit 1401 is further configured to:
[0312] acquire time lyric information of the to-be-synthesized variable voice;
[0313] Optionally, the processing unit 1403 is further configured to:
[0314] perform a duet time period marking on the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice and the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice, to generate a duet phrase group time period template of the target song.
[0315] Optionally, the processing unit 1403 is specifically configured to:
[0316] perform an alignment operation on the timestamp information of the effective voice phrase group in the to-be-synthesized variable voice and the timestamp information in the time lyric information of the to-be-synthesized variable voice, to obtain corrected timestamp information of the effective voice phrase group;
[0317] perform a duet time period marking on the to-be-synthesized variable voice according to the corrected timestamp information of the effective voice phrase group, to generate a duet phrase group time period template of the target song.
[0318] Optionally, the audio processing apparatus further includes:
[0319] a setting unit 1406 configured to set the loudness of the target object phrase group to be less than the loudness of the accompaniment of the target song and lower than the loudness of the accompaniment of the target song by a preset value.
[0320] Optionally, the acquisition unit 1401 is further configured to:
[0321] acquire time lyric information of the to-be-synthesized variable voice;
[0322] The processing unit 1403 is further configured to:
[0323] perform marking on time lyric information of the target object phrase group in the time lyric information of the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice, the retained target object phrase group and the user phrase group after the mute processing, to generate time lyric information with target object marking.
[0324] Optionally, the processing unit 1403 is specifically configured to:
[0325] perform marking on a timestamp position of lyrics of the target object phrase group in the time lyric information of the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice, the retained target object phrase group and the user phrase group after the mute processing, by using a note file or a midi file, to generate time lyric information with target object marking.
[0326] Optionally, the audio processing apparatus further comprises:
[0327] a voice changing unit 1407, configured to perform a pitch change invariant speed processing on original dry voices of a target song according to a voice tone of a target object, to obtain a target object voice changing set of the target song;
[0328] a screening unit 1408, configured to screen out to-be-synthesized changed voices meeting a preset standard from the target object voice changing set of the target song;
[0329] Optionally, the screening unit 1408 is further configured to:
[0330] screen out an original dry voice set meeting a preset melody standard and a preset tone standard from a dry voice set of the target song;
[0331] The screening unit 1408 is specifically configured to:
[0332] screen out to-be-synthesized changed voices meeting a preset standard from the target object voice changing set of the target song according to a naturalness and pleasantness standard of a listening.
[0333] Optionally, the preset melody standard comprises a preset pitch standard and a preset rhythm standard.
[0334] The naturalness and pleasantness standard of the listening comprises at least one of a smoothness of an audio breath, a cosine similarity of a fundamental frequency sequence and a stability of an audio beat distribution.
[0335] Optionally, the pitch change invariant speed processing comprises a speed and pitch change processing of resampling and a speed change invariant pitch processing of a phase frequency vocoder.
[0336] Optionally, the voice changing unit 1407 is specifically configured to:
[0337] perform the speed and pitch change processing of resampling on a speech signal of the original dry voice first, and then perform the speed change invariant pitch processing of the phase frequency vocoder;
[0338] or,
[0339] perform the speed change invariant pitch processing of the phase frequency vocoder on the speech signal of the original dry voice first, and then perform the speed and pitch change processing of resampling.
[0340] Optionally, the voice changing unit 1407 is specifically configured to:
[0341] perform resampling on a time domain speech signal of the original dry voice, to obtain a second changed voice after speed and pitch change;
[0342] frame and window the second changed voice, to decompose the second changed voice into a plurality of analysis frames;
[0343] performing time-domain to frequency-domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time-domain functions to frequency-domain functions;
[0344] keeping the spectral amplitudes of the analysis frames unchanged, modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0345] overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a third variable sound subjected to variable pitch invariant speed processing on the original dry sound.
[0346] Optionally, the variable sound unit 1407 is specifically configured to:
[0347] calculating a phase difference of adjacent analysis frames;
[0348] obtaining a variable pitch parameter of the second variable sound to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, a reciprocal of the variable pitch parameter of the second variable sound, and the any synthesis frame;
[0349] constructing a frequency-domain function of the target synthesis frame according to the spectral amplitudes of the analysis frames and the synthesis phase of the target synthesis frame to obtain the plurality of synthesis frames.
[0350] Optionally, the variable sound unit 1407 is specifically configured to:
[0351] windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time-domain functions;
[0352] performing time-domain to frequency-domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time-domain functions to frequency-domain functions;
[0353] keeping the spectral amplitudes of the analysis frames unchanged, modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0354] overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a fourth variable sound subjected to variable speed invariant pitch processing on the original dry sound.
[0355] performing resampling on a time-domain signal of the fourth variable sound according to a target sampling rate to obtain a fifth variable sound subjected to variable pitch invariant speed processing on the original dry sound, wherein a variable pitch parameter obtained according to the target sampling rate and a variable speed parameter of the fourth variable sound are reciprocals of each other.
[0356] Optionally, the variable sound unit 1407 is specifically configured to:
[0357] calculating a phase difference of adjacent analysis frames;
[0358] obtaining a variable speed parameter of the fourth variable voice, and constructing a synthesis phase of a target synthesis frame adjacent to the any synthesis frame according to a phase difference of the adjacent analysis frames, the variable speed parameter of the fourth variable voice and the any synthesis frame;
[0359] constructing a frequency domain function of the target synthesis frame according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame, to obtain a plurality of synthesis frames.
[0360] Optionally, if the audio in the target object variable voice set of the target song is a variable voice of raising the pitch of the original dry voice according to the target object's vocal tone, the variable voice unit 1407 is specifically configured to:
[0361] performing the variable speed and constant tone strategy of the phase frequency vocoder on the original dry voice in the original dry voice set first, and then performing the variable speed and variable tone strategy of the resampling.
[0362] Optionally, if the audio in the target object variable voice set of the target song is a variable voice of lowering the pitch of the original dry voice according to the target object's vocal tone, the variable voice unit 1407 is specifically configured to:
[0363] performing the variable speed and constant tone strategy of the resampling on the original dry voice in the original dry voice set first, and then performing the variable speed and variable tone strategy of the phase frequency vocoder.
[0364] It should be noted that the functions of the above units are similar to those described in the embodiments of the present application, and will not be described here. Figures 1 to 13
[0365] In the embodiments of the present application, the variable voice unit 1407 first performs the variable tone and constant speed processing on the original dry voice of the target song according to the target object's vocal tone to obtain the target object variable voice set of the target song, and the filtering unit 1408 filters the target object variable voice set of the target song according to the preset standard to obtain the to-be-synthesized variable voice that meets the preset standard, and the decomposition unit 1402 decomposes the to-be-synthesized variable voice into the target object sentence group and the user sentence group according to the singing duet sentence group period template of the target song, and the processing unit 1403 retains the target object sentence group in the to-be-synthesized variable voice and performs the mute processing on the user sentence group in the to-be-synthesized variable voice, and finally the mixing unit 1404 mixes the retained target object sentence group with the accompaniment of the target song according to the preset energy ratio, thereby obtaining the accompaniment audio with the target object variable voice, thereby improving the convenience of obtaining the accompaniment audio with the target object variable voice.
[0366] The above describes the audio processing device in the embodiments of the present application from the perspective of modular functional entities, and the following describes the computer device in the embodiments of the present application from the perspective of hardware processing:
[0367] The computer device is used to realize the function of the audio processing device, and one embodiment of the computer device in the embodiment of the application comprises:
[0368] a processor and a memory;
[0369] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the following steps can be realized:
[0370] obtaining to-be-synthesized variable sound, wherein the to-be-synthesized variable sound is an audio obtained by performing variable-pitch constant-speed processing on original dry sound corresponding to a target song according to the pitch of a target object;
[0371] obtaining a duet sentence group time period template of the target song, and decomposing the to-be-synthesized variable sound into a target object sentence group and a user sentence group according to the duet sentence group time period template;
[0372] retaining the target object sentence group in the to-be-synthesized variable sound, and performing mute processing on the user sentence group in the to-be-synthesized variable sound;
[0373] mixing the retained target object sentence group with a target song accompaniment according to a preset energy ratio, so that the loudness of the volume of the mixed accompaniment is the same as the original loudness of the target song accompaniment, and is not lower than a preset loudness standard;
[0374] outputting the mixed accompaniment audio with the target object variable sound.
[0375] In some embodiments of the application, the processor can also be used to realize the following steps:
[0376] performing first processing on the to-be-synthesized variable sound to generate a duet sentence group time period template of the target song, wherein the first processing is used to obtain timestamp information of an effective sound sentence group in the to-be-synthesized variable sound.
[0377] In some embodiments of the application, the processor can also be used to realize the following steps:
[0378] performing smoothing processing on a speech signal of the to-be-synthesized variable sound to obtain an amplitude envelope of the to-be-synthesized variable sound;
[0379] processing the amplitude envelope of the to-be-synthesized variable sound by using a threshold function to obtain timestamp information of an effective sound sentence group in the to-be-synthesized variable sound;
[0380] performing duet time period marking on the to-be-synthesized variable sound according to the timestamp information of the effective sound sentence group in the to-be-synthesized variable sound to generate a duet sentence group time period template of the target song.
[0381] In some embodiments of the application, the processor can also be used to realize the following steps:
[0382] Low-pass filter the voice signal to be synthesized to obtain a target frequency band voice signal to be synthesized.
[0383] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0384] Normalizing the amplitude of the voice signal to be synthesized to obtain a normalized voice signal to be synthesized.
[0385] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0386] Obtaining time lyric information of the voice signal to be synthesized;
[0387] Performing a duet time period marking on the voice signal to be synthesized according to the time lyric information of the voice signal to be synthesized and the timestamp information of the effective voice sentence group in the voice signal to be synthesized to generate a duet sentence group time period template of the target song.
[0388] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0389] Aligning the timestamp information of the effective voice sentence group in the voice signal to be synthesized with the timestamp information in the time lyric information of the voice signal to be synthesized to obtain corrected timestamp information of the effective voice sentence group;
[0390] Performing a duet time period marking on the voice signal to be synthesized according to the corrected timestamp information of the effective voice sentence group to generate a duet sentence group time period template of the target song.
[0391] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0392] Setting the loudness of the target object sentence group to be less than the loudness of the accompaniment of the target song and lower than the loudness of the accompaniment of the target song by a preset value.
[0393] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0394] Obtaining time lyric information of the voice signal to be synthesized;
[0395] In some embodiments of the present application, the processor can further be configured to perform the following steps:
[0396] Marking the time lyric information of the target object sentence group in the time lyric information of the voice signal to be synthesized according to the time lyric information of the voice signal to be synthesized, the retained target object sentence group and the user sentence group after the mute processing to generate time lyric information with a target object marker.
[0397] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0398] According to the time lyrics information of the to-be-synthesized variable voice, the reserved target object sentence group and the user sentence group after the mute processing, the lyrics of the target object sentence group in the time stamp position in the time lyrics information of the to-be-synthesized variable voice is marked by using a note file or a midi file, so as to generate the time lyrics information with the target object mark.
[0399] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0400] Performing variable pitch and constant speed processing on the original dry voice of the target song according to the pitch of the target object, so as to obtain a target object variable voice set of the target song.
[0401] Screening the to-be-synthesized variable voice meeting the preset standard from the target object variable voice set of the target song.
[0402] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0403] Screening an original dry voice set meeting the preset melody standard and the preset timbre standard from the dry voice set of the target song.
[0404] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0405] According to the natural degree and the pleasant degree of listening, screening the to-be-synthesized variable voice meeting the preset standard from the target object variable voice set of the target song.
[0406] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0407] Performing the variable speed and variable pitch strategy of the resampling on the speech signal of the original dry voice first, and then performing the variable speed and constant pitch strategy of the phase frequency vocoder.
[0408] Or,
[0409] Performing the variable speed and constant pitch strategy of the phase frequency vocoder on the speech signal of the original dry voice first, and then performing the variable speed and variable pitch strategy of the resampling.
[0410] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0411] Performing resampling on the time domain speech signal of the original dry voice, so as to obtain a second variable voice after variable speed and variable pitch.
[0412] windowing the second variable sound to decompose the second variable sound into a plurality of analysis frames;
[0413] performing time domain to frequency domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time domain functions to frequency domain functions;
[0414] keeping the spectral amplitudes of the analysis frames unchanged and modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0415] overlapping and adding the plurality of synthesis frames by a synthesis window function to obtain a third variable sound performing pitch variation invariant speed processing on the original dry sound.
[0416] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0417] calculating the phase difference of adjacent analysis frames;
[0418] obtaining the pitch variation parameter of the second variable sound to construct the synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the reciprocal of the pitch variation parameter of the second variable sound and the any synthesis frame;
[0419] constructing the frequency domain function of the target synthesis frame according to the spectral amplitudes of the analysis frames and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0420] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0421] windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time domain functions;
[0422] performing time domain to frequency domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time domain functions to frequency domain functions;
[0423] keeping the spectral amplitudes of the analysis frames unchanged and modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0424] overlapping and adding the plurality of synthesis frames by a synthesis window function to obtain a fourth variable sound performing speed variation invariant pitch processing on the original dry sound;
[0425] performing resampling on the time domain signal of the fourth variable sound according to a target sampling rate to obtain a fifth variable sound performing pitch variation invariant speed processing on the original dry sound, wherein the pitch variation parameter obtained according to the target sampling rate and the speed variation parameter of the fourth variable sound are reciprocal to each other.
[0426] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0427] calculating a phase difference of adjacent analysis frames;
[0428] obtaining a time-varying parameter of the fourth time-varying sound, to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the time-varying parameter of the fourth time-varying sound and the any synthesis frame;
[0429] constructing a frequency domain function of the target synthesis frame according to the spectral amplitude of the analysis frame and the synthesis phase of the target synthesis frame, to obtain a plurality of synthesis frames.
[0430] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0431] The time-varying and constant pitch processing of the phase vocoder is performed on the original dry sound in the original dry sound set first, and then the time-varying and variable pitch processing of the resampling is performed.
[0432] In some embodiments of the present application, the processor can be further configured to implement the following steps:
[0433] The time-varying and variable pitch processing of the resampling is performed on the original dry sound in the original dry sound set first, and then the time-varying and constant pitch processing of the phase vocoder is performed.
[0434] It can be understood that the processor in the computer device described above can also implement the functions of each unit in the corresponding device embodiments described above when the processor executes the computer program, and details are not repeated here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the audio processing device. For example, the computer program can be divided into the units in the audio processing device described above, and each unit can implement the specific functions as described above in the corresponding audio processing device.
[0435] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The computer device can include but is not limited to a processor and a memory. Those skilled in the art can understand that the processor and the memory are only examples of the computer device, and do not constitute a limitation on the computer device, and can include more or fewer components, or combine certain components, or different components, for example, the computer device can also include an input / output device, a network access device, a bus and the like.
[0436] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device and connects various parts of the entire computer device using various interfaces and lines.
[0437] The memory can be used to store the computer programs and / or modules, and the processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0438] The present invention further provides a computer-readable storage medium, which is used to implement the functions of the audio processing device and stores a computer program. When the computer program is executed by a processor, the processor can be configured to perform the following steps:
[0439] Obtaining a modified voice to be synthesized, wherein the modified voice to be synthesized is audio obtained by performing pitch-shifting without speed-shifting processing on an original dry voice corresponding to a target song according to the pitch of a target object;
[0440] Obtaining a duet sentence group time period template of the target song, and decomposing the voice-changing to-be-synthesized into a target object sentence group and a user sentence group according to the duet sentence group time period template;
[0441] retaining the target sentence group in the voice-changing synthesis to be performed, and performing a muting process on the user sentence group in the voice-changing synthesis to be performed;
[0442] The reserved target object sentence group is mixed with the target song accompaniment according to a preset energy ratio, so that the loudness of the mixed accompaniment volume is the same as the original loudness of the target song accompaniment and is not lower than a preset loudness standard.
[0443] The accompaniment audio with the target object voice change after mixing is output.
[0444] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0445] The first processing is performed on the to-be-synthesized voice change to generate the duet sentence group period template of the target song, wherein the first processing is used to acquire the timestamp information of the effective sound sentence group in the to-be-synthesized voice change.
[0446] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0447] The smoothing processing is performed on the voice signal of the to-be-synthesized voice change to acquire the amplitude envelope of the to-be-synthesized voice change.
[0448] The threshold function is used to process the amplitude envelope of the to-be-synthesized voice change to acquire the timestamp information of the effective sound sentence group in the to-be-synthesized voice change.
[0449] According to the timestamp information of the effective sound sentence group in the to-be-synthesized voice change, the duet period marking is performed on the to-be-synthesized voice change to generate the duet sentence group period template of the target song.
[0450] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0451] The low-pass filtering is performed on the voice signal of the to-be-synthesized voice change to acquire the to-be-synthesized voice change signal of the target frequency band.
[0452] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0453] The amplitude of the voice signal of the to-be-synthesized voice change is normalized to obtain the normalized to-be-synthesized voice change.
[0454] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0455] acquire time lyric information of the to-be-synthesized variable voice;
[0456] perform a duet time period marking on the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice and the timestamp information of the effective voice sentence group in the to-be-synthesized variable voice, to generate a duet sentence group time period template of the target song.
[0457] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used for performing the following steps:
[0458] perform an alignment operation on the timestamp information of the effective voice sentence group in the to-be-synthesized variable voice and the timestamp information in the time lyric information of the to-be-synthesized variable voice, to obtain corrected timestamp information of the effective voice sentence group;
[0459] perform a duet time period marking on the to-be-synthesized variable voice according to the corrected timestamp information of the effective voice sentence group, to generate a duet sentence group time period template of the target song.
[0460] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used for performing the following steps:
[0461] set the loudness of the target object sentence group to be less than the loudness of the accompaniment of the target song, and lower than the loudness of the accompaniment of the target song by a preset value.
[0462] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used for performing the following steps:
[0463] acquire time lyric information of the to-be-synthesized variable voice;
[0464] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used for performing the following steps:
[0465] perform a marking on the time lyric information of the target object sentence group in the time lyric information of the to-be-synthesized variable voice according to the time lyric information of the to-be-synthesized variable voice, the retained target object sentence group and the user sentence group after the mute processing, to generate time lyric information with a target object marking.
[0466] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used for performing the following steps:
[0467] According to the time lyrics information to be synthesized, the reserved target object sentence group and the user sentence group after the mute processing, the note file or the midi file is used to mark the timestamp position of the lyrics of the target object sentence group in the time lyrics information to be synthesized, so as to generate the time lyrics information with the target object mark.
[0468] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0469] Performing the variable pitch and constant speed processing on the original dry sound of the target song according to the pitch of the target object to obtain a target object variable sound set of the target song;
[0470] Screening the to-be-synthesized variable sound meeting the preset standard from the target object variable sound set of the target song.
[0471] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0472] Screening the original dry sound set meeting the preset melody standard and the preset sound quality standard from the dry sound set of the target song;
[0473] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0474] Screening the to-be-synthesized variable sound meeting the preset standard from the target object variable sound set of the target song according to the natural degree and the pleasant degree of listening;
[0475] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0476] Performing the variable speed and variable pitch strategy of the resampling on the speech signal of the original dry sound first, and then performing the variable speed and constant pitch strategy of the phase frequency vocoder;
[0477] Or,
[0478] Performing the variable speed and constant pitch strategy of the phase frequency vocoder on the speech signal of the original dry sound first, and then performing the variable speed and variable pitch strategy of the resampling.
[0479] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically used to execute the following steps:
[0480] performing resampling on a time-domain speech signal of the original dry sound to obtain a second variable sound after variable pitch and variable speed processing;
[0481] windowing the second variable sound to decompose the second variable sound into a plurality of analysis frames;
[0482] performing time-domain to frequency-domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time-domain functions to frequency-domain functions;
[0483] keeping the spectral amplitudes of the analysis frames unchanged and modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0484] overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a third variable sound after variable pitch and constant speed processing on the original dry sound.
[0485] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically configured to perform the following steps:
[0486] calculating the phase difference of adjacent analysis frames;
[0487] obtaining the variable pitch parameter of the second variable sound to construct the synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the reciprocal of the variable pitch parameter of the second variable sound and the any synthesis frame;
[0488] constructing the frequency-domain function of the target synthesis frame according to the spectral amplitude of the analysis frame and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0489] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically configured to perform the following steps:
[0490] windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time-domain functions;
[0491] performing time-domain to frequency-domain conversion on the plurality of analysis frames to convert the plurality of analysis frames from time-domain functions to frequency-domain functions;
[0492] keeping the spectral amplitudes of the analysis frames unchanged and modifying the phase information of the analysis frames to obtain a plurality of synthesis frames;
[0493] overlapping and adding the plurality of synthesis frames through a synthesis window function to obtain a fourth variable sound after constant pitch and variable speed processing on the original dry sound;
[0494] Resampling the fourth time-domain signal according to a target sampling rate to obtain a fifth time-domain signal, wherein the fifth time-domain signal is processed by pitch shifting and constant speed variation according to the target sampling rate, and the pitch shifting parameter according to the target sampling rate and the speed variation parameter of the fourth time-domain signal are reciprocal to each other.
[0495] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically configured to execute the following steps:
[0496] Calculating the phase difference of adjacent analysis frames;
[0497] Obtaining the speed variation parameter of the fourth time-domain signal, and constructing the synthesis phase of the target synthesis frame adjacent to any synthesis frame according to the phase difference of the adjacent analysis frames, the speed variation parameter of the fourth time-domain signal and the any synthesis frame;
[0498] Constructing the frequency domain function of the target synthesis frame according to the spectrum amplitude of the analysis frame and the synthesis phase of the target synthesis frame to obtain a plurality of synthesis frames.
[0499] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically configured to execute the following steps:
[0500] First performing the speed variation and constant pitch processing of the phase-frequency vocoder on the original dry sound in the original dry sound set, and then performing the resampling speed variation and pitch variation processing.
[0501] In some embodiments of the present application, when the computer program stored in the computer readable storage medium is executed by the processor, the processor can be specifically configured to execute the following steps:
[0502] First performing the resampling speed variation and pitch variation processing on the original dry sound in the original dry sound set, and then performing the speed variation and constant pitch processing of the phase-frequency vocoder.
[0503] It can be understood that the integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a corresponding computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned corresponding embodiment methods of the present application can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0504] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, another division mode can be used. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0505] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place, or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0506] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0507] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced by equivalent technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio processing method, characterized in that: include: Obtaining a modified voice to be synthesized, wherein the modified voice to be synthesized is audio obtained by performing pitch-shifting without speed-shifting processing on an original dry voice corresponding to a target song according to the pitch of a target object; Obtaining a duet sentence group time period template of the target song, and decomposing the voice-changing to-be-synthesized into a target object sentence group and a user sentence group according to the duet sentence group time period template; retaining the target sentence group in the voice-changing synthesis to be performed, and performing a muting process on the user sentence group in the voice-changing synthesis to be performed; Mixing the retained target object sentence group with the target song accompaniment according to a preset energy ratio, so that the loudness of the accompaniment after mixing is the same as the original loudness of the target song accompaniment and is not lower than a preset loudness standard; Output the mixed accompaniment audio with the changed voice of the target object.
2. The method according to claim 1, characterized in that Before obtaining the duet phrase group time period template of the target song, the method further includes: A first process is performed on the voice change to be synthesized to generate a duet phrase group time period template of the target song, wherein the first process is used to obtain timestamp information of the valid voice phrase group in the voice change to be synthesized.
3. The method according to claim 2, characterized in that The performing of a first processing on the voice-changing to-be-synthesized to generate a duet phrase group time period template of the target song includes: Performing smoothing processing on the speech signal to be synthesized and modified to obtain the amplitude envelope of the voice to be synthesized; Processing the amplitude envelope of the voice to be synthesized using a threshold function to obtain timestamp information of valid voice sentence groups in the voice to be synthesized; According to the timestamp information of the valid voice sentence group in the voice change to be synthesized, the duet period is marked on the voice change to be synthesized to generate a duet sentence group period template of the target song.
4. The method according to claim 3, characterized in that Before performing smoothing processing on the speech signal to be synthesized and modified to obtain the amplitude envelope of the voice to be synthesized and modified, the method further includes: The speech signal to be synthesized and voice-changed is low-pass filtered to obtain a voice-changed signal to be synthesized in a target frequency band.
5. The method according to claim 4, characterized in that Before low-pass filtering the voice signal to be synthesized and voice-changed, the method further includes: The amplitude of the speech signal to be synthesized and modified is normalized to obtain a normalized voice to be synthesized and modified.
6. The method according to claim 3, characterized in that The method of marking the duet period of the voice-changing song to be synthesized according to the timestamp information of the valid voice sentence group in the voice-changing song to be synthesized, so as to generate a duet sentence group period template of the target song, includes: Obtaining the time lyrics information of the voice-changing to be synthesized; According to the time lyrics information of the voice change to be synthesized and the timestamp information of the valid sound sentence group in the voice change to be synthesized, the duet period is marked on the voice change to be synthesized to generate a duet sentence group period template of the target song.
7. The method according to claim 6, characterized in that The method of performing duet period marking on the voice change to be synthesized according to the time lyrics information of the voice change to be synthesized and the timestamp information of the valid voice sentence group in the voice change to be synthesized to generate a duet sentence group period template of the target song includes: Aligning the timestamp information of the valid sound sentence group in the voice-changing to-be-synthesized with the timestamp information in the time lyrics information of the voice-changing to-be-synthesized to obtain the corrected timestamp information of the valid sound sentence group; According to the corrected timestamp information of the valid voice sentence group, the duet period is marked on the voice to be synthesized to generate a duet sentence group period template of the target song.
8. The method according to claim 1, characterized in that Before mixing the retained target sentence group with the target song accompaniment according to a preset energy ratio, the method further includes: The loudness of the target sentence group is set to be smaller than the loudness of the target song accompaniment and lower by a preset value than the loudness of the target song accompaniment.
9. The method according to claim 1, characterized in that The method further comprises: Obtaining the time lyrics information of the voice-changing to be synthesized; After retaining the target sentence group in the voice-changing synthesis to be performed and performing muting processing on the user sentence group in the voice-changing synthesis to be performed, the method further includes: According to the time lyrics information of the voice-changed to be synthesized, the retained target sentence group and the user sentence group after the mute processing, the time lyrics information of the target sentence group in the time lyrics information of the voice-changed to be synthesized is marked to generate time lyrics information with the target object mark.
10. The method according to claim 9, characterized in that The method of marking the time lyrics information of the target object sentence group in the time lyrics information of the voice-changed song to be synthesized, the retained target object sentence group, and the user sentence group after the mute processing to generate time lyrics information with the target object mark includes: According to the time lyrics information of the voice-changed to be synthesized, the retained target sentence group and the user sentence group after the mute processing, the lyrics of the target sentence group are marked with the timestamp position in the time lyrics information of the voice-changed to be synthesized using a note file or a midi file to generate time lyrics information with the target object mark.
11. The method according to any one of claims 1 to 10, characterized in that Before obtaining the voice to be synthesized, the method further includes: Performing pitch-shifting but speed-invariant processing on the original dry voice of the target song according to the pitch of the target object to obtain a target object voice-shifting set of the target song; The voice changes to be synthesized that meet the preset standards are screened out from the target object voice change set of the target song.
12. The method according to claim 11, characterized in that Before performing pitch-changing processing on the original dry voice of the target song according to the pitch of the target object without changing the speed to obtain the target object voice-changing set of the target song, the method further includes: Filtering out an original dry audio set that meets a preset melody standard and a preset sound quality standard from a dry audio set of a target song; Screening out the target object voice variation set of the target song that meets the preset criteria for synthesizing the target voice variation includes: According to the naturalness and pleasure of listening standards, the voice changes to be synthesized that meet the preset standards are screened from the target object voice change set of the target song.
13. The method according to claim 12, characterized in that The preset melody standard includes a preset pitch standard and a preset rhythm standard; The naturalness and pleasure of listening standards include at least one of the smoothness of the audio breath, the cosine similarity of the fundamental frequency sequence, and the stability of the audio beat distribution.
14. The method according to claim 11, characterized in that The pitch-changing without speed-invariant processing includes: speed-changing pitch-changing processing of resampling and speed-changing pitch-changing processing of phase-frequency vocoder.
15. The method according to claim 14, characterized in that The step of performing pitch-changing processing on the original dry sound of the target song according to the pitch of the target object without changing the speed includes: Firstly executing the resampled speed-changing and pitch-changing strategy on the original dry voice signal, and then executing the speed-changing and pitch-invariant strategy of the phase-frequency vocoder; or, The original dry voice signal is first subjected to the variable speed but non-invariant pitch strategy of the phase-frequency vocoder, and then subjected to the variable speed and pitch strategy of the resampling.
16. The method according to claim 15, characterized in that The resampling speed-changing and pitch-changing strategy is first executed on the original dry voice signal, and then the speed-changing and pitch-changing strategy of the phase-frequency vocoder is executed, including: Resampling the original dry voice time domain speech signal to obtain a second modified voice with a changed speed and pitch; framing and windowing the second voice change to decompose the second voice change into a plurality of analysis frames; Performing a time domain to frequency domain conversion on the multiple analysis frames to convert the multiple analysis frames from a time domain function to a frequency domain function; Maintaining the spectrum amplitude of the analysis frame unchanged, and modifying the phase information of the analysis frame to obtain a plurality of synthetic frames; The plurality of synthesized frames are overlapped and added through a synthesis window function to obtain a third modified sound by performing pitch shifting but not speed shifting processing on the original dry sound.
17. The method according to claim 16, characterized in that The modifying of the phase information of the analysis frame to obtain a plurality of synthetic frames includes: Calculate the phase difference between adjacent analysis frames; Acquire the pitch shift parameter of the second voice change, so as to construct a synthetic phase of a target synthetic frame adjacent to any synthetic frame according to the phase difference between the adjacent analysis frames, the inverse of the pitch shift parameter of the second voice change, and any synthetic frame; A frequency domain function of the target synthetic frame is constructed according to the frequency spectrum amplitude of the analysis frame and the synthetic phase of the target synthetic frame to obtain a plurality of synthetic frames.
18. The method according to claim 15, characterized in that The method of first performing the variable speed but non-invariant pitch strategy of the phase-frequency vocoder on the original dry voice signal and then performing the variable speed and non-invariant pitch strategy of the resampling includes: framing and windowing the original dry sound to decompose the original dry sound into a plurality of analysis frames, wherein the plurality of analysis frames are time domain functions; Performing a time domain to frequency domain conversion on the multiple analysis frames to convert the multiple analysis frames from a time domain function to a frequency domain function; Maintaining the spectrum amplitude of the analysis frame unchanged, and modifying the phase information of the analysis frame to obtain a plurality of synthetic frames; Overlapping and adding the plurality of synthesized frames through a synthesis window function to obtain a fourth modified sound by performing speed-variable but pitch-invariant processing on the original dry sound; The time domain signal of the fourth modified sound is resampled according to a target sampling rate to obtain a fifth modified sound by performing pitch shifting without speed change processing on the original dry sound, wherein the pitch shift parameter obtained according to the target sampling rate and the speed change parameter of the fourth modified sound are reciprocals of each other.
19. The method according to claim 18, characterized in that The modifying of the phase information of the analysis frame to obtain a plurality of synthetic frames includes: Calculate the phase difference between adjacent analysis frames; Acquiring a speed parameter of the fourth sound change to construct a synthesis phase of a target synthesis frame adjacent to any synthesis frame according to the phase difference between the adjacent analysis frames, the speed parameter of the fourth sound change and any synthesis frame; A frequency domain function of the target synthetic frame is constructed according to the frequency spectrum amplitude of the analysis frame and the synthetic phase of the target synthetic frame to obtain a plurality of synthetic frames.
20. The method according to claim 14, wherein If the audio in the target object voice change set of the target song is a voice change performed on the original dry voice according to the tone of the target object, then performing the pitch change without changing the speed of the original dry voice of the target song according to the tone of the target object includes: The original dry sounds in the original dry sound set are first subjected to the speed-variable but pitch-invariant processing of the phase-frequency vocoder, and then subjected to the speed-variable and pitch-invariant processing of the resampling.
21. The method according to claim 14, wherein If the audio in the target object voice change set of the target song is a voice change performed on the original dry voice according to the tone of the target object, then the pitch change without speed change processing is performed on the original dry voice of the target song according to the tone of the target object, including: The resampled speed-variable pitch-variable processing is first performed on the original dry sounds in the original dry sound set, and then the speed-variable but non-pitch-variable processing of the phase-frequency vocoder is performed.
22. A computer device comprising a processor, characterized in that: When executing the computer program stored in the memory, the processor is configured to implement the audio processing method according to any one of claims 1 to 21.
23. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the audio processing method according to any one of claims 1 to 21.
Citation Information
Patent Citations
Voice synthesis method and system
CN108269560A
Singing sound converter
CN110782866A