Audio adjustment method, computer device and program product
By selecting audio templates, recording human voice audio, obtaining basic frequency information and text timestamps, audio adjustments are made based on the duration and pitch information of the audio template, and integrating them with the template accompaniment, the problem of insufficient audio adjustment effect in the existing technology is solved, and more complex and interesting audio effects are achieved.
Patent Information
- Application Number
- CN202211090743.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-07
AI Technical Summary
When generating audio audio, the existing audio adjustment method is insufficient to achieve complex audio effects.
By selecting the audio template, recording the human voice audio, obtaining the basic frequency information and text timestamps, audio adjustments are made based on the duration and pitch information of the audio template, and fusing it with the template accompaniment to generate the target audio.
Improves the effect of ghost audio adjustments, achieving more complex and interesting audio effects.
Smart Images

Figure CN116312425B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to an audio adjustment method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] With the development of computer technology, it is currently possible to listen to and process audio through various terminal devices, such as listening to and processing songs. With the development of audio editing technology, ghost video and audio have gradually emerged. Ghosting refers to using the sounds in the materials for editing, splicing, tuning, and combining with the song accompaniment to obtain a complete video and audio work. Therefore, when a user needs to obtain a ghost audio, the audio needs to be adjusted. Currently, the way to adjust the audio and generate a ghost audio is usually to manually edit the audio in multiple time periods to obtain an adjusted ghost audio. However, adjusting the audio by time period editing can only achieve a simple repetition effect of the audio.
[0003] Therefore, the current audio adjustment method has the defect of insufficient adjustment effect. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an audio adjustment method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the adjustment effect.
[0005] In a first aspect, the present application provides an audio adjustment method, and the method includes:
[0006] Select an audio template and record the human voice audio corresponding to the audio template;
[0007] Obtain the fundamental frequency information corresponding to the human voice audio and identify the text in the human voice audio and the time stamps corresponding to the text characters;
[0008] Determine the fundamental frequency information corresponding to the text characters based on the fundamental frequency information and the time stamps corresponding to the text characters;
[0009] Adjust the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio part corresponding to each text character in the audio template to obtain target fundamental frequency information;
[0010] Perform pitch adjustment on the target fundamental frequency information according to the pitch information corresponding to each text character of the audio template and the pitch change trend of the audio template to determine the adjusted target human voice audio;
[0011] Obtain the template accompaniment corresponding to the audio template, and perform a fusion process on the target human voice audio and the template accompaniment to obtain the adjusted target audio.
[0012] In one embodiment, the selection of the audio template and the recording of the human voice audio corresponding to the template include:
[0013] Display at least one audio template to be selected;
[0014] Receive a selection instruction for the at least one audio template to be selected, and determine the selected audio template;
[0015] Display the text corresponding to the audio template, and record the human voice audio input by the user based on the text corresponding to the audio template.
[0016] In one embodiment, the recognition of the text and the time stamps corresponding to the text characters in the human voice audio includes:
[0017] Identify the original text corresponding to the human voice audio according to the fundamental frequency information;
[0018] Modify the original text according to the matching result between the original text and the text corresponding to the audio template to obtain the text corresponding to the human voice audio; each text character in the text corresponding to the human voice audio matches each text character in the text corresponding to the audio template;
[0019] Determine the time stamps of each text character in the human voice audio according to the duration of the audio corresponding to each text character in the text corresponding to the human voice audio.
[0020] In one embodiment, the obtaining of the fundamental frequency information corresponding to the human voice audio includes:
[0021] Obtain the pitch information, timbre information, and vocalization feature information corresponding to the human voice audio as the fundamental frequency information.
[0022] In one embodiment, the adjustment of the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio part corresponding to each text character in the audio template to obtain the target fundamental frequency information includes:
[0023] Obtain the pitch information, timbre information, and vocalization feature information of the audio part of each text character in the human voice audio corresponding to the corresponding text character in the audio template as the fundamental frequency information corresponding to each text character in the human voice audio;
[0024] Perform one-dimensional linear interpolation processing on the pitch information, timbre information, and vocalization feature information of each text character in the human voice audio respectively according to the duration information of the audio part corresponding to each text character in the audio template, so that the duration of the pitch information, the duration of the timbre information, and the duration of the vocalization feature information respectively match the duration information of the audio part corresponding to each text character in the audio template;
[0025] The interpolated target pitch information, target timbre information, and target vocalization feature information are used as target fundamental frequency information.
[0026] In one embodiment, the pitch adjustment of the target fundamental frequency information according to the pitch information corresponding to each text word of the audio template and the pitch change trend of the audio template includes:
[0027] Obtain the target pitch information in the target fundamental frequency information, and obtain the partial target pitch information corresponding to each text word in the human voice audio in the target pitch information;
[0028] For each partial target pitch information, obtain a preset number of adjacent frames adjacent to each frame of the partial target pitch information;
[0029] According to the average value of the partial target pitch information of each frame and the adjacent partial target pitch information of the adjacent frames corresponding to each frame, obtain the average pitch information corresponding to the partial target pitch information of each frame; the average pitch information represents the pitch change trend of the partial target pitch information of each frame;
[0030] According to the average pitch information and the pitch information corresponding to each text word in the audio template, perform pitch adjustment on the partial target pitch information in the target audio feature.
[0031] In one embodiment, the audio template further includes: a reference pitch and a reference frequency corresponding to the reference pitch; the obtaining of the pitch information, timbre information, and vocalization feature information corresponding to the human voice audio includes:
[0032] Obtain the text words corresponding to each character of the audio template in the text of the human voice audio;
[0033] Obtain the partial human voice audio corresponding to each text word in the human voice audio;
[0034] For the partial human voice audio corresponding to each text word, obtain the partial fundamental frequency corresponding to the partial human voice audio; according to the partial fundamental frequency corresponding to the partial human voice audio and the reference frequency, determine the pitch offset value of the pitch information corresponding to the partial human voice audio, and determine the pitch information corresponding to the partial human voice audio according to the reference pitch and the pitch offset value;
[0035] Construct an envelope matrix according to the envelope vectors of a preset number of frequency points in the partial human voice audio to obtain the timbre information;
[0036] Obtain the vocalization feature information according to the aperiodic information in a preset number of frequency bands in the partial human voice audio.
[0037] In one embodiment, the determination of the adjusted target human voice audio includes:
[0038] Determine the frequency adjustment value of the pitch-adjusted target pitch information according to the pitch-adjusted target pitch information and the reference pitch of the audio template;
[0039] Determine the adjusted target frequency corresponding to the human voice audio according to the reference frequency of the audio template and the frequency adjustment value;
[0040] Determine the adjusted target human voice audio according to the target frequency, the target timbre information, and the target vocalization feature information.
[0041] In one embodiment, the fusion processing of the target human voice audio and the template accompaniment to obtain the adjusted target audio includes:
[0042] Obtain the template beat of the template accompaniment;
[0043] Match the target human voice audio with the template accompaniment according to the template beat;
[0044] Perform mixing processing on the matched audio to obtain the adjusted target audio.
[0045] In one embodiment, the obtaining of the template beat of the template accompaniment includes:
[0046] Obtain the audio energy values at each time point in the template accompaniment, and use the time points corresponding to the audio energy values greater than the preset energy threshold as the downbeat timestamps of the template accompaniment;
[0047] Determine the template beat of the template accompaniment according to multiple downbeat timestamps.
[0048] In one embodiment, the performing of mixing processing on the matched audio to obtain the adjusted target audio includes:
[0049] Adjust the audio energy of the target human voice audio in the matched audio according to the audio energy of the template accompaniment in the matched audio, so that the audio energy of the target human voice audio is less than the audio energy of the template accompaniment;
[0050] Superimpose and mix the target human voice audio with adjusted audio energy and the template accompaniment to obtain the adjusted target audio.
[0051] In a second aspect, the present application provides an audio adjustment device, and the device includes:
[0052] A recording module, configured to select an audio template and record a human voice audio corresponding to the audio template;
[0053] An identification module, configured to obtain fundamental frequency information corresponding to the human voice audio and identify text in the human voice audio and time stamps corresponding to the text characters;
[0054] A determination module, configured to determine fundamental frequency information corresponding to the text characters based on the fundamental frequency information and the time stamps corresponding to the text characters;
[0055] A matching module, configured to adjust the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio part corresponding to each text character in the audio template to obtain target fundamental frequency information;
[0056] An adjustment module, configured to perform pitch adjustment on the target fundamental frequency information according to the pitch information corresponding to each text character of the audio template and the pitch change trend of the audio template to determine an adjusted target human voice audio;
[0057] A fusion module, configured to obtain a template accompaniment corresponding to the audio template and perform a fusion process on the target human voice audio and the template accompaniment to obtain a target audio with adjustment completed.
[0058] In a third aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0059] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0060] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0061] The above audio adjustment method, device, computer equipment, storage medium and computer program product record a human voice audio after selecting an audio template, obtain the fundamental frequency information corresponding to the human voice audio and identify the text in the human voice audio, determine the fundamental frequency information corresponding to the text characters based on the fundamental frequency and the time stamps corresponding to the text characters, obtain the fundamental frequency duration of the text characters in the human voice audio based on the duration information of each text character corresponding to the audio part in the audio template, obtain the target fundamental frequency information, perform pitch adjustment on the target fundamental frequency information according to the pitch information and change trend of each text character, determine the adjusted target human voice audio, and then fuse the target human voice audio with the template accompaniment to obtain the adjusted target audio. Compared with the traditional adjustment method of manually editing the audio in multiple time periods, this solution processes the human voice audio based on the audio template, including duration adjustment, pitch adjustment and fusion, etc., improving the audio adjustment effect when performing ghost voice audio adjustment. Description of the Drawings
[0062] Figure 1 It is a schematic flow chart of the audio adjustment method in an embodiment;
[0063] Figure 2 It is a schematic flow chart of the human voice audio adjustment steps in an embodiment;
[0064] Figure 3 It is a schematic flow chart of the interpolation step in an embodiment;
[0065] Figure 4 It is a schematic diagram of the pitch interpolation step in an embodiment;
[0066] Figure 5 It is a schematic diagram of the timbre and vocalization feature interpolation step in an embodiment;
[0067] Figure 6 It is a schematic diagram of the pitch adjustment step in an embodiment;
[0068] Figure 7 It is a schematic flow chart of the audio adjustment method in another embodiment;
[0069] Figure 8 It is a schematic flow chart of the audio adjustment method in yet another embodiment;
[0070] Figure 9 It is a structural block diagram of the audio adjustment device in an embodiment;
[0071] Figure 10 It is an internal structure diagram of the computer equipment in an embodiment. Detailed Embodiments
[0072] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0073] In one embodiment, as Figure 1 shown, an audio adjustment method is provided. In this embodiment, it is exemplified that this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server, including the following steps:
[0074] Step S202: Select an audio template and record the human voice audio corresponding to the audio template.
[0075] Among them, the audio template can be a template for adjusting human voice audio. The audio template can be pre-constructed. The audio template includes the text corresponding to the template. The text includes multiple text characters. The above audio template also includes the template pitch information corresponding to each character in the text of the template and the duration information of the audio part corresponding to each text character in the text of the template, etc. Among them, the template pitch information can include the pitch information of each character in the text of the audio template. The pitch information of each character can be determined according to the numbered musical notation corresponding to each character in the audio template. The human voice audio can be the human voice recorded by the user through the terminal. For example, the user can click on a relevant button in the terminal to open the ghost voice recording function. At this time, the terminal can respond to the user's selection and display at least one audio template to be selected. The user can select the corresponding audio template. After the user selects the corresponding template, the terminal can receive the selection instruction for the above at least one audio template to be selected, determine the selected audio template, so that the terminal can display the text corresponding to the selected audio template, and the user can record the corresponding human voice audio according to the text. The terminal can record the human voice audio input by the user based on the text corresponding to the audio template.
[0076] Step S204: Obtain the fundamental frequency information corresponding to the human voice audio and identify the text in the human voice audio and the time stamps corresponding to the text characters.
[0077] Among them, the human voice audio can be the voice information input by the user, which is a piece of audio to be adjusted. The terminal can identify the text information contained in the human voice audio. That is, the terminal can perform speech recognition on the human voice audio to identify the text information therein. Among them, the human voice audio may contain other non-human voice noises. Therefore, the terminal can extract the fundamental frequency information from the human voice audio and identify the text in the human voice audio by identifying the fundamental frequency information.
[0078] In addition, in some embodiments, the user can first input human voice audio, and then the terminal searches for the corresponding audio template based on the human voice audio. For example, after the user inputs the human voice audio, the terminal can query the template library according to the above human voice audio. When it is detected that the text in the audio template in the template library corresponds to the above human voice audio, the audio template can be output as the query result. Among them, the text in the voice input by the user can be the same as the number of characters and the specific pronunciation of each character in the text of each audio template included in the template library, or can be a voice with the same number of characters as the text of the audio template in the audio template. Thus, the terminal can determine whether to query the audio template according to the number of characters and the specific character form or according to the number of characters based on the input situation.
[0079] In addition, the above audio template can also include a template pitch and a template beat. Among them, the template pitch can be the reference pitch of the audio template. For example, it can be pitches such as "CDEFGAB", and each template pitch can correspond to a reference pitch. The above audio template can be a kind of ghost voice template, mainly composed of beats and fundamental frequency sequences. Ghost voice means using the sounds (speaking, sound effects) in the material for editing and piecing together, tuning, and combining with the song accompaniment to obtain a complete audio-visual work, which can be popular songs, electronic music, Rap and other songs. Each audio template can be regarded as the relevant information of a preset song. Then each preset song will have a pitch, and the pitch value corresponding to the text in the audio template can be the pitch value corresponding to the pitch of the preset song. For example, the pitch value corresponding to pitch F is 65, the template beat is the beat information of the preset song, the text of the audio template is the text information of some or all of the lyrics in the preset song, and the duration information of the audio part corresponding to each text character in the text of the audio template can be the duration of the audio part corresponding to the text of the audio template in its original template audio. For example, the above audio template can correspond to a preset song, the text of the audio template is some or all of the lyrics in the preset song, and the terminal can match the lyrics of the text of the audio template with the preset song corresponding to the text of the audio template to determine the duration of the audio part corresponding to the text of the audio template, so as to obtain the duration information of the audio part corresponding to each text character in the text of the audio template.
[0080] Among them, when the terminal obtains the text in the human voice audio, it can be obtained by identifying each character in the human voice audio. For example, after the terminal obtains the above-mentioned human voice audio, it can first obtain the fundamental frequency information in the human voice audio and identify the original text corresponding to the human voice audio according to the fundamental frequency information. Among them, the original text can be the text whose correctness has not been checked. Since there may be cases of incorrect recording or missing recording when the user records the human voice audio, such as the user reversing the order of two words or forgetting to say a certain word, etc., the terminal can perform a correctness check on the original text. For example, the terminal can match the original text with the text corresponding to the audio template to obtain a matching result, and the terminal can modify the original text according to the matching result, correct or supplement the incorrect text characters and missing text characters therein, so as to obtain the text corresponding to the above-mentioned human voice audio. That is, each text character in the text corresponding to the human voice audio matches each text character in the text corresponding to the audio template. In addition, the text corresponding to the above-mentioned human voice audio may include multiple text characters, and each text character will occupy a certain amount of audio duration in the human voice audio. After the terminal identifies the text of the human voice audio, it can determine the time stamps of each text character in the human voice audio according to the duration of the audio corresponding to each text character in the text corresponding to the human voice audio. Among them, the text of the human voice audio includes at least one text character and the time stamps of each text character in at least one text character. The time stamp can be the appearance time and the end time of each text character in the human voice audio.
[0081] Specifically, the human voice audio can be a recording input by the user, which contains the lyric information sung by the user. The terminal can identify and segment the lyric information in the human voice audio. The terminal can use a speech recognition tool to identify the start timestamp and end timestamp of each character in the human voice audio input by the user, and use open-source tools such as pYin and crepe to extract the fundamental frequency information in the user's human voice audio. Specifically, it can determine the audio part of each character in the human voice audio based on the start timestamp and end timestamp of each character in the human voice audio, and then perform fundamental frequency extraction on the audio part of each character separately to obtain the fundamental frequency information corresponding to each character. Among them, when the user inputs the human voice audio, they can perform voice input according to the text of the set audio template. For example, if the text of the audio template is "Really want to go out and play", the user can input the voice containing "Really want to go out and play" as the human voice audio. The terminal can use a speech recognition tool to identify the character timestamp corresponding to each character in "Really want to go out and play" in the human voice audio, so as to determine the audio part corresponding to each character, that is, obtain the audio part of each character in the above five characters of "Really want to go out and play". After the terminal determines the audio part corresponding to each character, it can extract the fundamental frequency information of the human voice audio in the human voice audio, so as to obtain the fundamental frequency information corresponding to each character. For example, the fundamental frequency information of each character in the above five characters of "Really want to go out and play". Thus, the terminal can obtain the human voice audio according to the fundamental frequency information corresponding to each character and the character timestamp corresponding to each fundamental frequency information. After the terminal performs lyric recognition, segmentation processing, and fundamental frequency extraction on the above human voice audio, it can obtain the melody curve of the user's dry voice and the lyric timestamp position. Moreover, the terminal can also determine the target melody curve and the target lyric position by obtaining the corresponding audio template. The target melody curve can be determined according to the pitch information in the audio template, and the target lyric position can be determined according to the duration information of the audio part corresponding to each text word in the text of the audio template. The terminal can adjust the human voice audio in the user input human voice audio to be consistent with the melody in the audio template, so as to realize the audio adjustment of a kind of ghost voice audio.
[0082] Step S206, determine the fundamental frequency information corresponding to the text word based on the fundamental frequency information and the timestamp corresponding to the text word.
[0083] Among them, the fundamental frequency information can be the overall fundamental frequency information of the human voice audio. There can be corresponding text in the human voice audio, and the text includes multiple text words. Then the terminal can determine the fundamental frequency information corresponding to each text word based on the fundamental frequency information and the timestamp of each text word identified above. For example, the terminal can use the timestamp of each text word identified above as the fundamental frequency information of the corresponding time part for each text word. Thus, the terminal can adjust the fundamental frequency information of each text word to adjust the ghost voice audio.
[0084] Step S208: Adjust the fundamental frequency duration of the text words in the human voice audio based on the duration information of the audio part corresponding to each text word in the audio template to obtain the target fundamental frequency information.
[0085] Among them, the duration information may be the duration information of the audio part corresponding to each text word in the text of the above audio template. The above duration information may be determined based on the timestamps of the respective text words, and specifically may be the duration information of the audio part corresponding to the above text in the human voice audio. Since the duration of the human voice audio input by the user is not necessarily the same as the duration information in the audio template, the terminal needs to adjust the duration of the relevant feature information of the human voice audio. The terminal can obtain the duration information of the audio part corresponding to the text in the audio template, and obtain the fundamental frequency information of the audio part corresponding to each text word in the text of the audio template in the human voice audio. The terminal can adjust the fundamental frequency information of the audio part corresponding to each text word in the human voice audio according to the duration information of the audio part corresponding to each text word in the text of the above audio template to obtain the target fundamental frequency information, so that the duration of the above target fundamental frequency information matches the duration of the text of the above audio template. That is, the terminal needs to perform time-domain stretching and shrinking on the fundamental frequency information of the above human voice audio.
[0086] Among them, the fundamental frequency information of the above human voice audio may include various feature information. For example, the terminal can obtain the pitch information, timbre information, and vocalization feature information corresponding to the above human voice audio as the fundamental frequency information. And the terminal can adjust the above fundamental frequency duration based on the above various feature information to obtain the target fundamental frequency information. For example, in one embodiment, adjusting the fundamental frequency duration of the text words in the human voice audio based on the duration information of the audio part corresponding to each text word in the audio template to obtain the target fundamental frequency information includes: obtaining the pitch information, timbre information, and vocalization feature information of the audio part corresponding to each text word in the human voice audio and the corresponding text word in the audio template as the fundamental frequency information corresponding to each text word in the human voice audio; performing one-dimensional linear interpolation processing on the pitch information, timbre information, and vocalization feature information of each text word in the human voice audio respectively according to the duration information of the audio part corresponding to each text word in the audio template, so that the duration of the pitch information, the duration of the timbre information, and the duration of the vocalization feature information respectively match the duration information of the audio part corresponding to each text word in the audio template; using the interpolated target pitch information, target timbre information, and target vocalization feature information as the target fundamental frequency information. In this embodiment, the human voice audio may include fundamental frequency information and the character timestamps corresponding to each character in the fundamental frequency information. As Figure 2 shown Figure 2It is a schematic flowchart of the human voice audio adjustment steps in an embodiment. The terminal can obtain the pitch information, timbre information, and vocalization feature information of the audio part corresponding to each text word in the text of the audio template from the dry voice segment, that is, the above human voice audio, as the fundamental frequency information. Among them, the pitch information can be the intonation accuracy information of the user's pronunciation, the timbre information can be the timbre of the user's vocalization, and the vocalization feature information can be the information that cannot be displayed in the fundamental frequency but the user actually vocalizes. For example, when the user makes a soft sound, the fundamental frequency of this soft sound will not be displayed in the fundamental frequency information, but it is necessary to represent these soft sounds made by the user. Therefore, the terminal needs to obtain the vocalization features. The above user will carry vocalization features when making each sound. Therefore, the above vocalization feature information can exist throughout the human voice audio.
[0087] The above pitch information can be obtained based on the fundamental frequency of the human voice audio. That is, the terminal can determine the pitch information by performing fundamental frequency extraction on the human voice audio; for the timbre information, the terminal can obtain it according to the envelope matrix in the human voice audio; for the vocalization feature information, the terminal can obtain it through the aperiodic information. After the terminal obtains the above pitch information, timbre information, and vocalization feature information, it can adjust the corresponding pitch information, timbre information, and vocalization feature information of the human voice audio according to the duration information of the audio part corresponding to each text word in the text of the audio template, and obtain the target pitch information, target timbre information, and target vocalization feature information that match the duration information of the text of the above audio template. The terminal can use the above target pitch information, target timbre information, and target vocalization feature information as the target fundamental frequency information. That is, the above pitch information, timbre information, and vocalization feature information can all be in sequence form. The terminal can scale the pitch information, timbre information, and vocalization feature information in the time domain, so as to obtain the target pitch information, target timbre information, and target vocalization feature information that match the duration information of the text of the audio template.
[0088] Specifically, the terminal can match the duration information of the pitch information, timbre information, and vocalization feature information corresponding to the above human voice audio with the duration information of the audio template. The terminal can perform one-dimensional linear interpolation processing on the above pitch information, timbre information, and vocalization feature information respectively according to the duration information of the audio part corresponding to each text word in the text of the above audio template. So that the duration of the above pitch information, the duration of the timbre information, and the duration of the vocalization feature information respectively match the duration information of the audio part corresponding to each text word in the text of the audio template. That is, the terminal needs to scale the above pitch information, timbre information, and vocalization feature information in the time domain respectively. Thus, the terminal can use the target pitch information, target timbre information, and target vocalization feature information after interpolation processing as the above target fundamental frequency information. Specifically, as Figure 2As shown, the terminal can perform vector interpolation on the pitch information corresponding to the fundamental frequency information based on the duration information of the audio part corresponding to each text character in the text of the audio template, and use the interpolated pitch information sequence as the target pitch information. The interpolated pitch sequence can be expressed as: c*(m) = interp1(c(n)), where c*(m) represents the pitch information sequence with a length of M after interpolation processing, m = 0, 1, 2, …, M; interp1() represents one-dimensional linear interpolation processing, and the interpolation curve can use spline, cubic, linear, etc.
[0089] For another example Figure 2 As shown, for the above timbre information and vocalization feature information, the terminal needs to perform matrix interpolation on the timbre information and vocalization feature information. However, since only time scaling of the timbre information and vocalization feature information is required, the terminal can obtain the matrix interpolation result by using linear interpolation of each dimension signal. Specifically, the terminal can represent the interpolation result of the timbre information through the following formula: e*(m, k0) = interp1(e(n, k0)). Where e*(m, k0) represents the interpolation result of the envelope signal at the k0th frequency point of the mth output frame. Specifically, as Figure 3 shown Figure 3 FIG. is a schematic flowchart of the interpolation step in an embodiment. Through the above interpolation processing, the duration of the timbre information can be changed from the duration N to the duration M, thereby realizing time scaling and making the timbre information match the duration in the audio template. For the vocalization feature interpolation process, its specific formula can be as follows: a*(n, i0) = interp1(a(n, i0)). Where a*(n, i0) represents the interpolated vocalization feature information sequence at the i0th frequency band in the nth output frame. The terminal performs linear interpolation processing on each fundamental frequency information to obtain the target fundamental frequency information that matches the duration of the audio template after interpolation processing. Thus, the terminal can perform audio adjustment based on the target fundamental frequency information, improving the adjustment effect of the audio adjustment for the ghost voice audio.
[0090] Specifically, as Figure 4 shown Figure 4 FIG. is a schematic diagram of the pitch interpolation step in an embodiment. Taking the text of the audio template as "I really want to go out and play" as an example, Figure 4 it can be a comparison chart of the pitch curves before and after interpolation of the word "good". Wherein, the abscissa is the time information, the ordinate is the pitch value, and the curve 401 is the original pitch of the fundamental frequency information corresponding to the word "good" in the above human voice audio before interpolation, and its duration does not match the duration of the corresponding character in the text of the audio template. However, through the above interpolation processing, a pitch curve 403 that matches the duration of the corresponding character in the text of the above audio template can be obtained.
[0091] In addition, the terminal can also perform interpolation processing on the timbre information and vocalization feature information in the above fundamental frequency information in the time domain. For example, Figure 5 as shown Figure 5 in the schematic diagram of the timbre and vocalization feature interpolation steps in an embodiment. Figure 5 It can be a comparison chart of the timbre and vocalization features before and after interpolation of the word "good". Among them, Figure 5 in each coordinate axis, the abscissa is time, the ordinate of the timbre feature is frequency information, and the ordinate of the vocalization feature is frequency band information. 601 is the original timbre feature corresponding to the word "good" in the above fundamental frequency information. It can be known that its duration information does not match the duration of the corresponding character in the text of the audio template. After interpolation processing, 603 can be obtained, and only scaling processing is performed in the time domain to make it match the duration of the corresponding character in the text of the audio template. 605 can be the original vocalization feature corresponding to the word "good" in the above fundamental frequency information. It can be known that its duration information does not match the duration of the corresponding character in the text of the audio template. After interpolation processing, 607 can be obtained, and only scaling processing is performed in the time domain to make it match the duration of the corresponding character in the text of the audio template.
[0092] Step S210: According to the pitch information corresponding to each text word of the audio template and the pitch change trend of the audio template, perform pitch adjustment on the target fundamental frequency information to determine the adjusted target human voice audio.
[0093] Among them, the pitch information in the audio template can be the pitch information of the text in the audio template corresponding to the above-mentioned human voice audio. Among them, the text in the audio template can include at least one character, each character can be a text word, and the text of the audio template can be in the form of a sequence containing at least one character. Then, the pitch information in the above-mentioned audio template can be a pitch sequence, and the pitch sequence contains the pitch information corresponding to each character in the text of the audio template. The pitch information of each character can be determined according to the pitch information of the audio part corresponding to each character in the preset song corresponding to the audio template. The pitch information corresponding to each character in the pitch information of the above-mentioned audio template can be used to determine the pitch accuracy of the audio part corresponding to each text word in the human voice audio corresponding to the user. That is, each pitch information in the pitch information of the above-mentioned audio template can be regarded as a target value. Among them, the pitch information of the audio template is also called template pitch information. The terminal can perform pitch adjustment on the target feature information according to the pitch information corresponding to each character in the above-mentioned template pitch information and the pitch change trend of each pitch information. Among them, the pitch change trend can be the rising and falling situations between the pitch information in the above-mentioned audio template. And the terminal can also determine the adjusted target human voice audio based on the target fundamental frequency information after pitch adjustment. For example, the terminal can combine the fundamental frequency information items in the target fundamental frequency information after pitch adjustment, such as combining pitch, timbre, and vocalization characteristics, etc., to obtain the adjusted target human voice audio. Among them, the target human voice audio can be the audio information containing only the user's human voice obtained after performing pitch accuracy adjustment on the pronunciation of each character in the user's audio based on the audio template.
[0094] Specifically, the above pitch adjustment can be performed on the fundamental frequency information corresponding to each character of the text of the audio template. The above audio template can be a kind of ghost voice template. That is, the terminal can perform variable speed and pitch change on a character-by-character basis for the fundamental frequency information according to the target melody curve formed by the template pitch information in the audio template, and obtain the above target human voice audio by splicing the lyrics after variable speed and pitch change. Among them, the terminal can perform lyric segmentation on the above human voice audio to obtain the human voice audio corresponding to each character, and perform pitch adjustment including variable speed and pitch change on a character-by-character basis for the human voice audio based on the pitch information of the audio part corresponding to each text character in the preset song corresponding to the audio template and the text of the audio template. The terminal can implement the above pitch adjustment through a vocoder. The above variable speed and pitch change processing can be implemented by signal processing means such as TSM (Time scale modification) and world vocoder, or can be implemented by neural network methods such as WaveNet and LPCNet. Among them, the WaveNet model is a sequence generation model that can be used for speech generation modeling. LPCNet is a digital signal processing and neural network cleverly combined and applied to the vocoder in speech synthesis, and can synthesize high-quality speech in real time on an ordinary CPU.
[0095] Step S212, obtain the template accompaniment corresponding to the audio template, and perform fusion processing on the target human voice audio and the template accompaniment to obtain the adjusted target audio.
[0096] Among them, there is a corresponding preset song for the above audio template, and the preset song will carry corresponding accompaniment information, and the accompaniment information of the preset song can be used as the template accompaniment of the above audio template. Among them, the above accompaniment information can be an accompaniment mainly composed of percussion. For example, the accompaniment information can include but is not limited to relatively low-pitched sounds, such as bass drums and knocking on the table, which can be used as downbeat information; it can also include but is not limited to relatively crisp sounds, such as snare drums, clapping hands, applauding, and finger-snapping, which can be used as upbeat information. In addition, in some embodiments, the terminal can also guide the user to record instrument sounds as template accompaniment information, and the accompaniment recorded by the user can be mainly composed of percussion.
[0097] After the terminal obtains the template accompaniment corresponding to the audio template, it can fuse the target human voice audio with the template accompaniment to obtain the adjusted target audio. Among them, the above target human voice audio can be the human voice audio of the user obtained after processing such as speed change and pitch change. When the terminal fuses the target human voice audio with the template accompaniment, it can include operations such as beat matching and mixing processing. Thus, the terminal can perform beat matching and mixing processing on the target human voice audio with the template beat to obtain the adjusted target audio. Specifically, since the above target human voice audio can be the human voice audio corresponding to each character, the terminal also needs to self- splice the target human voice audio of each character to obtain the target human voice audio containing the complete text, so that the terminal can perform the above fusion processing based on the target human voice audio of the complete text.
[0098] In the above audio adjustment method, by selecting an audio template and then recording the human voice audio, obtaining the fundamental frequency information corresponding to the human voice audio and identifying the text in the human voice audio, determining the fundamental frequency information corresponding to the text characters based on the fundamental frequency and the time stamps corresponding to the text characters, obtaining the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio part corresponding to each text character in the audio template to obtain the target fundamental frequency information, adjusting the pitch of the target fundamental frequency information according to the pitch information and change trend of each text character, determining the adjusted target human voice audio, and then fusing the target human voice audio with the template accompaniment to obtain the adjusted target audio. Compared with the traditional adjustment method of manually editing the audio in multiple time periods, this solution improves the audio adjustment effect when performing audio adjustment for ghost voice audio by performing processing such as duration adjustment, pitch adjustment, and fusion on the human voice audio based on the audio template.
[0099] In one embodiment, obtaining the pitch information, timbre information, and vocalization feature information corresponding to the human voice audio includes: obtaining the text characters corresponding to each character of the audio template in the text included in the human voice audio; obtaining the partial human voice audio corresponding to each text character in the human voice audio; for the partial human voice audio corresponding to each text character, obtaining the partial fundamental frequency corresponding to the partial human voice audio; according to the partial fundamental frequency corresponding to the partial human voice audio and the reference frequency, determining the pitch offset value of the pitch information corresponding to the partial human voice audio, and determining the pitch information corresponding to the partial human voice audio according to the reference pitch and the pitch offset value; constructing an envelope matrix according to the envelope vectors of a preset number of frequency points in the partial human voice audio to obtain the timbre information; and obtaining the vocalization feature information according to the aperiodic information in a preset number of frequency bands in the partial human voice audio.
[0100] In this embodiment, the terminal can obtain the pitch information, timbre information, and vocalization feature information in the above-mentioned human voice audio as the fundamental frequency information. For different types of feature information in the fundamental frequency information, the terminal can obtain them from the human voice audio in different ways. Among them, the human voice audio can be a type of fundamental frequency information. The terminal can obtain the above-mentioned various fundamental frequency information for each part of the human voice audio corresponding to each character in the human voice audio. Among them, the above-mentioned part of the human voice audio can be the human voice audio corresponding to a text character. The above audio template can include a reference pitch and a reference frequency corresponding to the reference pitch; the terminal can first obtain the reference pitch in the audio template and the reference frequency corresponding to the reference pitch. Among them, the reference pitch can be the pitch corresponding to the tone of the audio template. For example, the reference pitch corresponding to tone F is 65, and the reference pitch corresponding to tone A is 69. The above reference pitch can be pre-stored in the audio template, or the terminal can obtain the reference pitch of the above audio template by querying a preset tone-pitch mapping table. The terminal can also obtain the reference frequency corresponding to the above reference pitch. Among them, the reference frequency can be the frequency of the reference pitch of the above tone, and the terminal can obtain the reference frequency corresponding to the above reference pitch through a preset pitch-frequency mapping table, where the preset pitch-frequency mapping table can include multiple corresponding relationships between pitch and frequency. For example, the reference pitch corresponding to tone A is 69, and the reference frequency corresponding to the reference pitch 69 is 440Hz.
[0101] The above-mentioned part of the human voice audio can be the human voice audio corresponding to each text character in the human voice audio. Among them, the text characters of the human voice audio are the texts corresponding to each character of the text included in the human voice audio and the text of the audio template. For the part of the human voice audio corresponding to each text character in the text of the audio template in the human voice audio, the terminal can obtain the fundamental frequency corresponding to this part of the human voice audio, which is called the partial fundamental frequency. Among them, the partial fundamental frequency can be a type of frequency information. To fit the human ear's hearing, the terminal can convert the frequency information into pitch information. For example, the terminal can determine the pitch offset value of the pitch information corresponding to this part of the human voice audio according to the partial fundamental frequency corresponding to this part of the human voice audio and the above reference frequency, and determine the pitch information corresponding to the human voice audio according to the reference pitch and the above pitch offset value. Specifically, the terminal can perform fundamental frequency extraction on the human voice audio to obtain the dry voice fundamental frequency information f(n), where n represents the frame index, and each frame corresponds to a preset duration, such as 5ms. The terminal can set the number of signal frames of the current human voice audio as N, then the sequence length of the above human voice audio is N * 5ms, and the value range of the above n is 0, 1, 2,..., N - 1. The above pitch information can be represented by c(n), and the terminal can determine the pitch information corresponding to the fundamental frequency information of the above human voice audio based on the following formula:
[0102] Among them, taking the template pitch as key A as an example, the reference pitch corresponding to key A is 69, and the corresponding reference frequency is 440 Hz. In the above formula, represents the above pitch offset value, where 12 represents the number of semitones in an octave space. The above pitch information can be in units of semitones. One semitone can correspond to 100 cents, and one octave is 1200 cents.
[0103] The terminal can also construct a matrix based on the envelope vectors of a preset number of frequency points in the partial human voice audio to obtain the above timbre information. That is, the terminal can obtain the user's timbre information through the envelope matrix. Among them, the above partial human voice audio can include multiple frequency points, each frequency point can represent a frequency, and the envelope matrix can be constructed from the envelope vectors in the spectrogram of the user's human voice audio. Specifically, the terminal can use e(n,k) to represent the nth frame signal of the above human voice audio and the envelope vector at the kth frequency point. Where k = 0, 1, 2, …, K. Where K represents the spectral width, and common values are: 1024, 2048, 4096, etc.
[0104] The terminal can obtain the above vocalization feature information based on the aperiodic information in a preset number of frequency bands in the partial human voice audio. Among them, there can be a preset number of frequency bands in the above partial human voice audio. A frequency band can be an interval composed of frequencies within a preset range. There are multiple frequency points in each frequency band. The vocalization feature information can be information that cannot be displayed in the fundamental frequency but the user actually vocalizes. For example, when the user makes a soft sound, the fundamental frequency of this soft sound will not be displayed in the fundamental frequency information, but it is necessary to represent these soft sounds made by the user. Therefore, the terminal needs to obtain the vocalization features. The above user will carry vocalization features when making each sound. Therefore, the above vocalization feature information can exist throughout the human voice audio. For the human voice with fundamental frequency information, its signal is usually periodic information, while for the vocalization feature information without fundamental frequency information, its signal is in an aperiodic form. The terminal can describe the above vocalization features through the aperiodic components. Specifically, the terminal can use a(n,i) to represent the aperiodic information at the i-th frequency band of the n-th frame signal, where i = 0, 1, …, I - 1 represents the i-th frequency band, and I represents the number of frequency bands. Common values are 4, 5, etc. That is, the terminal can divide the above human voice audio into I frequency bands.
[0105] After the terminal obtains the above pitch information, timbre information, and vocalization feature information, it can obtain the fundamental frequency information corresponding to the human voice audio. Specifically, the fundamental frequency information can be represented in the following form: P ana = [f(n)e(n,k)a(n,i)]. The terminal can adjust the audio of the above fundamental frequency information through a vocoder and can reconstruct the adjusted audio based on the adjusted human voice audio to achieve audio adjustment of the user's human voice.
[0106] Through this embodiment, the terminal can obtain multiple fundamental frequency information for audio adjustment based on different characteristic information in the human voice audio, so as to perform audio adjustment based on the multiple fundamental frequency information, which can improve the adjustment effect of audio adjustment.
[0107] In one embodiment, according to the pitch information corresponding to each text character of the audio template and the pitch change trend of the audio template, performing pitch adjustment on the target fundamental frequency information includes: obtaining the target pitch information in the target fundamental frequency information, and obtaining the partial target pitch information corresponding to each text character in the human voice audio in the target pitch information; for each partial target pitch information, obtaining a preset number of adjacent frames adjacent to each frame of the partial target pitch information; according to the average value of the partial target pitch information of each frame and the adjacent partial target pitch information of the adjacent frames corresponding to each frame, obtaining the average pitch information corresponding to the partial target pitch information of each frame; the average pitch information represents the pitch change trend of the partial target pitch information of each frame; according to the average pitch information and the pitch information corresponding to each text character in the audio template, performing pitch adjustment on the partial target pitch information in the target audio feature.
[0108] In this embodiment, after the terminal performs interpolation processing on the fundamental frequency information to obtain the target fundamental frequency information that matches the duration of the text in the above audio template, pitch adjustment can be performed on the target fundamental frequency information. Among them, the above target fundamental frequency information includes target pitch information, target timbre information, and target vocalization feature information, and the above pitch adjustment can be to adjust the target pitch information in the target fundamental frequency information. The terminal can perform pitch adjustment on the target pitch information in the target fundamental frequency information according to the pitch information in the above audio template. Among them, the pitch information of the audio template can be a kind of sequence information, including the pitch information corresponding to each text character in the text of the audio template. The pitch information of the audio template can be used as the target value of the above target pitch information. The terminal can obtain the pitch information corresponding to the character represented by the above target fundamental frequency information in the audio template, and perform pitch adjustment on the target pitch information of the character based on the pitch information in the audio template and the pitch change trend of the audio template.
[0109] Specifically, such as Figure 2As shown above, when performing pitch adjustment on the target pitch information in the target fundamental frequency information, it includes dry pitch trend estimation and calculation of the target pitch. The target pitch can be determined based on the pitch in the above audio template. The dry pitch trend estimation represents the pitch change trend of the target pitch information within its time span. The terminal can obtain a smoother pitch information sequence through the dry pitch trend. The above human voice audio can include human voice audio corresponding to multiple characters. The terminal can perform audio adjustment in units of characters. Then, each of the above characters can correspond to corresponding target fundamental frequency information. For each character in the text of the audio template, the terminal can obtain the target pitch information in the target fundamental frequency information and obtain the partial target pitch information corresponding to the text characters in each of the above human voice audios in the target pitch information. Among them, the text characters in the human voice audio are the texts in the human voice audio corresponding to each character in the text of the audio template. For each partial target pitch information, the terminal can obtain a preset number of adjacent frames adjacent to each frame of the partial target pitch information. The terminal can obtain the target pitch information of these preset number of adjacent frames as adjacent partial target pitch information. The terminal can obtain adjacent partial target pitch information for each of the above frames. The terminal can obtain the average pitch information corresponding to the partial target pitch information of each frame according to the average value of the partial target pitch information of each frame and the adjacent partial target pitch information corresponding to each frame. Among them, the average pitch information represents the pitch change trend of the partial target pitch information of each frame. Based on the average pitch information of the above multiple frames, the terminal can obtain a smoother pitch change curve corresponding to the character. After the terminal obtains the above average pitch information, it can perform pitch adjustment on the pitch information part corresponding to the character in the target audio feature based on the above average pitch information and the pitch information of the audio template.
[0110] Specifically, the terminal can first obtain smoothed pitch information for the target pitch information in the target fundamental frequency information. The terminal can use mean smoothing to estimate the average pitch information, and its calculation formula is as follows:
[0111] Among them, represents the smoothed pitch information sequence corresponding to each character; c*(n + l) represents the adjacent target pitch information of the nth frame in the target pitch information of each character; L = 0.3 * N, where 0.3 is an empirical value and can be set according to the actual situation, and N is the number of frames of the above target pitch information. The terminal can perform pitch adjustment on the pitch information part corresponding to the character in the target audio feature according to the above average pitch information and the pitch information of the audio template, and its adjustment formula is as follows: Among them, c*(n) represents the sequence of pitch information corresponding to each character in the above human voice audio, and c ref (n) represents the template pitch information sequence of the above audio template. represents the target pitch information sequence after the pitch adjustment. That is, the terminal can determine the offset of the pitch information corresponding to each character based on the above formula, and then determine the specific pitch of the adjusted pitch information based on the template pitch information and the offset corresponding to each character, so as to achieve the pitch adjustment of the pitch information of each character.
[0112] Specifically, Figure 6 As shown, Figure 6 FIG. 4 is a schematic diagram of a pitch adjustment step in an embodiment. Figure 6 The horizontal axis is the time information, and the vertical axis is the pitch value. Take the text of the audio template "I really want to go out and play" as an example. Figure 6 The pitch sequence of the fundamental frequency information before adjustment may be the pitch adjustment information corresponding to the word "good" in the fundamental frequency information. Figure 6 Curve 701 is the pitch change corresponding to the pronunciation of the word "good" in the fundamental frequency information before adjustment. The terminal can obtain curve 703 by estimating the pitch trend of each pitch of the word "good". Curve 705 is the pitch information of the word "good" in the above audio template. The terminal can adjust the pitch of the fundamental frequency information of the human voice audio according to the above method, for example, adjust the pitch of the fundamental frequency corresponding to the word "good", so as to obtain curve 707, that is, the pitch curve after pitch adjustment, so that the pitch curve after pitch adjustment is as close to the pitch in the audio template as possible without destroying the original sound characteristics of the human voice audio.
[0113] After the terminal adjusts the target pitch information of each of the above characters, it can obtain the target pitch information after pitch adjustment, so that the terminal can obtain the target fundamental frequency information after pitch adjustment based on the target pitch information after pitch adjustment, the above target timbre information and the target vocal feature information.
[0114] Through the above embodiments, the terminal can perform pitch adjustment including pitch trend estimation and pitch adjustment on the target pitch information in the target fundamental frequency information, so as to obtain target pitch information that conforms to the pitch of each character in the audio template, thereby improving the adjustment effect when performing ghost audio adjustment.
[0115] In one embodiment, determining the adjusted target vocal audio includes: determining the frequency adjustment value of the target pitch information after the pitch adjustment according to the target pitch information after the pitch adjustment and the reference pitch of the audio template; determining the adjusted target frequency corresponding to the vocal audio according to the reference frequency of the audio template and the frequency adjustment value; determining the adjusted target vocal audio according to the target frequency, target timbre information and target vocal feature information.
[0116] In this embodiment, in order to fit the human ear's hearing, the terminal can convert the frequency information of the above fundamental frequency into pitch information for adjustment. After the terminal finishes adjusting the pitch, the terminal can restore the target pitch information after pitch adjustment to frequency information, so that the terminal can determine the specific pronunciation effect in the human voice audio based on the frequency information. For example, the terminal can determine the frequency adjustment value of the target pitch information after pitch adjustment according to the target pitch information after the above pitch adjustment and the reference pitch of the above audio template. This frequency adjustment value can determine the offset of the frequency information corresponding to each target pitch information from the reference frequency of this pitch. The terminal can determine the adjusted target frequency corresponding to the human voice audio according to the reference frequency of the audio template and the above frequency adjustment value. That is, the terminal can perform the above determination of the target frequency for the human voice audio corresponding to each character, so that the terminal can splice the target frequencies of each character to obtain the target frequency sequence of the overall human voice audio. The terminal can determine the adjusted target human voice audio according to the above target frequency, target timbre information, and target vocalization feature information. Specifically, taking the reference pitch as A, the reference pitch as 69, and the reference frequency as 440 Hz as an example, the acquisition formula of the above target frequency can be shown as follows:
[0117] Among them, represents the sequence of target pitch information after pitch adjustment, f*(n) represents the target frequency, and 12 represents the number of semitones in an octave space. After the terminal obtains the target frequency through the above frequency calculation formula, together with the target frequency, target timbre information, and target vocalization feature information, it can determine the adjusted target human voice audio. Specifically, the target parameter sequence constituting the target human voice audio can be P ana =[f*(n)e*(m,k)a*(m,i)]. Thus, the terminal can obtain the target melody corresponding to each character in the audio adjustment through the synthesis processing of the above parameter sequence.
[0118] Through this embodiment, the terminal can obtain the target human voice audio based on the frequency conversion of the pitch information, combined with the converted target frequency, and the above interpolated target timbre information and target vocalization feature information. The adjustment effect during audio adjustment of the ghost audio is improved.
[0119] In one embodiment, fusing the target human voice audio with the template accompaniment to obtain the adjusted target audio includes: obtaining the template beat of the template accompaniment; matching the target human voice audio with the template accompaniment according to the template beat; and performing mixing processing on the matched audio to obtain the adjusted target audio.
[0120] In this embodiment, after the terminal obtains the adjusted target human voice audio, it can fuse the target human voice audio with the template accompaniment. This fusion process includes processing such as beat matching and mixing. Among them, the above template accompaniment can be a kind of instrumental accompaniment with a certain beat. The terminal can obtain the template beat of the template accompaniment and match the target human voice audio with the template accompaniment according to the template beat to obtain the target human voice audio and the template accompaniment after the matching is completed. The terminal can also perform a mixing process on the audio after the matching is completed to obtain the adjusted target audio.
[0121] Among them, the above target human voice audio can be the information obtained by splicing the adjusted target human voice audio corresponding to each character based on the character timestamp. When performing beat matching, the terminal can perform beat matching through the downbeat timestamp. For example, in one embodiment, obtaining the template beat of the template accompaniment includes: obtaining the audio energy values at each time point in the template accompaniment, and taking the time points corresponding to the audio energy values greater than the preset energy threshold as the downbeat timestamps of the template accompaniment; determining the template beat of the template accompaniment according to multiple downbeat timestamps. In this embodiment, the terminal can obtain the audio energy values at each time point in the above template accompaniment and take the time points corresponding to the audio energy values greater than the preset energy threshold as the downbeat timestamps of the template accompaniment, that is, this time point is the downbeat of the template accompaniment. The terminal can obtain the template beat of the template accompaniment according to multiple downbeat timestamps. Specifically, the above template accompaniment can be the single accompaniment percussion sound created by the user. The terminal estimates the downbeat timestamp using the position of the maximum energy value, and the terminal can also estimate the timestamp position of each beat using information such as the time signature and bar line in the template, and align the above target human voice audio to the corresponding time position of the above downbeat. The terminal can also fill the beat positions without human voice with the percussion sound of the user.
[0122] After the terminal determines the template beat and performs beat matching, it can also perform mixing processing. For example, in one embodiment, the mixed processing is performed on the matched audio to obtain the adjusted target audio, including: adjusting the audio energy of the target human voice audio in the matched audio according to the audio energy of the template accompaniment in the matched audio, so that the audio energy of the target human voice audio is less than the audio energy of the template accompaniment; superimposing and mixing the target human voice audio with adjusted audio energy and the template accompaniment to obtain the adjusted target audio. In this embodiment, the terminal can adjust the audio energy of the target human voice audio in the matched audio according to the audio energy of the template accompaniment in the above-mentioned matched audio, so that the audio energy of the target human voice audio is less than the audio energy of the template accompaniment. The terminal can also superimpose and mix the target human voice audio with adjusted audio energy and the above-mentioned template accompaniment to obtain the adjusted target audio. Specifically, the terminal can perform weighted superposition on the energy of the human voice audio after adjustment and matching and the template accompaniment corresponding to the above-mentioned audio template. To prevent the dry voice energy from being too large and covering the accompaniment, the terminal can make the loudness of the human voice audio lower than the loudness of the accompaniment by a preset value, such as 3dB. After the terminal adjusts the energy of the above-mentioned human voice audio, it can superimpose the adjusted human voice audio and the template accompaniment to obtain the finally adjusted target audio for output.
[0123] Through the above embodiments, the terminal can perform beat matching on the target human voice audio after pitch adjustment and the template accompaniment, and perform mixing processing on the adjusted target human voice audio and the template accompaniment based on the energy value, thereby improving the audio adjustment effect of the ghost voice audio.
[0124] In one embodiment, as Figure 7 shown, Figure 7 is a schematic flowchart of an audio adjustment method in another embodiment. In this embodiment, the above audio adjustment can be an adjustment of ghost voice audio. The method includes the following steps: The terminal can construct a ghost template, that is, the above audio template, through the above method. The terminal can obtain the user's dry voice material, such as the human voice audio input by the user. Specifically, the terminal can display the above audio templates, and the user can select an audio template from the multiple audio templates displayed by the terminal and record the human voice audio according to the text corresponding to the audio template, so that the terminal can record the human voice audio input by the user as the dry voice material. The terminal can perform lyric segmentation on the dry voice material, and perform audio adjustment for each lyric based on the above ghost template to obtain the dry voice melody information, that is, the above adjusted target human voice audio. The terminal can also obtain the accompaniment material and obtain the beat information therein, match the dry voice melody information with the beat through the ghost template, and the terminal can also mix the matched accompaniment and the user's dry voice melody to obtain a ghost voice audio, that is, the above adjusted target audio, and output it.
[0125] Through this embodiment, the terminal processes the human voice audio based on the audio template, including duration adjustment, pitch adjustment, and fusion, etc., improving the audio adjustment effect when performing audio adjustment for ghostly audio.
[0126] In addition, in some embodiments, as Figure 8 shown, Figure 8 FIG. 8 is a schematic flowchart of an audio adjustment method in yet another embodiment. In this embodiment, the terminal can obtain a specific song or a song segment specified by the user, or any random melody hummed by the user as the target melody. The terminal estimates the target melody through an automatic fundamental frequency extraction and beat detection tool, uses the target melody as the audio template, and performs the above audio adjustment process based on this audio template.
[0127] Through the above embodiments, the terminal can obtain an audio template based on the melody information determined by the user independently, improving the flexibility of audio adjustment. And the terminal can process the human voice audio based on the audio template, including duration adjustment, pitch adjustment, and fusion, etc., improving the audio adjustment effect when performing audio adjustment for ghostly audio.
[0128] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0129] Based on the same inventive concept, the embodiments of the present application also provide an audio adjustment device for implementing the above-mentioned audio adjustment method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following audio adjustment device can refer to the limitations on the audio adjustment method in the above text, and will not be repeated here.
[0130] In one embodiment, as Figure 9 shown, an audio adjustment device is provided, including: a recording module 500, an identification module 502, a determination module 504, a matching module 506, an adjustment module 508, and a fusion module 510, where:
[0131] A recording module 500, configured to select an audio template and record the human voice audio corresponding to the audio template;
[0132] An identification module 502, configured to obtain the fundamental frequency information corresponding to the human voice audio and identify the text in the human voice audio and the time stamps corresponding to the text characters;
[0133] A determination module 504, configured to determine the fundamental frequency information corresponding to the text characters based on the fundamental frequency information and the time stamps corresponding to the text characters;
[0134] A matching module 506, configured to adjust the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio part corresponding to each text character in the audio template to obtain the target fundamental frequency information.
[0135] An adjustment module 508, configured to perform pitch adjustment on the target fundamental frequency information according to the pitch information corresponding to each text character of the audio template and the pitch change trend of the audio template, and determine the adjusted target human voice audio.
[0136] A fusion module 510, configured to obtain the template accompaniment corresponding to the audio template, and perform a fusion process on the target human voice audio and the template accompaniment to obtain the adjusted target audio.
[0137] In one embodiment, the above-mentioned recording module 500 is specifically configured to display at least one audio template to be selected; receive a selection instruction for the at least one audio template to be selected, determine the selected audio template; display the text corresponding to the audio template, and record the human voice audio input by the user based on the text corresponding to the audio template.
[0138] In one embodiment, the above-mentioned identification module 502 is specifically configured to identify the original text corresponding to the human voice audio according to the fundamental frequency information; modify the original text according to the matching result between the original text and the text corresponding to the audio template to obtain the text corresponding to the human voice audio; each text character in the text corresponding to the human voice audio matches each text character in the text corresponding to the audio template; according to the duration of the audio corresponding to each text character in the text corresponding to the human voice audio in the human voice audio, determine the time stamps of each text character in the human voice audio.
[0139] In one embodiment, the above-mentioned identification module 502 is specifically configured to obtain the pitch information, timbre information, and vocalization feature information of the audio part corresponding to each text character in the text of the audio template in the human voice audio as the fundamental frequency information.
[0140] In one embodiment, the above-mentioned matching module 506 is specifically configured to obtain the pitch information, timbre information, and vocalization feature information of the audio parts corresponding to the respective text characters in the human voice audio and the corresponding text characters in the audio template as the fundamental frequency information corresponding to the respective text characters in the human voice audio; perform one-dimensional linear interpolation processing on the pitch information, timbre information, and vocalization feature information of the respective text characters in the human voice audio according to the duration information of the audio parts corresponding to each text character in the audio template, so that the duration of the pitch information, the duration of the timbre information, and the duration of the vocalization feature information respectively match the duration information of the audio parts corresponding to each text character in the audio template; and use the interpolated target pitch information, target timbre information, and target vocalization feature information as the target fundamental frequency information.
[0141] In one embodiment, the above-mentioned recognition module 502 is specifically configured to obtain the text characters corresponding to each character of the audio template in the text of the human voice audio; obtain the partial human voice audio corresponding to each text character in the human voice audio; for the partial human voice audio corresponding to each text character, obtain the partial fundamental frequency corresponding to the partial human voice audio; determine the pitch offset value of the pitch information corresponding to the partial human voice audio according to the partial fundamental frequency corresponding to the partial human voice audio and the reference frequency, and determine the pitch information corresponding to the partial human voice audio according to the reference pitch and the pitch offset value; construct an envelope matrix based on the envelope vectors of a preset number of frequency points in the partial human voice audio to obtain the timbre information; and obtain the vocalization feature information according to the aperiodic information in a preset number of frequency bands in the partial human voice audio.
[0142] In one embodiment, the above-mentioned adjustment module 508 is specifically configured to obtain the target pitch information in the target fundamental frequency information, and obtain the partial target pitch information corresponding to each text character in the human voice audio in the target pitch information; for each partial target pitch information, obtain a preset number of adjacent frames adjacent to each frame of the partial target pitch information; obtain the average pitch information corresponding to the partial target pitch information of each frame according to the average value of the partial target pitch information of each frame and the adjacent partial target pitch information of the adjacent frames corresponding to each frame; the average pitch information represents the pitch change trend of the partial target pitch information of each frame; and perform pitch adjustment on the partial target pitch information in the target audio feature according to the average pitch information and the pitch information corresponding to each text character in the audio template.
[0143] In one embodiment, the above-mentioned adjustment module 508 is specifically configured to determine the frequency adjustment value of the target pitch information after pitch adjustment according to the target pitch information after pitch adjustment and the reference pitch of the audio template; determine the adjusted target frequency corresponding to the human voice audio according to the reference frequency of the audio template and the frequency adjustment value; and determine the adjusted target human voice audio according to the target frequency, target timbre information, and target vocalization feature information.
[0144] In one embodiment, the above-mentioned fusion module 510 is specifically configured to obtain the template beat of the template accompaniment; match the target human voice audio with the template accompaniment according to the template beat; and perform mixing processing on the matched audio to obtain the adjusted target audio.
[0145] In one embodiment, the above-mentioned fusion module 510 is specifically configured to obtain the audio energy values at each time point in the template accompaniment, and use the time points corresponding to the audio energy values greater than the preset energy threshold as the downbeat timestamps of the template accompaniment; determine the template beat of the template accompaniment according to multiple downbeat timestamps.
[0146] In one embodiment, the above-mentioned fusion module 510 is specifically configured to adjust the audio energy of the target human voice audio in the matched audio according to the audio energy of the template accompaniment in the matched audio, so that the audio energy of the target human voice audio is less than the audio energy of the template accompaniment; superimpose and mix the target human voice audio with adjusted audio energy and the template accompaniment to obtain the adjusted target audio.
[0147] In one embodiment, the above-mentioned acquisition module 500 is specifically configured to perform speech recognition on the human voice audio to determine the audio text corresponding to the human voice audio; the audio text includes at least one character and the character timestamps of each character; extract the fundamental frequency of the audio part corresponding to each character in the audio text to obtain the fundamental frequency information corresponding to each character; and obtain the human voice audio according to the fundamental frequency information corresponding to each character and the character timestamps corresponding to each fundamental frequency information.
[0148] Each module in the above-mentioned audio adjustment device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0149] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 10As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements an audio adjustment method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0150] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0151] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned audio adjustment method is implemented.
[0152] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the above-mentioned audio adjustment method is implemented.
[0153] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the above-mentioned audio adjustment method is implemented.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties.
[0155] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0156] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0157] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An audio adjustment method, characterized in that The method includes: Select an audio template and record the human voice audio corresponding to the audio template; Obtain the fundamental frequency information corresponding to the human voice audio and identify the text in the human voice audio and the timestamps corresponding to the text characters; Determine the fundamental frequency information corresponding to the text characters based on the fundamental frequency information and the timestamps corresponding to the text characters; Adjust the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio parts corresponding to each text character in the audio template to obtain the target fundamental frequency information; the target fundamental frequency information includes target pitch information, and the target pitch information includes partial target pitch information of each text character in multiple frames; Determine the average pitch information of each text character in each frame according to the partial target pitch information of each text character in each frame and the average value of the partial target pitch information of adjacent frames of each frame, so as to obtain the average pitch information of each text character in multiple frames; Perform pitch adjustment on the partial target pitch information of each text character in multiple frames according to the offset of the partial target pitch information of each text character in multiple frames compared with the average pitch information of the corresponding frame, and the template pitch information corresponding to each text character, so as to determine the adjusted target human voice audio; Obtain the template accompaniment corresponding to the audio template, and perform a fusion process on the target human voice audio and the template accompaniment to obtain the adjusted target audio.
2. The method according to claim 1, characterized in that The selection of the audio template and the recording of the human voice audio corresponding to the audio template include: Display at least one audio template to be selected; Receive a selection instruction for the at least one audio template to be selected, and determine the selected audio template; Display the text corresponding to the audio template, and record the human voice audio input by the user based on the text corresponding to the audio template.
3. The method according to claim 1, wherein The identification of the text in the human voice audio and the timestamps corresponding to the text characters includes: Identify the original text corresponding to the human voice audio according to the fundamental frequency information; Modify the original text according to the matching result between the original text and the text corresponding to the audio template to obtain the text corresponding to the human voice audio; each text character in the text corresponding to the human voice audio matches each text character in the text corresponding to the audio template; Determine the timestamps of each text character in the human voice audio according to the duration of the audio corresponding to each text character in the text corresponding to the human voice audio in the human voice audio.
4. The method according to claim 1, characterized in that The obtaining of the fundamental frequency information corresponding to the human voice audio includes: Obtain the pitch information, timbre information, and vocalization feature information corresponding to the human voice audio as the fundamental frequency information.
5. The method according to claim 4, wherein The adjustment of the fundamental frequency duration of the text characters in the human voice audio based on the duration information of the audio parts corresponding to each text character in the audio template to obtain the target fundamental frequency information includes: Obtain the pitch information, timbre information, and vocalization feature information of the audio parts corresponding to each text character in the human voice audio and the corresponding text characters in the audio template as the fundamental frequency information corresponding to each text character in the human voice audio; Perform one-dimensional linear interpolation processing on the pitch information, timbre information, and vocalization feature information of each text word in the human voice audio respectively according to the duration information of the audio part corresponding to each text word in the audio template, so that the duration of the pitch information, the duration of the timbre information, and the duration of the vocalization feature information respectively match the duration information of the audio part corresponding to each text word in the audio template; Use the interpolated target pitch information, target timbre information, and target vocalization feature information as the target fundamental frequency information.
6. The method according to claim 5, characterized in that, The determining the average pitch information of each text word in each frame according to the partial target pitch information of each text word in each frame and the average value of the partial target pitch information of adjacent frames of each frame includes: Among the partial pitch information of each text word in multiple frames, determine a preset number of adjacent frames adjacent to each frame, and use the partial target pitch information of the adjacent frames as the adjacent partial target pitch information; Obtain the average pitch information of each text word in each frame according to the average value of the partial target pitch information of each text word in each frame and the corresponding adjacent partial target pitch information.
7. The method according to claim 4, wherein The audio template further includes: a reference pitch and a reference frequency corresponding to the reference pitch; the obtaining the pitch information, timbre information, and vocalization feature information corresponding to the human voice audio includes: Obtain the text words corresponding to each character of the audio template in the text of the human voice audio; Obtain the partial human voice audio corresponding to each of the text words in the human voice audio; For the partial human voice audio corresponding to each text word, obtain the partial fundamental frequency corresponding to the partial human voice audio; determine the pitch offset value of the pitch information corresponding to the partial human voice audio according to the partial fundamental frequency corresponding to the partial human voice audio and the reference frequency, and determine the pitch information corresponding to the partial human voice audio according to the reference pitch and the pitch offset value; Construct an envelope matrix according to the envelope vectors of a preset number of frequency points in the partial human voice audio to obtain the timbre information; Obtain the vocalization feature information according to the aperiodic information in a preset number of frequency bands in the partial human voice audio.
8. The method according to claim 7, wherein The determining the adjusted target human voice audio includes: Determine the frequency adjustment value of the target pitch information after pitch adjustment according to the target pitch information after pitch adjustment and the reference pitch of the audio template; Determine the adjusted target frequency corresponding to the human voice audio according to the reference frequency of the audio template and the frequency adjustment value; Determine the adjusted target human voice audio according to the target frequency, target timbre information, and target vocalization feature information.
9. The method according to claim 1, wherein The fusing the target human voice audio with the template accompaniment to obtain the adjusted target audio includes: Obtain the template beat of the template accompaniment; Match the target human voice audio with the template accompaniment according to the template beat; Perform a mixing process on the matched audio to obtain the adjusted target audio.
10. The method according to claim 9, wherein The obtaining the template beat of the template accompaniment includes: Obtain the audio energy values at each time point in the template accompaniment, and use the time points corresponding to the audio energy values greater than the preset energy threshold as the downbeat timestamps of the template accompaniment; Determine the template beat of the template accompaniment according to multiple said reshooting timestamps.
11. The method according to claim 9, wherein Perform mixing processing on the audio with matching completed to obtain the adjusted target audio, including: Adjust the audio energy of the target human voice audio in the audio with matching completed according to the audio energy of the template accompaniment in the audio with matching completed, so that the audio energy of the target human voice audio is less than the audio energy of the template accompaniment; Superimpose and mix the target human voice audio with adjusted audio energy and the template accompaniment to obtain the adjusted target audio.
12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Synthetic method of songs and terminal
CN106898340A