Sound modification method, computer device and computer readable storage medium

By aligning and editing the lyric information and basic frequency sequence of the original audio data, using the template audio data as a reference, the pitch of the user's singing is efficiently improved, and the problem of long training time in the prior art is solved.

CN115101080BActive Publication Date: 2025-08-19TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210698969.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-08-19
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

The existing sound-recommended technology relies on model training to take a long time and cannot efficiently improve the pitch of user singing.

Method used

By obtaining the lyrics information and the basic frequency sequence of the original audio data, the template audio data is used for alignment processing, the target lyrics information and the basic frequency sequence are obtained, and the sound editing is performed based on the basic frequency sequence of the template audio data, and the modified fundamental frequency sequence is obtained.

Benefits of technology

It realizes efficient sound editing, improves the pitch of user singing, simplifies the sound editing process, and reduces training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115101080B_ABST
    Figure CN115101080B_ABST
Patent Text Reader

Abstract

This application discloses a method for audio correction, a computer device, and a computer-readable storage medium. The method comprises: obtaining lyrics information corresponding to original audio data and a fundamental frequency sequence corresponding to the original audio data; aligning the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information; aligning the fundamental frequency sequence corresponding to the original audio data based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence; and performing audio correction on the target fundamental frequency sequence based on the fundamental frequency sequence corresponding to the template audio data to obtain a corrected fundamental frequency sequence corresponding to the original audio data. Using the solution provided by this application, audio correction can be completed efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a sound modification method, a computer device, and a computer-readable storage medium. Background Art

[0002] Singing is a popular way to express emotions and entertain people. However, most non-professional singers lack the ability to accurately grasp pitch like professional singers or vocalists. Therefore, existing singing software often includes features such as pitch correction to improve the pitch of users' singing, making their performances more beautiful and melodious.

[0003] Existing pitch correction technologies typically use network-based model training to adjust the user's singing audio data to achieve the purpose of changing the pitch. This model-based pitch correction method requires a lot of time to train the model in the early stage. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for audio correction, which can efficiently perform audio correction.

[0005] In a first aspect, an embodiment of the present application provides a method for tuning audio, the method comprising:

[0006] Obtaining lyrics information corresponding to the original audio data and a fundamental frequency sequence corresponding to the original audio data;

[0007] Aligning the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information, wherein a time position corresponding to any word in the target lyrics information in the target lyrics information is the same as a time position corresponding to the lyrics information corresponding to the template audio data;

[0008] Based on the target lyrics information, aligning the fundamental frequency sequence corresponding to the original audio data with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence;

[0009] Based on the fundamental frequency sequence corresponding to the template audio data, the target fundamental frequency sequence is subjected to tone modification processing to obtain a tone modified fundamental frequency sequence corresponding to the original audio data.

[0010] In a possible implementation, the aligning the fundamental frequency sequence corresponding to the original audio data based on the target lyrics information to obtain the target fundamental frequency sequence includes:

[0011] removing the fundamental frequencies whose fundamental frequency values are less than the effective threshold from the beginning and the end of the fundamental frequency sequence corresponding to the original audio data to obtain a first effective fundamental frequency sequence;

[0012] Based on the target lyrics information, the first effective fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0013] In a possible implementation, the aligning the first effective fundamental frequency sequence based on the target lyrics information to obtain a target fundamental frequency sequence includes:

[0014] Determining whether there is a fundamental frequency in the first valid fundamental frequency sequence whose fundamental frequency value is less than the valid threshold;

[0015] If there is a fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence, reassigning the fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence to obtain a second effective fundamental frequency sequence, where the second effective fundamental frequency sequence is a continuous sequence;

[0016] Based on the target lyrics information, the second effective fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0017] In a possible implementation, the aligning the second effective fundamental frequency sequence based on the target lyrics information to obtain a target fundamental frequency sequence includes:

[0018] If the length of the second effective fundamental frequency sequence is different from the length of the fundamental frequency sequence corresponding to the template audio data, performing a stretching process on the second effective fundamental frequency sequence to obtain an effective fundamental frequency sequence of equal length, wherein the length of the effective fundamental frequency sequence of equal length is the same as the length of the fundamental frequency sequence corresponding to the template audio data;

[0019] Based on the target lyrics information, the effective fundamental frequency sequences of equal length are aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0020] In a possible implementation, the aligning the effective fundamental frequency sequences of equal length based on the target lyrics information to obtain a target fundamental frequency sequence includes:

[0021] Based on the target lyrics information, the effective fundamental frequency sequence of equal length is shifted in the time domain to obtain a target fundamental frequency sequence corresponding to the target lyrics information.

[0022] In a possible implementation, performing pitch trimming on the target pitch sequence based on the pitch sequence corresponding to the template audio data to obtain the trimmed pitch sequence corresponding to the original audio data includes:

[0023] Determining a first mean of a fundamental frequency sequence corresponding to the template audio data and a second mean of the target fundamental frequency sequence;

[0024] The target fundamental frequency sequence is subjected to a tone-modification process according to the first mean, the second mean, and the fundamental frequency sequence of the template audio data to obtain a tone-modified fundamental frequency sequence corresponding to the original audio data.

[0025] In a possible implementation, performing pitch trimming on the target pitch sequence based on the first mean, the second mean, and the pitch sequence of the template audio data to obtain a pitch trimmed pitch sequence corresponding to the original audio data includes:

[0026] Calculating a difference between the first mean and the second mean as a mean difference;

[0027] Adjusting the fundamental frequency sequence of the template audio data according to the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data;

[0028] The target fundamental frequency sequence is subjected to a tone-modification process according to the adjusted fundamental frequency sequence of the template audio data to obtain a tone-modified fundamental frequency sequence corresponding to the original audio data.

[0029] In a possible implementation, performing pitch trimming on the target pitch sequence according to the adjusted pitch sequence of the template audio data to obtain the pitch trimmed pitch sequence corresponding to the original audio data includes:

[0030] Smoothing the target fundamental frequency sequence to obtain a trend corresponding to the target fundamental frequency sequence;

[0031] According to the trend and the adjusted fundamental frequency sequence of the template audio data, the target fundamental frequency sequence is subjected to a tuning process to obtain a tuned fundamental frequency sequence corresponding to the original audio data.

[0032] In a second aspect, the present application provides a sound correction device, comprising:

[0033] An acquisition unit, configured to acquire lyrics information corresponding to the original audio data and a fundamental frequency sequence corresponding to the original audio data;

[0034] a processing unit configured to align the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information, wherein a time position corresponding to any word in the target lyrics information is the same as a time position corresponding to the lyrics information corresponding to the template audio data;

[0035] The processing unit is further configured to align the fundamental frequency sequence corresponding to the original audio data based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence;

[0036] The processing unit is further configured to perform tone trimming on the target tone sequence based on the tone sequence corresponding to the template audio data to obtain a tone trimmed tone sequence corresponding to the original audio data.

[0037] In a possible implementation, the processing unit is further configured to remove fundamental frequencies having fundamental frequency values less than a valid threshold at the beginning and end of the fundamental frequency sequence corresponding to the original audio data, to obtain a first valid fundamental frequency sequence;

[0038] The processing unit is further configured to perform alignment processing on the first effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0039] In a possible implementation, the processing unit is further configured to determine whether there is a fundamental frequency in the first valid fundamental frequency sequence whose fundamental frequency value is less than the valid threshold;

[0040] The processing unit is further configured to, if there is a fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence, reassign the fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence to obtain a second effective fundamental frequency sequence, where the second effective fundamental frequency sequence is a continuous sequence;

[0041] The processing unit is further configured to perform alignment processing on the second effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0042] In a possible implementation, the processing unit is further configured to, if the length of the second effective fundamental frequency sequence is different from that of the fundamental frequency sequence corresponding to the template audio data, perform scaling processing on the second effective fundamental frequency sequence to obtain an effective fundamental frequency sequence of equal length, where the length of the effective fundamental frequency sequence of equal length is the same as that of the fundamental frequency sequence corresponding to the template audio data;

[0043] The processing unit is further configured to perform alignment processing on the effective fundamental frequency sequences of equal length based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0044] In a possible implementation, the processing unit is further configured to shift the effective fundamental frequency sequence of equal length in the time domain based on the target lyrics information to obtain a target fundamental frequency sequence corresponding to the target lyrics information.

[0045] In a possible implementation, the processing unit is further configured to determine a first mean of a fundamental frequency sequence corresponding to the template audio data and a second mean of the target fundamental frequency sequence;

[0046] The processing unit is further configured to perform tone trimming on the target fundamental frequency sequence according to the first mean, the second mean, and the fundamental frequency sequence of the template audio data to obtain a tone trimmed fundamental frequency sequence corresponding to the original audio data.

[0047] In a possible implementation manner, the processing unit is further configured to calculate a difference between the first mean and the second mean as a mean difference;

[0048] The processing unit is further configured to adjust the fundamental frequency sequence of the template audio data according to the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data;

[0049] The processing unit is further configured to perform pitch trimming processing on the target pitch sequence according to the adjusted pitch sequence of the template audio data to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0050] In a possible implementation, the processing unit is further configured to perform smoothing on the target fundamental frequency sequence to obtain a trend corresponding to the target fundamental frequency sequence;

[0051] The processing unit is further configured to perform pitch trimming on the target pitch sequence according to the trend and the adjusted pitch sequence of the template audio data to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0052] The present application provides a computer device, which includes: a processor, a memory, and a network interface; the processor is connected to the memory and the network interface, wherein the network interface is used to provide network communication functions, the memory is used to store program code, and the processor is used to call the program code to execute the method described in the first aspect.

[0053] The present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the method described in the first aspect is executed.

[0054] The embodiment of the present application first obtains the lyrics information corresponding to the original audio data and the fundamental frequency sequence corresponding to the original audio data. Then, the lyrics information corresponding to the original audio data is aligned according to the lyrics information corresponding to the template audio data, and the lyrics information corresponding to the aligned original audio data is used as the target lyrics information. Through this alignment process, the time position corresponding to any word in the target lyrics information is the same as the time position corresponding to the lyrics information corresponding to the template audio data. Then, based on the target lyrics information, the fundamental frequency sequence corresponding to the original audio data is aligned to obtain a target fundamental frequency sequence that is aligned with the fundamental frequency sequence corresponding to the template audio data. Finally, according to the fundamental frequency sequence corresponding to the template audio data, the target fundamental frequency sequence is tuned to obtain the tuned fundamental frequency sequence. Through the above method, tuning can be completed efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0056] Figure 1 1 is a flow chart of the audio correction method provided in an embodiment of the present application;

[0057] Figure 2 This is a schematic diagram of an embodiment provided by the present application;

[0058] Figure 3 This is another embodiment provided by the present invention;

[0059] Figure 4 This is a schematic diagram of another embodiment provided in the embodiments of the present application;

[0060] Figure 5 This is another embodiment provided by the present application;

[0061] Figure 6 Schematic diagram of the structure of the tuning device provided in an embodiment of the present application;

[0062] Figure 7 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0064] In order to better understand the embodiments of the present application, some professional terms involved in the embodiments of the present application are introduced below:

[0065] Fundamental frequency: In sound, fundamental frequency refers to the frequency of the fundamental tone in a complex sound. Among the multiple tones that make up a complex sound, the fundamental tone has the lowest frequency and the highest intensity. The fundamental frequency determines the pitch of a note. In music, pitch refers to the human psychological perception of the fundamental frequency of a note. Commonly referred to as "off-tune" refers to a mismatch between the singer's pitch and the note.

[0066] The audio correction method provided in the embodiment of the present application can be executed by a terminal device, including but not limited to: smart phones, tablet computers, laptop computers and other devices. Alternatively, it can be executed by a chip or a chip system, the chip can perform audio correction processing on the original audio data, and the chip can be embedded in the terminal device. Alternatively, it can be executed by a server, including but not limited to: an independent physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms and other basic cloud computing services. Alternatively, it can be executed by other devices, which is not limited in this application.

[0067] See also Figure 1 , is a flow chart of a method for tuning a sound provided in an embodiment of the present application. The method for tuning a sound includes S101 to S104. Among them:

[0068] S101: Acquire lyrics information corresponding to original audio data and a fundamental frequency sequence corresponding to the original audio data.

[0069] The original audio data may be a song sung by a user, audio data downloaded by a user from the Internet, or audio data acquired in real time by a microphone of a terminal device.

[0070] Optionally, original audio data is obtained, fundamental frequency is extracted from the original audio data to obtain a fundamental frequency sequence corresponding to the original audio data, speech recognition is performed on the original audio data to obtain lyrics information corresponding to the original audio data.

[0071] Optionally, the lyric information corresponding to the original audio data includes the lyrics and the timestamps corresponding to each word in the lyrics in the original audio data. For example, a certain original audio data includes a lyric: "Still remember your smile", and the lyric information corresponding to this lyric is: [264,2686] yet(264,188) still(453,268) remember(721,289) your(1337,207) smile(1936,245) face(2181,769). The 264 within the square brackets indicates that the start time of this lyric in the original audio data is the 264th ms, and 2686 indicates that the time occupied by this lyric during playback is 2686 ms. Taking the word "still" as an example, its corresponding 453 indicates that the start time of the word "still" in the original audio data is the 453rd ms, and 268 indicates that the time occupied by the word "still" during the playback of the lyric "Still remember your smile" is 268 ms.

[0072] S102. Align the lyric information corresponding to the original audio data based on the lyric information corresponding to the template audio data to obtain the target lyric information, where the time position of any word in the target lyric information is the same as the time position in the lyric information corresponding to the template audio data.

[0073] After splitting the lyric information corresponding to the original audio data by word, align the split lyric information corresponding to the original audio data with the lyric information corresponding to the template audio data word by word to obtain the target lyric information.

[0074] Exemplarily, the lyric information corresponding to the template audio data is: your(1337,207) smile(1936,245) face(2181,769), and the lyric information corresponding to the original audio data is: your(1339,209) smile(1934,242) face(2185,762). Align the lyric information corresponding to the original audio data with the lyric information corresponding to the template audio data to obtain the target lyric information: your(1337,207) smile(1936,245) face(2181,769). It can be seen that the time position of any word in the target lyric information is the same as the time position in the lyric information corresponding to the template audio data.

[0075] The song corresponding to the template audio data and the song corresponding to the original audio data are the same song (e.g., if the song sung in the original audio data is "Qinghai-Tibet Plateau," the song corresponding to the template audio data is also "Qinghai-Tibet Plateau"). The template audio data can be audio data obtained by a professional singer singing the same lyrics, or the template audio data can be audio data generated based on the music score of the same song. In other words, the pitch accuracy of the template audio data is higher than that of the original audio data.

[0076] Optionally, the format of the lyrics information corresponding to the template audio data is the same as the format of the lyrics information corresponding to the original audio data. The format of the target lyrics information is the same as the format of the lyrics information corresponding to the original audio data.

[0077] S103 : Based on the target lyrics information, align the fundamental frequency sequence corresponding to the original audio data with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0078] Optionally, based on the fundamental frequency sequence corresponding to each word in the target lyrics information, the fundamental frequency sequence corresponding to the original audio data is aligned to obtain a target fundamental frequency sequence. The time position of the fundamental frequency sequence corresponding to any word in the target fundamental frequency sequence is the same as the time position in the fundamental frequency sequence corresponding to the template audio data. It can be understood that the fundamental frequency sequence corresponding to any word is composed of the fundamental frequencies of several sounds that constitute the word. For example, the fundamental frequency sequence corresponding to the word "I" in the target lyrics information in the target fundamental frequency sequence is: a first fundamental frequency sequence. The time position of the first fundamental frequency sequence in the target fundamental frequency sequence is: 400ms, and the duration is 50ms. In the lyrics information of the template audio data, the fundamental frequency sequence corresponding to the same word "I" in the fundamental frequency sequence of the template audio data is: a second fundamental frequency sequence. The time position of the second fundamental frequency sequence in the fundamental frequency sequence of the template audio data is also: 400ms, and the duration is 50ms.

[0079] In a possible embodiment, based on the target lyric information, the fundamental frequency sequence corresponding to the original audio data is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence. Specifically, the fundamental frequencies whose fundamental frequency values are less than a valid threshold at the beginning and end of the fundamental frequency sequence corresponding to the original audio data are removed to obtain a first valid fundamental frequency sequence; based on the target lyric information, the first valid fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0080] Among them, the fundamental frequencies whose fundamental frequency values are less than the effective threshold at the beginning and end of the fundamental frequency sequence corresponding to the original audio data may be unvoiced sounds (sounds in which the vocal cords do not vibrate when pronounced).

[0081] Songs often begin and end with an intro and an outro, which are usually silent. Removing the silent portions (those with a fundamental frequency value below a valid threshold) from the beginning and end of the fundamental frequency sequence corresponding to the original audio data before performing alignment can more efficiently obtain the target fundamental frequency sequence.

[0082] In a possible embodiment, based on the target lyric information, an alignment process is performed on a first effective fundamental frequency sequence with a fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence. Specifically, the alignment process comprises: determining whether there is a fundamental frequency with a fundamental frequency value less than a valid threshold value in the first effective fundamental frequency sequence; if there is a fundamental frequency with a fundamental frequency value less than the valid threshold value in the first effective fundamental frequency sequence, reassigning the fundamental frequency with a fundamental frequency value less than the valid threshold value in the first effective fundamental frequency sequence to obtain a second effective fundamental frequency sequence, where the second effective fundamental frequency sequence is a continuous sequence; and based on the target lyric information, an alignment process is performed on the second effective fundamental frequency sequence with a fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0083] Among them, the fundamental frequencies whose fundamental frequency values are less than the effective threshold in the first effective fundamental frequency sequence can be reassigned by linear interpolation, thereby obtaining a continuous second effective fundamental frequency sequence, such as Figure 2 As shown in the figure marked with 21.

[0084] When singing a song, people often take a breath during certain words, and the fundamental frequency value during this breath is usually 0. By assigning a value to the fundamental frequencies in the first valid fundamental frequency sequence whose fundamental frequency values are less than the valid threshold (such as the fundamental frequency at the breath moment mentioned above) and then performing alignment processing, the final sound correction result is more beautiful.

[0085] In a possible embodiment, based on the target lyric information, the second effective baseband sequence is aligned with the baseband sequence corresponding to the template audio data as a reference to obtain the target baseband sequence. Specifically, if the second effective baseband sequence is different in length from the baseband sequence corresponding to the template audio data, the second effective baseband sequence is stretched and contracted to obtain an effective baseband sequence of equal length, where the length of the effective baseband sequence of equal length is the same as the length of the baseband sequence corresponding to the template audio data; and the effective baseband sequences of equal length are aligned based on the target lyric information to obtain the target baseband sequence.

[0086] Among them, the length of the base frequency sequence corresponding to any word in the equal-length effective base frequency sequence is the same as the length of the base frequency sequence corresponding to the arbitrary word in the template audio data. For example, the length of the word "I" in the equal-length effective base frequency sequence is 80ms, and the length of the base frequency sequence corresponding to the same word "I" in the template audio data is also 80ms. That is to say, not only is the length of the equal-length effective base frequency sequence the same as the length of the base frequency sequence corresponding to the template audio data, but the length of the base frequency sequence corresponding to each word in the equal-length effective base frequency sequence is also the same as the length of the base frequency sequence corresponding to the template audio data. For example, Figure 2 As shown in the figure marked with 22, the length of the equal-length effective base frequency sequence is the same as the length of the base frequency sequence corresponding to the template audio data.

[0087] Optionally, when the length of the baseband sequence corresponding to any character in the second effective baseband sequence differs from the length of the baseband sequence corresponding to the same character in the template audio data, the second effective baseband sequence is stretched to obtain an effective baseband sequence of equal length. In other words, in addition to determining whether the second effective baseband sequence and the baseband sequence corresponding to the template audio data are of the same overall length, determining whether stretching is required can also be determined by determining whether the lengths of the baseband sequences corresponding to any character are the same.

[0088] Optionally, when the length of the second effective fundamental frequency sequence is different from that of the fundamental frequency sequence corresponding to the template audio data, the second effective fundamental frequency sequence is stretched and contracted word by word to obtain an effective fundamental frequency sequence of equal length.

[0089] For example, the lyrics information corresponding to the template audio data is: But (264, 188) still (453, 268) remember (721, 289) your (1009, 328) (1337, 207) (1545, 391) smile (1936, 245) appearance (2181, 769). Taking the first word "Que" in the lyrics information as an example, the word "Que" corresponds to baseband sequence 1 (264, 188) in the baseband sequence corresponding to the template audio data. The baseband sequence 1 starts at 264ms and lasts for 188ms. In the second valid baseband sequence, the word "Que" corresponds to baseband sequence 2 (300, 198). The baseband sequence 2 starts at 300ms and lasts for 198ms. For the character "que," shorten the duration of baseband sequence 2 to 188, making it the same as the duration of baseband sequence 1. Similarly, if the duration of baseband sequence 2 is shorter than that of baseband sequence 1, extend it to the same length as baseband sequence 1. The same logic applies to the baseband sequences corresponding to the remaining characters. Once all the characters have the same baseband sequence length, the length of the second valid baseband sequence is the same as the baseband sequence corresponding to the template audio data.

[0090] By aligning the second effective fundamental frequency sequence with the fundamental frequency sequence corresponding to the template audio data, no deviation will occur during subsequent audio correction, and the audio correction result is more accurate.

[0091] In a possible embodiment, based on the target lyric information, the equal-length effective fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain the target fundamental frequency sequence. Specifically, based on the target lyric information, the equal-length effective fundamental frequency sequence is shifted in the time domain to obtain the target fundamental frequency sequence corresponding to the target lyric information.

[0092] The effective base frequency sequence of equal length is horizontally shifted in the time domain (leftward / rightward) so that the starting time of the effective base frequency sequence of equal length is the same as the starting time of the base frequency sequence corresponding to the template audio data. Figure 2 As shown in the figure marked with 23 in the figure. It should be noted that, because the length of the base frequency sequence corresponding to any word in the equal-length effective base frequency sequence is the same as the length of the base frequency sequence corresponding to the any word in the template audio data. When the starting time of the equal-length effective base frequency sequence is aligned with the base frequency sequence corresponding to the template audio data, the starting time of the base frequency sequence corresponding to any word will also be aligned. Figure 2 As shown in the figure marked with 23.

[0093] Exemplarily, the lyric information corresponding to the template audio data is: Your (1337,207) smile (1936,245) appearance (2181,769). The fundamental frequency sequences corresponding to each character in the lyric information corresponding to the template audio data are: fundamental frequency sequence 1, fundamental frequency sequence 2, fundamental frequency sequence 3, and fundamental frequency sequence 4. The fundamental frequency sequence corresponding to the template audio data is: fundamental frequency sequence 1(1337,207) - fundamental frequency sequence 2(1545,391) - fundamental frequency sequence 3(1936,245) - fundamental frequency sequence 4(2181,769). The start time of this fundamental frequency sequence is 1337 ms. The equal-length valid fundamental frequency sequence is: fundamental frequency sequence 5(1339,209) - fundamental frequency sequence 6(1547,393) - fundamental frequency sequence 7(1938,247) - fundamental frequency sequence 8(2183,771). The start time of this equal-length valid sequence is 1339 ms. After aligning this equal-length valid fundamental frequency sequence with the fundamental frequency sequence corresponding to the template audio data, the target fundamental frequency sequence is obtained: fundamental frequency sequence 5(1337,207) - fundamental frequency sequence 6(1545,391) - fundamental frequency sequence 7(1936,245) - fundamental frequency sequence 8(2181,769). It can be seen that the start time of the target fundamental frequency sequence is the same as that of the fundamental frequency sequence corresponding to the template audio data. Moreover, any character's corresponding fundamental frequency sequence is also aligned. Taking the character "笑" as an example, the character "笑" corresponds to fundamental frequency sequence 3 in the fundamental frequency sequence corresponding to the template audio data, and the character "笑" corresponds to fundamental frequency sequence 7 in the target fundamental frequency sequence. It can be seen that the start time and duration of fundamental frequency sequence 3(1936,245) and fundamental frequency sequence 7(1936,245) are the same.

[0094] By aligning the equal-length valid fundamental frequency sequence with the fundamental frequency sequence corresponding to the template audio data, the target fundamental frequency sequence is obtained. Subsequently, pitch correction is performed on this target fundamental frequency sequence to make the result after pitch correction more accurate.

[0095] For a better understanding of how to obtain the target fundamental frequency sequence after aligning the fundamental frequency sequence corresponding to the original audio data. The following combines Figure 2 To further explain the process of obtaining the target fundamental frequency sequence.

[0096] Refer to Figure 2 the figure marked 21 in Figure 2 First, remove the parts with a fundamental frequency value of 0 at the beginning and end of the fundamental frequency sequence corresponding to the original audio data to obtain the first valid fundamental frequency sequence. Then, determine whether there is a fundamental frequency with a value of 0 in this first valid fundamental frequency sequence, and reassign values to the fundamental frequencies with a value of 0 in this first valid fundamental frequency sequence to obtain Figure 2As can be seen from the figure marked with 21, the length of the second effective base frequency sequence is shorter than the length of the base frequency sequence corresponding to the template audio data, so the second effective base frequency sequence is lengthened to obtain Figure 2 The effective frequency sequence of equal length in the figure marked with 22. Figure 2 As can be seen from the figure marked with 22, the equal-length effective base frequency sequence is not aligned with the base frequency sequence corresponding to the template audio data, so the equal-length effective base frequency sequence is left-shifted in the time domain so that the starting time of the equal-length effective base frequency sequence and the starting time of the base frequency sequence corresponding to the template audio data are both 0ms. The equal-length effective base frequency sequence after the shift alignment is used as Figure 2 The target base frequency sequence in the figure marked with 23.

[0097] Through the above steps, the fundamental frequency sequence corresponding to the original audio data is processed to obtain a target fundamental frequency sequence, which is then tuned so that the fundamental frequency sequence corresponding to each word can be accurately tuned.

[0098] S104 : Based on the fundamental frequency sequence corresponding to the template audio data, perform pitch trimming on the target fundamental frequency sequence to obtain a pitch trimmed fundamental frequency sequence corresponding to the original audio data.

[0099] Based on the fundamental frequency sequence corresponding to the template audio data, the fundamental frequency values in the target fundamental frequency sequence are adjusted frame by frame, and the adjusted target fundamental frequency sequence is used as the tuned fundamental frequency sequence corresponding to the original audio data.

[0100] Optionally, the fundamental frequency sequence corresponding to the template audio data is used as a reference standard for audio correction, and the fundamental frequency value in the target fundamental frequency sequence is adjusted to be the same as the fundamental frequency value in the fundamental frequency sequence corresponding to the template audio data, so that the fundamental frequency sequence after correction is more accurate.

[0101] In a possible embodiment, based on the fundamental frequency sequence corresponding to the template audio data, a target fundamental frequency sequence is subjected to audio trimming processing to obtain a modified fundamental frequency sequence corresponding to the original audio data. Specifically, the first mean of the fundamental frequency sequence corresponding to the template audio data and a second mean of the target fundamental frequency sequence are determined; and based on the first mean, the second mean, and the fundamental frequency sequence of the template audio data, the target fundamental frequency sequence is subjected to audio trimming processing to obtain a modified fundamental frequency sequence corresponding to the original audio data.

[0102] The first mean may be the average of the fundamental frequency values of the fundamental frequency sequence corresponding to the template audio data, and the second mean may be the average of the fundamental frequency values of the target fundamental frequency sequence. The first mean may also be the median of the fundamental frequency values of the fundamental frequency sequence corresponding to the template audio data, and the second mean may be the median of the fundamental frequency values of the target fundamental frequency sequence. It is understood that after sorting the fundamental frequency values in the fundamental frequency sequence corresponding to the template audio data from small to large or from large to small, the middle value of the sorted fundamental frequency sequence is taken as the first mean. Similarly, after sorting the fundamental frequency values in the target fundamental frequency sequence from small to large or from large to small, the middle value of the sorted fundamental frequency sequence is taken as the second mean. For example, if the fundamental frequency sequence corresponding to the template audio data is [120, 125, 120, 130, 131], and the fundamental frequency sequence corresponding to the template audio data is sorted from small to large to obtain the fundamental frequency sequence [120, 120, 125, 130, 131], the middle value of the sorted fundamental frequency sequence, 125, is taken as the first mean. The first mean and the second mean may also be other values related to the fundamental frequency sequence and the target fundamental frequency sequence corresponding to the template audio data, which is not limited here.

[0103] Optionally, a first mean of the fundamental frequency sequence corresponding to the template audio data and a second mean of the fundamental frequency sequence corresponding to the original audio data are determined; the target fundamental frequency sequence is trimmed according to the first mean and the second mean and the fundamental frequency sequence of the template audio data to obtain a trimmed fundamental frequency sequence corresponding to the original audio data.

[0104] Optionally, a first mean of the fundamental frequency sequence corresponding to the template audio data and a second mean of the first effective fundamental frequency sequence are determined; and the target fundamental frequency sequence is trimmed according to the first mean and the second mean and the fundamental frequency sequence of the template audio data to obtain a trimmed fundamental frequency sequence corresponding to the original audio data.

[0105] Optionally, a first mean of the fundamental frequency sequence corresponding to the template audio data and a second mean of the second effective fundamental frequency sequence are determined; and the target fundamental frequency sequence is trimmed according to the first mean and the second mean and the fundamental frequency sequence of the template audio data to obtain a trimmed fundamental frequency sequence corresponding to the original audio data.

[0106] In a possible embodiment, a target fundamental frequency sequence is trimmed based on the first mean, the second mean, and the fundamental frequency sequence of the template audio data to obtain a trimmed fundamental frequency sequence corresponding to the original audio data. Specifically, the difference between the first mean and the second mean is calculated as the mean difference; the fundamental frequency sequence of the template audio data is adjusted based on the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data; and the target fundamental frequency sequence is trimmed based on the adjusted fundamental frequency sequence of the template audio data to obtain a trimmed fundamental frequency sequence corresponding to the original audio data.

[0107] The difference between the first mean and the second mean is calculated as the mean difference, and the fundamental frequency sequence of the template audio data is shifted on the vertical axis according to the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data. A pitch-modification process is performed on the target fundamental frequency sequence based on the adjusted fundamental frequency sequence of the template audio data to obtain a pitch-modified fundamental frequency sequence corresponding to the original audio data. The shifting of the fundamental frequency sequence of the template audio data on the vertical axis according to the mean difference can be upward or downward.

[0108] Among them, the distribution range of the fundamental frequency value of the adjusted fundamental frequency sequence of the template audio data is closer to the distribution range of the fundamental frequency value of the target fundamental frequency sequence. That is to say, the overall pitch of the fundamental frequency sequence corresponding to the template audio data is higher / lower than the overall pitch of the target fundamental frequency sequence. After the mean adjustment, the overall pitch of the adjusted fundamental frequency sequence of the template audio data is closer to the overall pitch of the target fundamental frequency sequence. Exemplarily, the fundamental frequency sequence corresponding to the template audio data may be obtained by singing by a female voice, and the target fundamental frequency sequence may be obtained by singing by a male voice. Generally speaking, the pitch of a female voice is higher than that of a male voice, so if the work sung by a male voice is tuned based on the fundamental frequency sequence corresponding to the template audio data sung by a female voice, the tuned result may cause the tuned sound to be unnatural and fail to retain the characteristics of the singer. Therefore, it is necessary to adjust the fundamental frequency sequence corresponding to the template audio data sung by a female voice so that the fundamental frequency sequence corresponding to the template audio data sung by a female voice is closer to the fundamental frequency sequence corresponding to the work sung by a male voice, so as to achieve the purpose of retaining the characteristics of the male voice singing and making the tuned work more natural. See Figure 3 ,Depend on Figure 3 It can be seen that the adjusted fundamental frequency sequence of the template audio data is closer to the target fundamental frequency sequence than the fundamental frequency sequence corresponding to the template audio data.

[0109] For example, the fundamental frequency sequence corresponding to the template audio data is: [65, 70, 78, 80, 60]. The first mean of the fundamental frequency sequence corresponding to the template audio data is (65+70+78+80+60) / 5=70.6 notes. The target fundamental frequency sequence is: [44, 50, 56, 59, 40]. The second mean of the target fundamental frequency sequence is (44+50+56+59+40) / 5=49.8. The difference between the first mean and the second mean is: 70.6-49.8=20.8 notes. It can be seen that the first mean is greater than the second mean, that is, the pitch of the fundamental frequency sequence corresponding to the template audio data is higher than the target fundamental frequency sequence, so the fundamental frequency sequence of the template audio data is shifted downward based on this difference. In order to ensure the melodic sense of the sound, the unit is octave (12 notes, i.e. 12 semitones). Therefore, the 20.8 notes are two octaves, i.e. 24 semitones (24 notes). To this end, the fundamental frequency sequence corresponding to the template audio data is shifted down by 24 notes to obtain the adjusted fundamental frequency sequence of the template audio data: [41, 46, 54, 56, 36].

[0110] Optionally, a specific formula for calculating the first mean and the second mean may be as follows:

[0111]

[0112]

[0113] Among them, note ref (i) represents the fundamental frequency sequence corresponding to the template audio data, Represents the target base frequency sequence. N ref The first mean of the fundamental frequency sequence corresponding to the template audio data, N rec Represents the second mean of the target fundamental frequency sequence.

[0114] Optionally, a specific formula for calculating the mean difference between the first mean and the second mean may be as follows:

[0115] ΔN=[(N rec -N ref ) / 12]·12

[0116] Where ΔN represents the mean difference between the first mean and the second mean, N rec Represents the second mean of the target fundamental frequency sequence, N ref The first mean value of the fundamental frequency sequence corresponding to the template audio data.

[0117] Optionally, a specific formula for calculating the adjusted fundamental frequency sequence of the template audio data may be as follows:

[0118] note des (i)=noteref (i)+ΔN

[0119] Among them, note des (i) represents the adjusted fundamental frequency sequence of the template audio data, N ref The first mean of the fundamental frequency sequence corresponding to the template audio data, ΔN represents the mean difference between the first mean and the second mean.

[0120] By adjusting the fundamental frequency sequence corresponding to the template audio data, an adjusted fundamental frequency sequence of the template audio data that is closer to the target fundamental frequency sequence is obtained, and then the target fundamental frequency sequence is tuned according to the adjusted fundamental frequency sequence of the template audio data, so that the final tuned result retains more of the singer's characteristics and is more natural and beautiful.

[0121] In a possible embodiment, a target fundamental frequency sequence is subjected to pitch trimming processing based on the adjusted fundamental frequency sequence of the template audio data to obtain a pitch trimmed fundamental frequency sequence corresponding to the original audio data. Specifically, the target fundamental frequency sequence is smoothed to obtain a trend corresponding to the target fundamental frequency sequence; and the target fundamental frequency sequence is pitch trimmed based on the trend and the adjusted fundamental frequency sequence of the template audio data to obtain a pitch trimmed fundamental frequency sequence corresponding to the original audio data.

[0122] When singing, especially in the high notes, singers are prone to vibrato and other factors that may cause misjudgment of the fundamental frequency direction, such as Figure 4 The broken line marked by 401 in the figure has an overall trend of a downward arc, but there is also an upward part in the broken line, so it is necessary to Figure 4 The broken line marked by 401 is used to analyze the fundamental frequency trend. Figure 4 The arc line marked with 402 is the trend line corresponding to the broken line 401 .

[0123] Optionally, a spline function is used to smooth the target base frequency sequence to obtain the trend of the target base frequency sequence. The calculation formula for smoothing using the spline function is as follows:

[0124]

[0125] Among them, note trend (i) indicates the trend of the target base frequency sequence, smooth indicates smoothing, Represents the target base frequency sequence, M represents the point division required for smoothing, and M is defined here as: M = [0.3·L], where L represents the length of the target base frequency sequence. “·” represents rounding to the nearest integer. b (m) represents the spline function. Define the spline function:

[0126]

[0127] By analyzing the trend of the target fundamental frequency sequence, the fundamental frequency sequence obtained after tuning is made more accurate.

[0128] In a possible embodiment, the target fundamental frequency sequence is subjected to pitch correction processing based on the trend and the adjusted fundamental frequency sequence of the template audio data. The calculation formula for obtaining the pitch-corrected fundamental frequency sequence corresponding to the original audio data can be as follows:

[0129]

[0130] Among them, note mod (i) represents the modified fundamental frequency sequence corresponding to the original audio data, represents the target base frequency sequence, note trend (i) indicates the trend of the target base frequency sequence, note des (i) represents the adjusted fundamental frequency sequence of the template audio data.

[0131] In order to better understand the method for modifying the sound provided by this application, Figure 5 For further explanation. Figure 5 As shown in step 501, the original audio data is subjected to fundamental frequency extraction to obtain a fundamental frequency sequence corresponding to the original audio data. As shown in step 502, speech recognition is performed on the original audio data to obtain lyrics information corresponding to the original audio data. This process is described in S101 above. As shown in step 503, the lyrics information corresponding to the original audio data is aligned with the lyrics information corresponding to the template audio data to obtain target lyrics information. This alignment process is described in S102 above. The start time and duration of any word in the target lyrics information are the same as the start time and duration in the lyrics information corresponding to the template audio data. As shown in step 504, the fundamental frequency sequence corresponding to the template audio data is adjusted using the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data. The fundamental frequency sequence corresponding to the original audio data is then aligned with the fundamental frequency sequence corresponding to the template audio data to obtain a target fundamental frequency sequence. This alignment process is described in S103 above. As shown in step 505, the target fundamental frequency sequence is pitch-modified to obtain a pitch-modified fundamental frequency sequence corresponding to the original audio data. This process is described in S104 above.

[0132] See also Figure 6 , Figure 6 6 is a schematic diagram of the structure of the tuning device provided in the embodiment of the present application. The tuning device provided in the embodiment of the present application includes: an acquisition unit 601 and a processing unit 602.

[0133] An acquisition unit 601 is configured to acquire lyrics information corresponding to the original audio data and a fundamental frequency sequence corresponding to the original audio data;

[0134] A processing unit 602 is configured to align the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information, wherein the time position of any word in the target lyrics information is the same as the time position of any word in the lyrics information corresponding to the template audio data;

[0135] The processing unit 602 is further configured to align the fundamental frequency sequence corresponding to the original audio data with the fundamental frequency sequence corresponding to the template audio data as a reference based on the target lyrics information to obtain a target fundamental frequency sequence;

[0136] The processing unit 602 is further configured to perform pitch trimming on the target pitch sequence based on the pitch sequence corresponding to the template audio data, to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0137] In another implementation, the processing unit 602 is further configured to remove the fundamental frequencies whose fundamental frequency values are less than the effective threshold from the beginning and the end of the fundamental frequency sequence corresponding to the original audio data, to obtain a first effective fundamental frequency sequence;

[0138] The processing unit 602 is further configured to perform alignment processing on the first effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0139] In another implementation, the processing unit 602 is further configured to determine whether there is a fundamental frequency in the first valid fundamental frequency sequence whose fundamental frequency value is less than a valid threshold;

[0140] The processing unit 602 is further configured to, if there are fundamental frequencies in the first effective fundamental frequency sequence whose fundamental frequency values are less than the effective threshold, reassign the fundamental frequencies in the first effective fundamental frequency sequence whose fundamental frequency values are less than the effective threshold to obtain a second effective fundamental frequency sequence, where the second effective fundamental frequency sequence is a continuous sequence;

[0141] The processing unit 602 is further configured to perform alignment processing on the second effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0142] In another implementation, the processing unit 602 is further configured to, if the lengths of the second effective fundamental frequency sequence and the fundamental frequency sequence corresponding to the template audio data are different, perform stretching processing on the second effective fundamental frequency sequence to obtain an effective fundamental frequency sequence of equal length, wherein the length of the effective fundamental frequency sequence of equal length is the same as the length of the fundamental frequency sequence corresponding to the template audio data;

[0143] The processing unit 602 is further configured to perform alignment processing on effective fundamental frequency sequences of equal length based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

[0144] In another implementation, the processing unit 602 is further configured to shift the effective fundamental frequency sequence of equal length in the time domain based on the target lyrics information to obtain a target fundamental frequency sequence corresponding to the target lyrics information.

[0145] In another implementation, the processing unit 602 is further configured to determine a first mean value of a fundamental frequency sequence corresponding to the template audio data and a second mean value of a target fundamental frequency sequence;

[0146] The processing unit 602 is further configured to perform pitch trimming on the target pitch sequence according to the first mean, the second mean, and the pitch sequence of the template audio data to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0147] In another implementation, the processing unit 602 is further configured to calculate a difference between the first mean and the second mean as the mean difference;

[0148] The processing unit 602 is further configured to adjust the fundamental frequency sequence of the template audio data according to the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data;

[0149] The processing unit 602 is further configured to perform pitch trimming on the target pitch sequence according to the adjusted pitch sequence of the template audio data to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0150] In another implementation, the processing unit 602 is further configured to perform smoothing on the target fundamental frequency sequence to obtain a trend corresponding to the target fundamental frequency sequence;

[0151] The processing unit 602 is further configured to perform pitch trimming on the target pitch sequence according to the trend and the adjusted pitch sequence of the template audio data, to obtain a pitch trimmed pitch sequence corresponding to the original audio data.

[0152] It can be understood that the functions of each functional unit of the tuning device provided in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description in the above method embodiment, which will not be repeated here.

[0153] In other feasible embodiments, the tuning device provided in the embodiments of the present application may also be implemented in a combination of software and hardware. As an example, the tuning device provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the tuning method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0154] See Figure 7 , is a structural diagram of a computer device provided in an embodiment of the present application, and the computer device 70 may include a processor 701, a memory 702, a network interface 703 and at least one communication bus 704. Among them, the processor 701 is used to schedule computer programs, and may include a central processing unit, a controller, and a microprocessor; the memory 702 is used to store computer programs, and may include a high-speed random access memory RAM, a non-volatile memory, such as a disk storage device, a flash memory device; the network interface 703 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), to provide data communication functions, and the communication bus 704 is responsible for connecting various communication elements. The computer device 70 may also include an audio interface, which is used to collect audio, and the audio interface may optionally include a microphone. The computer device 70 may correspond to the data processing device 70 mentioned above. The memory 702 is used to store a computer program, which includes program instructions, and the processor 701 is used to execute the program instructions stored in the memory 702 to execute the process described in steps S101 to S104 in the above embodiment, and perform the following operations:

[0155] In one implementation, lyrics information corresponding to the original audio data and a fundamental frequency sequence corresponding to the original audio data are obtained;

[0156] Aligning the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information, wherein the time position of any word in the target lyrics information is the same as the time position of any word in the lyrics information corresponding to the template audio data;

[0157] Based on the target lyrics information, the fundamental frequency sequence corresponding to the original audio data is aligned with the fundamental frequency sequence corresponding to the template audio data to obtain the target fundamental frequency sequence;

[0158] Based on the fundamental frequency sequence corresponding to the template audio data, the target fundamental frequency sequence is modified to obtain the modified fundamental frequency sequence corresponding to the original audio data.

[0159] In a specific implementation, the above computer device can execute the above-mentioned functions through its built-in functional modules. Figure 1 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.

[0160] The present invention also provides a computer-readable storage medium that stores a computer program. The computer program includes program instructions that are executed by a processor to implement Figure 1 For the tuning methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.

[0161] The above-mentioned computer-readable storage medium can be the recommendation model training device provided by any of the aforementioned embodiments or the internal storage unit of the above-mentioned terminal device, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the electronic device. Furthermore, the computer-readable storage medium can also include both the internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0162] The terms "first," "second," "third," "fourth," and the like in the claims, specification, and drawings of this application are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0163] In the specific implementation of this application, data related to user information (such as original audio data, etc.) is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0164] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present application. The appearance of such a phrase in various locations in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, that the embodiments described herein may be combined with other embodiments. The term "and / or" as used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, including such combinations. Those skilled in the art will appreciate that the elements and algorithmic steps of the various examples described in connection with the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0165] The methods and related devices provided by the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided by the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.

Claims

1. A method for tuning a sound, characterized in that: The method comprises: Obtaining lyrics information corresponding to the original audio data and a fundamental frequency sequence corresponding to the original audio data; Aligning the lyrics information corresponding to the original audio data based on the lyrics information corresponding to the template audio data to obtain target lyrics information, wherein a time position of any word in the target lyrics information is the same as a time position of any word in the lyrics information corresponding to the template audio data; Based on the target lyrics information, aligning the fundamental frequency sequence corresponding to the original audio data with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence; Determining a first mean of a fundamental frequency sequence corresponding to the template audio data and a second mean of the target fundamental frequency sequence; The target fundamental frequency sequence is subjected to a tone-modification process according to the first mean, the second mean, and the fundamental frequency sequence of the template audio data to obtain a tone-modified fundamental frequency sequence corresponding to the original audio data.

2. The method according to claim 1, characterized in that The step of aligning the fundamental frequency sequence corresponding to the original audio data based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence includes: removing the fundamental frequencies whose fundamental frequency values are less than the effective threshold from the beginning and the end of the fundamental frequency sequence corresponding to the original audio data to obtain a first effective fundamental frequency sequence; Based on the target lyrics information, the first effective fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

3. The method according to claim 2, characterized in that The step of aligning the first effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence includes: Determining whether there is a fundamental frequency in the first valid fundamental frequency sequence whose fundamental frequency value is less than the valid threshold; If there is a fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence, reassigning the fundamental frequency with a fundamental frequency value less than the effective threshold in the first effective fundamental frequency sequence to obtain a second effective fundamental frequency sequence, where the second effective fundamental frequency sequence is a continuous sequence; Based on the target lyrics information, the second effective fundamental frequency sequence is aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

4. The method according to claim 3, characterized in that The step of aligning the second effective fundamental frequency sequence based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence includes: If the length of the second effective fundamental frequency sequence is different from the length of the fundamental frequency sequence corresponding to the template audio data, performing a stretching process on the second effective fundamental frequency sequence to obtain an effective fundamental frequency sequence of equal length, wherein the length of the effective fundamental frequency sequence of equal length is the same as the length of the fundamental frequency sequence corresponding to the template audio data; Based on the target lyrics information, the effective fundamental frequency sequences of equal length are aligned with the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence.

5. The method according to claim 4, characterized in that The step of aligning the effective fundamental frequency sequences of equal length based on the target lyrics information and taking the fundamental frequency sequence corresponding to the template audio data as a reference to obtain a target fundamental frequency sequence includes: Based on the target lyrics information, the effective fundamental frequency sequence of equal length is shifted in the time domain to obtain a target fundamental frequency sequence.

6. The method according to any one of claims 1, wherein: The performing tone trimming processing on the target tone sequence according to the first mean, the second mean, and the tone trimming sequence of the template audio data to obtain a tone trimmed tone sequence corresponding to the original audio data includes: Calculating a difference between the first mean and the second mean as a mean difference; Adjusting the fundamental frequency sequence corresponding to the template audio data according to the mean difference to obtain an adjusted fundamental frequency sequence of the template audio data; The target fundamental frequency sequence is subjected to a tone-modification process according to the adjusted fundamental frequency sequence of the template audio data to obtain a tone-modified fundamental frequency sequence corresponding to the original audio data.

7. The method according to any one of claim 6, characterized in that The step of performing tone trimming on the target tone sequence according to the adjusted tone sequence of the template audio data to obtain the tone trimmed tone sequence corresponding to the original audio data includes: Smoothing the target fundamental frequency sequence to obtain a trend corresponding to the target fundamental frequency sequence; According to the trend and the adjusted fundamental frequency sequence of the template audio data, the target fundamental frequency sequence is subjected to a tuning process to obtain a tuned fundamental frequency sequence corresponding to the original audio data.

8. A computer device, characterized in that: include: A processor, a communication interface and a memory, wherein the processor, the communication interface and the memory are connected to each other, wherein the memory stores an executable program code, and the processor is used to call the executable program code to execute the audio correction method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the audio modification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and apparatus for determining pitch deviation of audio content

    CN108206026A

  • Fundamental frequency extraction method and device

    CN109346109A

  • Audio adjustment method and computer equipment

    CN114582306A