Audio adjustment method and computer device

The audio adjustment method analyzes user singing difficulty to create personalized templates for pitch and lyric adjustments, enhancing the alignment and quality of user performances with the original audio.

CN114582306BActive Publication Date: 2025-07-15TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210171012.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-07-15
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

The existing audio sound editing method has a stiff effect and cannot effectively adjust the difference between user singing and original song.

Method used

By obtaining the difficulty information of the audio to be adjusted and its standard audio, determining the melody template and lyric template, performing pitch adjustments, including pitch translation, replacement and compression operations, and combining tone change processing, targeted adjustments to the audio are achieved.

Benefits of technology

Improve the audio adjustment effect, allowing users to sing close to the original song, and avoiding the rigidity of fixed sound editing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582306B_ABST
    Figure CN114582306B_ABST
Patent Text Reader

Abstract

The present application relates to an audio adjustment method, device, computer device, storage medium, and computer program product. An adjustment template is determined based on the standard melody information, standard lyric information, and difficulty information of the standard audio corresponding to the audio to be adjusted. The standard lyric information is matched with the audio to be adjusted to obtain an audio sequence to be adjusted including class note units corresponding to each lyric in the standard lyric information. Based on the melody template, lyric template, and identification information for identifying the difficulty information in the adjustment template, pitch adjustment is performed on multiple class note units in the audio sequence to be adjusted, and then the audio to be adjusted is adjusted based on the obtained adjusted audio sequence to obtain the adjusted audio. Compared with the traditional audio adjustment based on a fixed method, this solution analyzes the difficulty information of the standard audio and performs targeted adjustment on the audio to be adjusted according to the difficulty information of the standard audio, improving the adjustment effect of audio adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and particularly to an audio adjustment method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] At present, users can already sing songs through mobile devices such as mobile phones. Since the singing skills of each user are different, there will be a situation where the sung song by the user is different from the original song. At this time, it is necessary to tune the sung song by the user so that the sung song by the user is as close as possible to the original song. Currently, the method of tuning the sung song by the user is usually to tune the sung song by the user based on a fixed method. However, through the fixed tuning method, the tuning effect will be relatively rigid.

[0003] Therefore, the current tuning method has the defect of insufficient tuning effect. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an audio adjustment method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the tuning effect.

[0005] In a first aspect, the present application provides an audio adjustment method, and the method includes:

[0006] Obtain the audio to be adjusted and its corresponding standard audio, and obtain the difficulty information of the standard audio, where the difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyric information;

[0007] Determine an adjustment template based on the standard melody information, the standard lyric information, and the difficulty information of the standard audio; the adjustment template includes a melody template and a lyric template; at least one of the melody template and the lyric template includes identification information for identifying the difficulty information;

[0008] Match the lyric template with the audio to be adjusted to obtain an audio sequence to be adjusted; the audio sequence to be adjusted includes the class note units corresponding to each lyric in the lyric template;

[0009] Based on the melody template, perform pitch adjustment on multiple class note units in the audio sequence to be adjusted to obtain an adjusted audio sequence; adjust the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio.

[0010] In one of the embodiments, the determining an adjustment template based on the standard melody information, the standard lyric information, and the difficulty information of the standard audio includes:

[0011] Based on the standard melody information, obtain a melody template; the target intervals in the melody template are marked with replacement intervals; the target intervals are determined based on the first difficulty information and represent the difficult intervals in the standard melody information.

[0012] Based on the standard lyric information, obtain a lyric template; the target lyric information in the lyric template is marked with ornamentation identification information; the target lyric information is determined based on the second difficulty information and represents the lyrics corresponding to the ornamented melody in the standard lyric information.

[0013] Determine an adjustment template according to the melody template and the lyric template.

[0014] In one embodiment, the obtaining the melody template based on the standard melody information includes:

[0015] Obtain the difficult intervals in the standard melody information according to the pitch difference between adjacent notes in the standard melody information.

[0016] Obtain a replacement interval corresponding to the difficult interval; the interval difference between the replacement interval and the difficult interval is less than a preset interval difference threshold, and the pitch of adjacent notes in the replacement interval is the same as that in the difficult interval.

[0017] Mark the replacement interval at the position of the difficult interval in the standard melody information to obtain a melody template.

[0018] In one embodiment, the obtaining the lyric template based on the standard lyric information includes:

[0019] Obtain the ornamented melody in the standard melody information.

[0020] Add ornamentation identification information to the lyrics corresponding to the ornamented melody in the standard lyric information to obtain a lyric template.

[0021] In one embodiment, the obtaining the difficult intervals in the standard melody information according to the pitch difference between adjacent notes in the standard melody information includes:

[0022] If the magnitude of the pitch difference is greater than or equal to a preset pitch difference threshold, determine that the adjacent notes in the standard melody information are difficult intervals.

[0023] In one embodiment, the adding ornamentation identification information to the lyrics corresponding to the ornamented melody in the standard lyric information to obtain a lyric template includes:

[0024] Obtain the number of characters of the lyrics corresponding to the ornamented melody in the standard melody information.

[0025] Obtain the ratio of the number of words in the lyrics corresponding to the grace note melody to the total number of lyrics in the standard lyrics information;

[0026] If the ratio is greater than a preset grace note threshold, add grace note identification information to the lyrics corresponding to the grace note melody in the standard lyrics information to obtain a lyrics template.

[0027] In one embodiment, before matching the standard lyrics information with the audio to be adjusted, it further includes:

[0028] Obtain the number of long vowel notes with pitch change and pronunciation length greater than a preset length threshold in the standard melody information;

[0029] According to the number of long vowel notes, obtain the proportion of long vowel notes in the standard melody information;

[0030] If the proportion is greater than a preset long vowel probability threshold, obtain the extended vowel corresponding to the long vowel note;

[0031] According to the pronunciation length of the long vowel note and the extended vowel, expand the phoneme of the long vowel note in the lyrics template to obtain the extended lyrics information corresponding to the lyrics template;

[0032] Match based on the extended lyrics information and the audio to be adjusted.

[0033] In one embodiment, the obtaining the proportion of long vowel notes in the standard melody information according to the number of long vowel notes includes:

[0034] Obtain the number of lines of lyrics in the standard lyrics information;

[0035] Based on the ratio of the number of long vowel notes to the number of lines of lyrics, determine the proportion of long vowel notes in the standard melody information.

[0036] In one embodiment, the matching the standard lyrics information with the audio to be adjusted to obtain an audio sequence to be adjusted includes:

[0037] Perform fundamental frequency detection on the audio to be adjusted to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and convert the fundamental frequency sequence into a sequence of note-like units;

[0038] Align each note-like unit in the sequence of note-like units with each lyric in the lyrics template to obtain the audio sequence to be adjusted after word-by-word mapping.

[0039] In one embodiment, based on the adjustment template, performing pitch adjustment on multiple note-like units in the audio sequence to be adjusted to obtain an adjusted audio sequence, including:

[0040] For each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit contains ornamentation identification information, perform pitch translation on the notes in the note-like unit to conform to the melody template;

[0041] If it is detected that the note-like unit contains a replacement interval, traverse each replacement interval of the note-like unit, and replace the note corresponding to the interval in the note-like unit with the replacement note in the replacement interval with the smallest interval difference degree from the interval of the note-like unit;

[0042] Based on the pitch adjustment results of multiple note-like units, obtain the adjusted audio sequence.

[0043] In one embodiment, based on the adjustment template, performing pitch adjustment on multiple note-like units in the audio sequence to be adjusted to obtain an adjusted audio sequence, including:

[0044] For each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit does not contain ornamentation identification information and replacement intervals, perform pitch translation and amplitude compression processing on the notes in the note-like unit to conform to the melody template in the adjustment template;

[0045] Based on the pitch adjustment results of multiple note-like units, obtain the adjusted audio sequence.

[0046] In one embodiment, based on the adjusted audio sequence, adjusting the audio to be adjusted to obtain an adjusted audio, including:

[0047] Based on the adjusted audio sequence, perform key adjustment on the audio to be adjusted to obtain the adjusted audio.

[0048] In a second aspect, the present application provides an audio adjustment device, and the device includes:

[0049] An acquisition module, configured to acquire the audio to be adjusted and its corresponding standard audio, and acquire the difficulty information of the standard audio, where the difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyric information;

[0050] A determination module, configured to determine an adjustment template based on the standard melody information, the standard lyric information, and the difficulty information of the standard audio; the adjustment template includes a melody template and a lyric template; at least one of the melody template and the lyric template includes identification information for identifying the difficulty information;

[0051] An alignment module for matching the lyric template with the audio to be adjusted to obtain an audio sequence to be adjusted; the audio sequence to be adjusted includes class-note units corresponding to each lyric in the lyric template.

[0052] An adjustment module for adjusting the pitch of multiple class-note units in the audio sequence to be adjusted based on the melody template to obtain an adjusted audio sequence; and adjusting the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio.

[0053] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0054] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0055] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0056] The above audio adjustment method, device, computer device, storage medium, and computer program product obtain the audio to be adjusted and its standard audio, and obtain the difficulty information in the standard audio. Determine an adjustment template based on the standard melody information, standard lyric information, and difficulty information of the standard audio, and match the standard lyric information with the audio to be adjusted to obtain an audio sequence to be adjusted including class-note units corresponding to each lyric in the standard lyric information. And based on the melody template, lyric template, and identification information for identifying the difficulty information in the adjustment template, adjust the pitch of multiple class-note units in the audio sequence to be adjusted to obtain an adjusted audio sequence, and then adjust the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio. Compared with the traditional audio adjustment based on a fixed method, this solution analyzes the difficulty information of the standard audio and makes targeted adjustments to the audio to be adjusted according to the difficulty information of the standard audio, improving the adjustment effect of audio adjustment. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is an application environment diagram of the audio adjustment method in an embodiment;

[0058] Figure 2 It is a flowchart of the audio adjustment method in an embodiment;

[0059] Figure 3 Schematic flowchart of steps for obtaining an adjustment template in an embodiment;

[0060] Figure 4 Schematic flowchart of steps for obtaining a melody template in an embodiment;

[0061] Figure 5 Schematic flowchart of steps for expanding vowels in an embodiment;

[0062] Figure 6 Schematic flowchart of steps for alignment in an embodiment;

[0063] Figure 7 Schematic flowchart of steps for alignment in another embodiment;

[0064] Figure 8 Schematic interface diagram of steps for audio adjustment in an embodiment;

[0065] Figure 9 Schematic flowchart of steps for audio adjustment in an embodiment;

[0066] Figure 10 Schematic flowchart of a method for audio adjustment in another embodiment;

[0067] Figure 11 Schematic block diagram of an audio adjustment device in an embodiment;

[0068] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0069] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] The audio adjustment method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. The terminal 102 can obtain the audio to be adjusted input by the user and send the audio to be adjusted to the server 104. The server 104 can perform difficulty analysis and determination of the adjustment method based on the obtained audio to be adjusted. Thus, the server 104 can perform targeted audio adjustment on the audio to be adjusted based on the difficulty information of the standard audio corresponding to the audio to be adjusted. The server 104 can transmit the adjusted audio to the terminal 102 to realize the adjustment of the audio to be adjusted. In addition, in some embodiments, the terminal 102 can also perform difficulty analysis and audio adjustment on the audio to be adjusted. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0071] In one embodiment, as Figure 2 shown, an audio adjustment method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0072] Step S202, obtain the audio to be adjusted and its corresponding standard audio, and obtain the difficulty information of the standard audio. The difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyric information.

[0073] Among them, the audio to be adjusted can be the audio input by the user, such as the song audio recorded by the user when singing. The standard audio can be the standard audio corresponding to the audio to be adjusted. Taking the adjustment of the user's singing audio as an example, the server 104 can obtain the original singing audio corresponding to the user's singing audio as the standard audio. For example, taking the terminal 102 as a mobile phone as an example, the mobile phone can collect the song recording generated after the user sings, and send the song recording to the server 104. The server 104 can obtain the id of the original singing song corresponding to the song recording. After the server 104 receives the sound adjustment task, it confirms that the materials required for the task, including the song recording, the song melody of the original song and the song lyrics file are all complete, and the user's song recording can be adjusted. Among them, the standard audio corresponding to the audio to be adjusted has diversity, and each standard audio has corresponding difficulty information. For example, for songs, there are many styles and categories of songs. Due to different melodies or the characteristics of the singer's voice, the singing difficulty of different songs is also different. The server 104 can grade the difficulty of songs in the sound adjustment scenario through a difficulty grading system to better realize the diversity of sound adjustment strategies. Therefore, after the server 104 obtains the standard audio corresponding to the audio to be adjusted, it can obtain the difficulty information of the standard audio. Among them, the difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyrics information. The first difficulty information and the second difficulty information can exist at the same time or any one of them can exist. The first difficulty information and the second difficulty information can be different difficulty information. Among them, the standard melody information can be the melody information of the standard audio, and the standard lyrics information can be the lyrics information of the standard audio. The server 104 can obtain the above-mentioned standard melody information and standard lyrics information from the standard audio. The above-mentioned first difficulty information can be a difficult example interval with a large pitch span in the standard audio melody. Among them, the interval refers to the difference in pitch between adjacent notes in the melody information of the standard audio. The above-mentioned second difficulty information can be the lyrics and melody of the standard audio using singing techniques such as ornaments. Among them, singing techniques can also be called vocal techniques, which belong to a kind of vocal technology and can be used to evaluate the singing skills of users. The vocal techniques mainly include true voice, false voice, strong voice, weak voice, breathy voice, vibrato, glissando, transposition, pharyngeal voice, mute voice, angry voice, choking voice, crying voice, etc. Different vocal techniques can have different emotional effects on singing.

[0074] When there is difficulty information in the audio, the server 104 can adjust the audio in a targeted manner based on the difficulty information. For example, the terminal 102 uploads the audio to be adjusted input by the user to the server 104, and the server 104 performs MIR (Music Information Retriveal) analysis, audio synthesis and other operations on the audio to be adjusted by the user to achieve the adjustment of the audio to be adjusted.

[0075] Step S204: Determine an adjustment template based on the standard melody information, standard lyric information, and difficulty information of the standard audio. The adjustment template includes a melody template and a lyric template; at least one of the melody template and the lyric template includes identification information for identifying the difficulty information.

[0076] Among them, the standard audio can be the standard audio corresponding to the audio to be adjusted input by the user. The standard audio includes its standard melody information, standard lyric information, and difficulty information. The server 104 can determine the adjustment template based on the standard melody information, standard lyric information, and difficulty information in the standard audio. Among them, the audio to be adjusted can be the original dry voice of the song recording input by the user, the standard audio can be the original singer audio of the song corresponding to the original dry voice, and the adjustment template can be a template for adjusting the audio to be adjusted. The adjustment template contains the standard melody and standard lyrics in the standard audio, as well as the identification information of the difficulty information marked in the standard audio after analyzing the difficulty of the standard audio, that is, the adjustment template includes a melody template and a lyric template, and at least one of the melody template and the lyric template includes identification information for identifying the difficulty information. Among them, there can be multiple types of the above-mentioned difficulty information, and the identification forms of each type of difficulty information can be inconsistent.

[0077] Step S206: Match the lyric template with the audio to be adjusted to obtain an audio sequence to be adjusted; the audio sequence to be adjusted includes note-like units corresponding to each lyric in the lyric template.

[0078] Among them, the standard lyric information can be the standard lyric information corresponding to the standard audio. For example, if the standard audio is the original singer audio corresponding to the audio to be adjusted above, then the original singer audio has corresponding standard lyric information. The server 104 can match the audio to be adjusted and the standard audio based on the standard lyric information to obtain an audio sequence to be adjusted. For example, the server 104 can take each character in the lyric template as a unit, and match the audio to be adjusted and the lyric template word by word to obtain the audio sequence to be adjusted after word-by-word mapping. Among them, the audio sequence to be adjusted includes NLU (Note-Like Unit) corresponding to each lyric in the standard lyric information. The note-like unit can be composed of a section of melody information corresponding to each lyric in the standard melody information of the above standard audio, and this section of melody information can be a waveform information. The above matching can be a lyric alignment technology, and the word-by-word timestamp of the audio corresponding to the lyric content sung by the user in the audio to be adjusted can be obtained.

[0079] Step S208: Based on the melody template, perform pitch adjustment on multiple note-like units in the audio sequence to be adjusted to obtain an adjusted audio sequence; adjust the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio.

[0080] Among them, the adjustment template can be a template obtained based on the standard melody information, standard lyric information, and difficulty information in the standard audio, which includes the melody template, lyric template, and identification information of the difficulty information corresponding to the standard audio. The audio sequence to be adjusted can be a sequence obtained by the server 104 after matching and aligning the audio to be adjusted and the standard audio based on the standard lyric information. The audio sequence to be adjusted includes multiple note-like units. The server 104 can perform pitch adjustment on the multiple note-like units in the audio sequence to be adjusted based on the above adjustment template to obtain an adjusted audio sequence. For example, the server 104 can adjust the multiple pitches included in each note-like unit, and the adjustment includes operations such as translation, compression, and replacement. After the server 104 adjusts all the note-like units in the audio sequence to be adjusted, an adjusted audio sequence can be obtained. After the server 104 obtains the above adjusted audio sequence, it can also adjust the audio to be adjusted based on the adjusted audio sequence, thereby obtaining an adjusted audio, and completing the adjustment of the audio to be adjusted input by the user.

[0081] For example, in one embodiment, adjusting the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio includes: performing pitch shifting on the audio to be adjusted according to the adjusted audio sequence to obtain an adjusted audio. In this embodiment, the server 104 can process the audio to be adjusted input by the user based on the adjusted audio sequence obtained after adjustment. The server 104 can perform pitch shifting on the audio to be adjusted based on the adjusted audio sequence, thereby obtaining an adjusted audio. For example, taking the audio to be adjusted as the original dry voice of the user's singing recording, the above adjusted audio sequence can be a pitch shifting sequence. After the server 104 obtains the processed pitch shifting sequence, it can perform pitch shifting on the user's original dry voice to complete the tuning adjustment of the user's singing recording. Among them, pitch shifting refers to a technology that can change the pitch of the audio.

[0082] That is, the server 104 can tune the user's singing recording. For example, in a KTV application, since the singing methods and techniques of each song are different, it is impossible to express the understanding of the song based on a fixed tuning method. Therefore, the server 104 can analyze the difficulty information of the song through the above method, and then adjust the user's input singing recording according to the difficulty information.

[0083] In the above audio adjustment method, by obtaining the audio to be adjusted and its standard audio, and obtaining the difficulty information in the standard audio, an adjustment template is determined based on the standard melody information, standard lyrics information, and difficulty information of the standard audio. The standard lyrics information is matched with the audio to be adjusted to obtain an audio sequence to be adjusted including class note units corresponding to each lyric in the standard lyrics information. Based on the melody template, lyrics template, and identification information for identifying the difficulty information in the adjustment template, pitch adjustment is performed on multiple class note units in the audio sequence to be adjusted to obtain an adjusted audio sequence. Then, based on the adjusted audio sequence, the audio to be adjusted is adjusted to obtain the adjusted audio. Compared with the traditional audio adjustment based on a fixed method, this solution analyzes the difficulty information of the standard audio and makes targeted adjustments to the audio to be adjusted according to the difficulty information of the standard audio, improving the adjustment effect of audio adjustment.

[0084] In one embodiment, determining an adjustment template based on the standard melody information, standard lyrics information, and difficulty information of the standard audio includes: obtaining a melody template based on the standard melody information; a target interval in the melody template is marked with a replacement interval; the target interval is determined according to the first difficulty information of the standard melody information; the first difficulty information represents difficult example intervals in the standard melody information; obtaining a lyrics template based on the standard lyrics information; the target lyrics information in the lyrics template is marked with ornamentation note identification information; the target lyrics information is determined according to the second difficulty information of the standard lyrics information; the second difficulty information represents the lyrics corresponding to the ornamentation melody in the standard lyrics information; determining the adjustment template according to the melody template and the lyrics template.

[0085] In this embodiment, the server 104 can determine an adjustment template for adjusting the audio to be adjusted based on the standard melody information, standard lyrics information, and difficulty information of the standard audio. Among them, as Figure 3 shown, Figure 3It is a schematic flowchart of the steps for obtaining an adjustment template in an embodiment. The adjustment template includes a melody template and a lyrics template. The server 104 obtains the melody template based on the standard melody information in the standard audio. Among them, the melody template includes multiple notes, and every two adjacent notes can form an interval, so the melody template can include multiple intervals. The server 104 can also replace the target interval in the melody template with a replacement interval, and this identification can be a kind of association process. Among them, the target interval can be determined based on the first difficulty information in the standard melody information, and the first difficulty information is a difficult example interval in the standard melody information, that is, the server 104 can detect the target interval belonging to the difficult example interval from the standard melody information and add a replacement interval for this target interval. Among them, the difficult example interval refers to an interval with a relatively large span between adjacent notes in the above standard melody information, that is, the server 104 can extract the pitch range in the melody, and the replacement interval can be an interval with the same tonality as the corresponding difficult example interval and a smaller interval. There can be multiple replacement intervals, that is, the above difficult example interval can be associated with multiple replacement intervals. The server 104 can replace the difficult example interval with a suitable associated replacement interval during the audio adjustment stage.

[0086] The server 104 can also obtain the lyrics template based on the standard lyrics information of the standard audio. Among them, there may be singing techniques such as grace notes in the standard audio. The server 104 can identify the grace notes in the standard melody information and mark the corresponding lyrics. Thus, the server 104 can identify the target lyrics information in the above lyrics template to obtain the target lyrics information marked with grace note identification information. Among them, the target lyrics information can be determined according to the second difficulty information of the standard lyrics information, and the second difficulty information represents the lyrics corresponding to the grace note melody in the standard lyrics information. Among them, the above grace notes can include various forms, such as turns, vibratos, and slides, etc. Then the server 104 can detect the type of grace notes in the standard melody information and give different grace note identification information to the corresponding lyrics information based on the type of grace notes, such as turn identification, vibrato identification, slide identification, etc. Thus, during the audio adjustment stage, the server 104 can adjust the corresponding pitch in the audio sequence to be adjusted based on the grace note identification information in the lyrics template. Specifically, taking the audio as a song as an example, such as Figure 3As shown, the server 104 can obtain the original audio corresponding to the song recording input by the user, and extract information such as the long vowels of the lyrics, the difficulty of the song melody, and the range span of the song from the original audio. Thus, based on these characteristic information, the server can perform difficulty analysis on the song, obtain corresponding melody templates and lyric templates containing identification information, and determine adjustment templates based on the melody templates and lyric templates. Additionally, in some embodiments, the above-mentioned grace note identification information can also be marked in the melody template. For example, the server 104 can add corresponding grace note identification information at the melody corresponding to the lyrics where grace notes appear in the melody template.

[0087] Through this embodiment, the server 104 can perform difficulty analysis on the standard audio, identify the difficult interval and grace note identification information therein, so that the server 104 can perform targeted adjustment on the audio to be adjusted based on the difficulty information, improving the effect of audio adjustment.

[0088] In one embodiment, obtaining a melody template based on standard melody information includes: obtaining the difficult interval in the standard melody information according to the pitch difference between adjacent notes in the standard melody information; obtaining a replacement interval corresponding to the difficult interval; the interval difference between the replacement interval and the difficult interval is less than a preset interval difference threshold, and the pitch is the same as that of the adjacent notes in the difficult interval; marking the replacement interval at the difficult interval in the standard melody information to obtain the melody template.

[0089] In this embodiment, the server 104 can judge and detect the difficult intervals in the standard melody information. The server 104 can detect and obtain the difficult intervals in the standard melody information based on the pitch difference between adjacent notes in the above-mentioned standard melody information. For example, in one embodiment, obtaining the difficult intervals in the standard melody information according to the pitch difference between adjacent notes in the standard melody information includes: obtaining the magnitude of the pitch difference between adjacent notes in the standard melody information. If the magnitude of the pitch difference is greater than or equal to a preset pitch difference threshold, it is determined that the adjacent notes in the standard melody information are difficult intervals. In this embodiment, the interval includes adjacent notes in the melody information. The server 104 can obtain the magnitude of the pitch difference between adjacent notes in the standard melody information, and obtain the comparison result between the magnitude of the pitch difference and the preset pitch difference threshold. If the server 104 detects that the magnitude of the pitch difference is greater than or equal to the preset pitch difference threshold, the server 104 can determine that the adjacent notes in the standard melody information are difficult intervals.

[0090] Specifically, the server 104 can obtain the magnitude of the pitch difference between adjacent notes through the following formula: Interval = Note i -Note i-1 , Interval >= 6. Where Interval is the magnitude of the pitch difference, Note iFor the i-th note in the standard melody information, as can be seen from the above formula, when the server 104 detects that the pitch difference is greater than or equal to 6, it can be determined as a difficult interval to sing. The server 104 can mark such difficult intervals in the melody template.

[0091] After the server 104 obtains and marks the difficult intervals, it can obtain the replacement intervals corresponding to the difficult intervals. Among them, the interval difference between the replacement interval and the difficult interval is less than the preset interval difference threshold, and the pitch of adjacent notes in the replacement interval is the same as that of adjacent notes in the difficult interval. That is, the server 104 can provide a replacement interval that satisfies the tonality of the difficult interval and has a smaller interval, and there can be multiple replacement intervals.

[0092] After the server 104 obtains the replacement intervals, it can mark the replacement intervals at the difficult intervals in the standard melody information to obtain the melody template. Specifically, as Figure 4 shown, Figure 4 FIG. is a schematic flowchart of obtaining a melody template in an embodiment. The identification information in the melody template can include information such as the pitch range of the standard audio, singing skills (such as ornamentation identification information), and interval difficulty information. Thus, the server 104 can adjust the audio to be adjusted based on the melody template, such as tuning a song.

[0093] Through the above embodiments, the server 104 can analyze the difficult intervals in the standard melody information and obtain the corresponding replacement intervals. Thus, the server 104 can perform targeted adjustment on the audio to be adjusted based on the difficult intervals and the replacement intervals, improving the effect of audio adjustment.

[0094] In an embodiment, based on the standard lyrics information, obtaining a lyrics template includes: obtaining the ornamentation melody in the standard melody information, and adding ornamentation identification information to the lyrics corresponding to the ornamentation melody in the standard lyrics information to obtain the lyrics template.

[0095] In this embodiment, the above template information further includes a lyrics template. The server 104 can obtain the lyrics template based on the standard lyrics information. Moreover, the server 104 can also identify the characters belonging to singing techniques such as grace notes, so that corresponding adjustments can be made to the melody with grace notes during the audio adjustment stage. The server 104 can detect the grace note melody belonging to the grace notes in the standard melody information, and the server 104 can determine the lyrics corresponding to these melodies belonging to the grace notes. Thus, the server 104 can add grace note identification information at the lyrics corresponding to the grace note melody in the standard lyrics information to obtain the lyrics template. Among them, there can be various types of the above grace notes, such as turns, vibratos, and slides, etc. Then the server 104 can add corresponding grace note identification information in the standard lyrics information based on the type of the grace note, such as turn marks, vibrato marks, and slide marks, etc. It should be noted that in some embodiments, the server 104 can also add the above grace note identification information to the grace note melody in the standard melody information. Thus, the server 104 can adjust the audio to be adjusted input by the user based on the melody template formed by the standard melody information containing the grace note identification information during the audio adjustment.

[0096] Among them, the server 104 can also count the frequency of using singing techniques in the standard audio, and add the grace note identification information only when the frequency of the appearance of the singing technique reaches a certain value. For example, in one embodiment, obtaining the grace note melody in the standard melody information and adding the grace note identification information to the lyrics corresponding to the grace note melody in the standard lyrics information to obtain the lyrics template includes: obtaining the number of characters of the lyrics corresponding to the grace note melody in the standard melody information; obtaining the ratio of the number of characters of the lyrics corresponding to the grace note melody to the total number of lyrics in the standard lyrics information. If the ratio is greater than the preset grace note threshold, obtaining the grace note melody in the standard melody information and adding the grace note identification information to the lyrics corresponding to the grace note melody in the standard lyrics information to obtain the lyrics template.

[0097] In this embodiment, the server 104 can first obtain the total number of lyrics in the standard lyric information in the standard audio as the first word count, and obtain the word count of the lyrics corresponding to the ornamented melody in the standard melody information. Thus, the server 104 can obtain the ratio of the word count of the lyrics corresponding to the ornamented melody to the total number of lyrics in the standard lyric information. If the server 104 detects that this ratio is greater than the preset ornamented tone threshold, the server 104 can obtain the ornamented melody in the standard melody information and add ornamented tone identification information to the lyrics corresponding to the ornamented melody in the standard lyric information, thereby obtaining a lyric template. That is, the server 104 can add the ornamented tone identification information only when it detects that the probability of the appearance of the ornamented tone in the standard melody information is greater than a certain value. Specifically, taking the above standard audio as the audio of the original singer's song as an example, the server 104 can obtain the usage frequency of the singing skills in the original singer's song. For example, the server 104 can calculate the average value of the word counts of the ornamented tones such as vibrato, portamento, and melisma used in the original singer's song relative to the total word count of the standard lyric information. The calculation formula is as follows: Then when the server 104 detects it can be determined that there is a high probability that the user will have ornamented tones when singing this song. Therefore, when making the melody template, it is necessary to add identification information of singing skills, including melisma marks, vibrato marks, portamento marks, etc. The server 104 can perform word-by-word audio adjustment based on the audio sequence. Therefore, when calculating the pitch shift sequence for tuning during the calculation strategy processing stage, the server 104 can pay attention to whether there are marks of singing skills such as ornamented tone identification information for the current word, and thus perform corresponding adjustments. Among them, the above ornamented tone identification information can be added to the melody template or the lyric template, so that when the server 104 performs word-by-word audio adjustment, it can determine the adjustment strategy according to whether each word in the lyrics has ornamented tone identification information.

[0098] Through the above embodiment, the server 104 can determine whether to add ornamented tone identification information based on the probability of the appearance of the ornamented tone in the standard audio, and can perform corresponding adjustment on the audio input by the user based on the ornamented tone identification information during the audio adjustment stage, improving the adjustment effect of the audio adjustment.

[0099] In one embodiment, before matching the lyric template with the audio to be adjusted, it further includes: obtaining the number of long vowel notes with pitch change and pronunciation length greater than the preset length threshold in the standard melody information; obtaining the proportion of the long vowel notes in the standard melody information according to the number of long vowel notes; if the proportion is greater than the preset long vowel probability threshold, obtaining the extended vowel corresponding to the long vowel note; expanding the phoneme of the long vowel note in the standard lyric information according to the pronunciation length of the long vowel note and the extended vowel to obtain the extended lyric information corresponding to the standard lyric information; matching the extended lyric information with the audio to be adjusted.

[0100] In this embodiment, before the server 104 can match the audio to be adjusted with the standard lyrics information, it can identify the long vowels with pitch changes in the standard audio. Long vowels are drawn-out sounds, which mostly appear at the end of each sentence when singing. For example, it can identify long vowels such as the character "ah" with multiple pitch changes. The occurrence of long vowels will affect the alignment accuracy between the human voice and the lyrics. When the probability of the occurrence of long vowels with pitch changes is greater than a certain value, the server 104 can perform phoneme expansion on the long vowels. For example, the server 104 can obtain the number of long vowel notes with pitch changes and a pronunciation length greater than a preset length threshold in the standard melody information, and obtain the proportion of long vowel notes in the standard melody information based on the number of long vowel notes. Among them, this proportion can be obtained by calculating the ratio. For example, in one embodiment, to obtain the proportion of long vowel notes in the standard melody information according to the number of long vowel notes, it includes: obtaining the number of lines of lyrics in the standard lyrics information, obtaining the ratio of the number of long vowel notes to the number of lines of lyrics in the standard lyrics information, and using it as the proportion of long vowel notes in the standard melody information. In this embodiment, the server 104 can calculate the proportion of long vowel notes in the standard melody information. For example, the server 104 can obtain the number of lines of lyrics in the standard lyrics information, and obtain the ratio of the number of the above-mentioned long vowel notes to the number of lines of lyrics in the above-mentioned standard lyrics information, and use it as the proportion of long vowel notes in the standard melody information.

[0101] After the server 104 obtains the proportion of long vowel notes, if the server 104 detects that this proportion is greater than the preset long vowel probability threshold, the server 104 can obtain the extended vowel corresponding to the long vowel note. Among them, the extended vowel can be obtained based on the pronunciation dictionary. The server 104 can query the pronunciation dictionary based on the long vowel note to obtain the extended vowel corresponding to the long vowel note. For example, the last pronunciation syllable in the long vowel note is used as the extended vowel. Thus, the server 104 can expand the phoneme of the long vowel note in the standard lyrics information based on the obtained extended vowel to obtain the extended lyrics information corresponding to the standard lyrics information. The server 104 can match the extended lyrics information with the audio to be adjusted to obtain the alignment result. Specifically, as Figure 5 shown Figure 5It is a schematic flow diagram of the step of expanding vowels in an embodiment. Taking the recorded voice input by the user as a singing recording as an example, the standard audio is the original singing audio of the singing recording. The server 104 can expand the vowels of the standard lyrics file based on the probability p of the long vowels with tone change in the standard audio and the pronunciation dictionary. The server 104 can first calculate the probability of the occurrence of long vowels in the standard audio, and its calculation formula can be as follows: P = the number of long vowels with tone change / the number of lyrics lines. Among them, the number of long vowels with tone change can be the number of long vowels with tone change in the standard audio, and the number of lyrics lines can be the number of lines of the standard lyrics. Specifically, for the lyrics in lrc format, it has information such as the time tags of each line of lyrics. Therefore, the server 104 can determine the number of lyrics lines based on this information. If P > 0.1, it means that there is a high probability of the occurrence of long vowels with tone change when the user sings this song. Then, when the server 104 makes the pronunciation dictionary, it will expand the vowel phonemes based on the long vowel part. For example, when the long vowel "ha" appears during singing, its corresponding phoneme sequence is expanded from the original Ha1 to Ha1a1a1, and the number of the a1 vowel phonemes will be expanded according to the length of the long vowel, and the number of vowel phonemes is in a direct proportion relationship with the length of the long vowel. Among them, the pronunciation dictionary contains the set of words that the alignment system needs to process and indicates their pronunciations. The server 104 can obtain the mapping relationship between the acoustic model modeling unit and the language model modeling unit through the pronunciation dictionary, so as to form a search state space for decoding. For example: for the character "ha" -> the pinyin "Ha" -> the phoneme sequence "Ha1a1..." (where 1 represents the tone in Chinese, and the first tone is 1).

[0102] Through the above embodiment, the server 104 can detect the probability of the occurrence of long vowels in the standard audio, and when the probability of the occurrence of long vowels is relatively large, expand the phoneme sequence of the long vowels, so that when matching the standard lyrics information with the audio to be adjusted by the user, the alignment accuracy of the long vowels can be improved, and thus the adjustment effect of the audio adjustment can be improved.

[0103] In an embodiment, matching the lyrics template with the audio to be adjusted to obtain the audio sequence to be adjusted includes: detecting the fundamental frequency of the audio to be adjusted to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and converting the fundamental frequency sequence into a sequence of note-like units; aligning each note-like unit in the sequence of note-like units with each lyric in the lyrics template to obtain the audio sequence to be adjusted after word-by-word mapping.

[0104] In this embodiment, the server 104 can match the audio to be adjusted with the standard lyrics information. Among them, the server 104 can perform conversion processing on the audio to be adjusted before matching and aligning it with the corresponding lyrics template of the standard lyrics information. For example, the server 104 can perform fundamental frequency detection on the audio to be adjusted to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and convert the fundamental frequency sequence into a sequence of note-like units. The fundamental frequency extraction technology refers to the ability to extract the fundamental frequency (F0) curve of the human voice in the user's dry voice. The sequence of note-like units includes multiple note-like units NLU. The server 104 can align each note-like unit in the sequence of note-like units with each lyric in the above-mentioned lyrics template to obtain the audio sequence to be adjusted after word-by-word mapping. Specifically, as Figure 6 shown Figure 6 Figure 4 is a schematic flowchart of the alignment step in an embodiment. The server 104 can first perform fundamental frequency detection on the audio input by the user. The fundamental frequency is generated by the vibration of the vocal cords. Generally, voiced sounds have a fundamental frequency. The server 104 can perform periodic analysis on the audio signal of the voiced segment in the audio to be adjusted to obtain the fundamental frequency sequence. After the server 104 obtains the fundamental frequency sequence, it can convert it into a Note sequence, that is, a sequence of note-like units, through a set formula, so as to facilitate subsequent comparison with the template. The server 104 can obtain a logarithmic function according to the ratio of the fundamental frequency sequence to the first value; obtain the product of the logarithmic function and the second value, and obtain the sum of the product and the third value to obtain the sequence of note-like units; among them, the first value, the second value, and the third value are all different. Specifically, its calculation formula is as follows: Note = 12 * log2(frequency / 440) + 69. Among them, Note can be a sequence of note-like units, and frequency can be the fundamental frequency sequence. Among them, the above-mentioned fundamental frequency sequence can also include the extended vowel content after the server 104 performs phoneme extension on long vowels, so that the converted sequence of note-like units can also include the extended vowel content, which is convenient for alignment. Another example is Figure 7 shown Figure 7 Figure 5 is a schematic flowchart of the alignment step in another embodiment. Taking the audio recorded by the user as a song as an example, after the server 104 obtains the above-mentioned sequence of note-like units, it can compare and align the audio to be adjusted input by the user with the lyrics template to obtain a word-by-word mapping relationship as the audio sequence to be adjusted.

[0105] Through this embodiment, the server 104 can convert the audio to be adjusted into a sequence, thereby aligning based on the sequence and the lyrics template, and performing audio adjustment based on the obtained audio sequence to be adjusted after word-by-word mapping, thereby improving the effect of audio adjustment.

[0106] In one embodiment, based on a melody template, pitch adjustment is performed on multiple note-like units in an audio sequence to be adjusted to obtain an adjusted audio sequence, including: for each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit contains ornamentation identification information, pitch translation is performed on the notes in the note-like unit to conform to the melody template; if it is detected that the note-like unit contains replacement intervals, each replacement interval of the note-like unit is traversed, and the note corresponding to the interval in the note-like unit is replaced with the replacement note in the replacement interval with the smallest interval difference degree from the interval of the note-like unit; according to the pitch adjustment results of multiple note-like units, an adjusted audio sequence is obtained.

[0107] In this embodiment, during the audio processing stage, the server 104 has added ornamentation identification information and replacement intervals of difficult intervals to the audio to be adjusted. The server 104 can convert the audio to be adjusted into an audio sequence to be adjusted, and the audio sequence to be adjusted includes multiple note-like units. For each note-like unit in the audio sequence to be adjusted, the server 104 can detect whether the note-like unit contains ornamentation identification information and replacement intervals. If the server 104 detects that the note-like unit includes ornamentation identification information, the server 104 can perform pitch translation on the notes in the note-like unit based on the melody information in the melody template, so that the waveform in the note-like unit conforms to the waveform at the corresponding position in the melody template. If the server 104 detects that the note-like unit contains replacement intervals, since there are multiple replacement intervals, the server 104 can traverse each replacement interval of the note-like unit and replace the note corresponding to the interval in the note-like unit with the replacement note in the replacement interval with the smallest frequency shift from the note-like unit, that is, the smallest interval difference degree.

[0108] In addition, in one embodiment, each note-like unit in the above audio sequence to be adjusted may not include ornamentation identification information and replacement intervals. When the server 104 detects that the note-like unit does not contain ornamentation identification information and replacement intervals, pitch translation and amplitude compression processing can be performed on the notes in the note-like unit to conform to the melody template in the adjustment template. Thus, the server 104 can obtain an adjusted audio sequence according to the pitch adjustment results of multiple note-like units in the audio sequence to be adjusted.

[0109] Specifically, as Figure 8 shown, Figure 8 is a schematic diagram of the interface for the audio adjustment steps in one embodiment. The note-like unit NLU is Figure 8The horizontal line in the figure shows that the note-like unit is a short-term stable note value obtained by smoothing the waveform of the fundamental frequency sequence. The server 104 can make two adjustments to the note-like unit, including: shifting the NLU (Note-Like Unit) of the user's fundamental frequency, and adjusting the dynamic range of the fundamental frequency jitter in the NLU. Among them, pitch shifting refers to the operation of returning the pitch value that deviates from the template to the standard value through the overall up-down operation; dynamic range adjustment refers to controlling the jitter amplitude of the fundamental frequency sequence in a single NLU, for example Figure 8 Rectangle 800 and rectangle 802 in the image. Among them, the degree of jitter in rectangle 800 is relatively large, while the amplitude of the fundamental frequency jitter in rectangle 802 is relatively stable. Server 104 can perform the above-mentioned translation and compression operations on the note-like units that do not contain the ornament identification information. It should be noted that, of course, if the fundamental frequency sequence has almost no jitter in the current NLU, it will produce a mechanical sense in the hearing sense. Therefore, it is not advisable to jitter too violently or to fix the pitch unchanged. Server 104 can make the fundamental frequency sequence close to the melody template.

[0110] In addition, if Figure 9 As shown, Figure 9 1 is a flow chart of an audio adjustment step in an embodiment. Taking the audio input by the user as a singing recording as an example, the standard audio of the singing recording is the original singing audio. The server 104 can perform a sequence-based word-by-word match based on the template melody and template lyrics of the original singing audio, and the dry voice melody and dry voice word-by-word information input by the user, to obtain the corresponding audio sequence to be adjusted, thereby performing word-by-word matching, and determining the corresponding sound modification strategy based on the identification of whether each word has corresponding difficulty information, and obtaining the final frequency shift sequence. For example, when the server 104 detects that there are multiple replacement intervals in the class note unit during the word-by-word matching process, the server 104 can traverse all replacement intervals, and select the target replacement interval with the least frequency shift degree for use, and replace it with the original difficult example interval. When the server 104 detects that there is an ornamental sound identification information in the class note unit, the user's audio sequence to be adjusted can be skill-detected. If no singing skills are detected, no special processing is performed; if singing skills such as ornamental sounds are detected, the server 104 can only perform a translation process on the user's pitch without compression processing, thereby realizing audio adjustment of the audio sequence to be adjusted.

[0111] Through the above-mentioned embodiment, the server 104 can adopt different audio adjustment methods for different note-like units based on the difficulty information identifier in the audio sequence to be adjusted, thereby improving the adjustment effect of the audio adjustment.

[0112] In one embodiment, Figure 10 Show, Figure 10Schematic flowchart of an audio adjustment method in another embodiment. This method can be applied to pitch correction of singing audio. The server 104 may include a pitch correction engine. The user can obtain the audio to be adjusted input by the user as the user's dry voice, and obtain the corresponding original singing audio based on the user's dry voice. After performing difficulty analysis on the original singing audio, such as lyric word frequency, song melody difficulty, and song pitch range span, etc., the server 104 can obtain the corresponding template melody and template lyrics after marking the corresponding difficulty information. The server 104 can also extract features from the user's dry voice, such as fundamental frequency extraction and obtaining the sequence of word-by-word mapping, etc. The server 104 can determine the pitch correction strategy for the audio sequence corresponding to the user's dry voice based on the audio sequence containing the features of the user's dry voice and the above-mentioned recognition results such as singing skills and pitch range. After the server 104 performs targeted pitch correction on the audio sequence, it can perform corresponding pitch shift processing on the user's original dry voice based on the pitch-corrected audio sequence to obtain the pitch-corrected audio.

[0113] Through the above embodiments, the server 104 can, based on the original singing of the song, add corresponding replacement intervals and singing skill annotation information on the basis of the melody template, avoid unnatural pitch correction effects caused by too many difficult intervals, and through analyzing the difficulty information of the standard audio, perform targeted adjustment on the audio to be adjusted according to the difficulty information of the standard audio, improving the adjustment effect of audio adjustment.

[0114] It should be understood that although the steps in the flowcharts of the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0115] Based on the same inventive concept, the embodiments of the present application also provide an audio adjustment device for implementing the above-mentioned audio adjustment method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the audio adjustment device can refer to the limitations on the audio adjustment method in the above text and will not be repeated here.

[0116] In one embodiment, as Figure 11As shown, an audio adjustment device is provided, including: an acquisition module 500, a determination module 502, an alignment module 504, and an adjustment module 506, where:

[0117] The acquisition module 500 is configured to acquire the audio to be adjusted and its corresponding standard audio, and acquire the difficulty information of the standard audio. The difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyric information.

[0118] The determination module 502 is configured to determine an adjustment template based on the standard melody information, the standard lyric information, and the difficulty information of the standard audio; the adjustment template includes a melody template and a lyric template; at least one of the melody template and the lyric template includes identification information for identifying the difficulty information.

[0119] The alignment module 504 is configured to match the lyric template with the audio to be adjusted to obtain an audio sequence to be adjusted; the audio sequence to be adjusted includes class-note units corresponding to each lyric in the lyric template.

[0120] The adjustment module 506 is configured to perform pitch adjustment on multiple class-note units in the audio sequence to be adjusted based on the melody template to obtain an adjusted audio sequence; and adjust the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio.

[0121] In one embodiment, the determination module 502 is specifically configured to obtain a melody template based on the standard melody information; the target interval in the melody template is marked with a replacement interval; the target interval is determined based on the first difficulty information and represents the difficult example interval in the standard melody information; obtain a lyric template based on the standard lyric information; the target lyric information in the lyric template is marked with ornamentation identification information; the target lyric information is determined based on the second difficulty information and represents the lyric corresponding to the ornamentation melody in the standard lyric information; and determine an adjustment template according to the melody template and the lyric template.

[0122] In one embodiment, the determination module 502 is specifically configured to obtain the difficult example interval in the standard melody information according to the pitch difference between adjacent notes in the standard melody information; obtain a replacement interval corresponding to the difficult example interval; the interval difference between the replacement interval and the difficult example interval is less than a preset interval difference threshold, and the tones of the adjacent notes in the difficult example interval are the same; mark the replacement interval at the difficult example interval in the standard melody information to obtain a melody template.

[0123] In one embodiment, the determination module 502 is specifically configured to obtain the ornamentation melody in the standard melody information, and add ornamentation identification information to the lyric corresponding to the ornamentation melody in the standard lyric information to obtain a lyric template.

[0124] In one embodiment, the determining module 502 is specifically configured to determine that adjacent notes in the standard melody information are difficult example intervals if the magnitude of the pitch difference is greater than or equal to a preset pitch difference threshold.

[0125] In one embodiment, the determining module 502 is specifically configured to obtain the number of characters of the lyrics corresponding to the ornamented melody in the standard melody information; obtain the ratio of the number of characters of the lyrics corresponding to the ornamented melody to the total number of lyrics in the standard lyric information; if the ratio is greater than a preset ornament threshold, add ornament identification information to the lyrics corresponding to the ornamented melody in the standard lyric information to obtain a lyric template.

[0126] In one embodiment, the apparatus further includes: an expansion module, configured to obtain the number of long vowel notes in the standard melody information that have a key change and a pronunciation length greater than a preset length threshold; obtain the proportion of the long vowel notes in the standard melody information according to the number of the long vowel notes; if the proportion is greater than a preset long vowel probability threshold, obtain the extended vowel corresponding to the long vowel note; expand the phoneme of the long vowel note in the lyric template according to the pronunciation length of the long vowel note and the extended vowel to obtain the extended lyric information corresponding to the lyric template; match the extended lyric information with the audio to be adjusted.

[0127] In one embodiment, the expansion module is specifically configured to obtain the number of lines of lyrics in the standard lyric information; determine the proportion of the long vowel notes in the standard melody information based on the ratio of the number of the long vowel notes to the number of lines of lyrics.

[0128] In one embodiment, the alignment module 504 is specifically configured to perform fundamental frequency detection on the audio to be adjusted to obtain a fundamental frequency sequence corresponding to the audio to be adjusted, and convert the fundamental frequency sequence into a sequence of note-like units; align each note-like unit in the sequence of note-like units with each lyric in the lyric template to obtain the audio sequence to be adjusted after word-by-word mapping.

[0129] In one embodiment, the adjustment module 506 is specifically configured to, for each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit contains ornament identification information, perform pitch translation on the notes in the note-like unit to fit the melody template; if it is detected that the note-like unit contains replacement intervals, traverse each replacement interval in the note-like unit, and replace the note corresponding to the interval in the note-like unit with the replacement note in the replacement interval with the smallest interval difference from the interval in the note-like unit; obtain the adjusted audio sequence according to the pitch adjustment results of multiple note-like units.

[0130] In one embodiment, the above adjustment module 506 is specifically configured to, for each class note unit in the audio sequence to be adjusted, if it is detected that the class note unit does not contain grace note identification information and replacement intervals, perform pitch translation and amplitude compression processing on the notes in the class note unit to fit the melody template in the adjustment template; and obtain the adjusted audio sequence according to the pitch adjustment results of multiple class note units.

[0131] In one embodiment, the above adjustment module 506 is specifically configured to perform pitch shifting processing on the audio to be adjusted according to the adjusted audio sequence to obtain the adjusted audio.

[0132] Each module in the above audio adjustment device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0133] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio adjustment method.

[0134] Those skilled in the art can understand that Figure 12 the structure shown in

[0135] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0136] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the above audio adjustment method.

[0137] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the above audio adjustment method.

[0138] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.

[0139] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in this application can include at least one of a relational database and a non-relational database. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0141] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An audio adjustment method, characterized in that, The method includes: Obtaining the audio to be adjusted and its corresponding standard audio, and obtaining the difficulty information of the standard audio, where the difficulty information includes at least one of the first difficulty information of the standard melody information of the standard audio and the second difficulty information of the standard lyric information; Determining an adjustment template based on the standard melody information, the standard lyric information, and the difficulty information of the standard audio; the adjustment template includes a melody template and a lyric template; at least one of the melody template and the lyric template includes identification information for identifying the difficulty information; Matching the lyric template with the audio to be adjusted to obtain an audio sequence to be adjusted; the audio sequence to be adjusted includes a class note unit corresponding to each lyric in the lyric template; the class note unit is composed of a section of melody information corresponding to each lyric in the standard melody information of the standard audio, and the melody information is waveform information; Based on the melody template, performing pitch adjustment on multiple class note units in the audio sequence to be adjusted to obtain an adjusted audio sequence; adjusting the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio.

2. The method according to claim 1, wherein The determining the adjustment template based on the standard melody information, the standard lyric information, and the difficulty information includes: Obtaining a melody template based on the standard melody information; a target interval in the melody template is marked with a replacement interval; the target interval is determined based on the first difficulty information and represents a difficult example interval in the standard melody information; Obtaining a lyric template based on the standard lyric information; target lyric information in the lyric template is marked with ornamentation identification information; the target lyric information is determined based on the second difficulty information and represents the lyric corresponding to the ornamentation melody in the standard lyric information; Determining an adjustment template according to the melody template and the lyric template.

3. The method according to claim 2, wherein The obtaining the melody template based on the standard melody information includes: Obtaining the difficult example intervals in the standard melody information according to the pitch difference between adjacent notes in the standard melody information; Obtaining a replacement interval corresponding to the difficult example interval; the interval difference between the replacement interval and the difficult example interval is less than a preset interval difference threshold, and the tones of adjacent notes in the difficult example interval are the same; Marking the replacement interval at the difficult example interval in the standard melody information to obtain a melody template.

4. The method according to claim 2, wherein The obtaining the lyric template based on the standard lyric information includes: Obtaining the ornamentation melody in the standard melody information; Adding ornamentation identification information to the lyric corresponding to the ornamentation melody in the standard lyric information to obtain a lyric template.

5. The method according to claim 3, wherein The obtaining the difficult example intervals in the standard melody information according to the pitch difference between adjacent notes in the standard melody information includes: If the magnitude of the pitch difference is greater than or equal to a preset pitch difference threshold, determining that the adjacent notes in the standard melody information are difficult example intervals.

6. The method according to claim 4, characterized in that The adding ornamentation identification information to the lyric corresponding to the ornamentation melody in the standard lyric information to obtain a lyric template includes: Obtaining the number of characters of the lyric corresponding to the ornamentation melody in the standard melody information; Obtain the ratio of the number of words in the lyrics corresponding to the grace note melody to the total number of lyrics in the standard lyrics information; If the ratio is greater than a preset grace note threshold, add grace note identification information to the lyrics corresponding to the grace note melody in the standard lyrics information to obtain a lyrics template.

7. The method according to claim 1, characterized in that Before matching the lyrics template with the audio to be adjusted, it further includes: Obtain the number of long vowel notes with pitch changes and pronunciation lengths greater than a preset length threshold in the standard melody information; According to the number of long vowel notes, obtain the proportion of long vowel notes in the standard melody information; If the proportion is greater than a preset long vowel probability threshold, obtain the extended vowel corresponding to the long vowel note; According to the pronunciation length of the long vowel note and the extended vowel, expand the phoneme of the long vowel note in the lyrics template to obtain the extended lyrics information corresponding to the lyrics template; Perform matching based on the extended lyrics information and the audio to be adjusted.

8. The method according to claim 7, characterized in that, The obtaining the proportion of long vowel notes in the standard melody information according to the number of long vowel notes includes: Obtain the number of lines of lyrics in the standard lyrics information; Based on the ratio of the number of long vowel notes to the number of lines of lyrics, determine the proportion of long vowel notes in the standard melody information.

9. The method according to claim 1, wherein The matching the lyrics template with the audio to be adjusted to obtain an audio sequence to be adjusted includes: Perform fundamental frequency detection on the audio to be adjusted to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and convert the fundamental frequency sequence into a sequence of note-like units; Align each note-like unit in the sequence of note-like units with each lyric in the lyrics template to obtain the audio sequence to be adjusted after word-by-word mapping.

10. The method according to claim 2, wherein The performing pitch adjustment on multiple note-like units in the audio sequence to be adjusted based on the melody template to obtain an adjusted audio sequence includes: For each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit contains grace note identification information, perform pitch translation on the notes in the note-like unit to fit the melody template; If it is detected that the note-like unit contains replacement intervals, traverse each replacement interval of the note-like unit, and replace the note corresponding to the interval in the note-like unit with the replacement note in the replacement interval with the smallest interval difference from the interval of the note-like unit; According to the pitch adjustment results of multiple note-like units, obtain the adjusted audio sequence.

11. The method according to claim 2, characterized in that The performing pitch adjustment on multiple note-like units in the audio sequence to be adjusted based on the melody template to obtain an adjusted audio sequence includes: For each note-like unit in the audio sequence to be adjusted, if it is detected that the note-like unit does not contain grace note identification information and replacement intervals, perform pitch translation and amplitude compression processing on the notes in the note-like unit to fit the melody template in the adjustment template; According to the pitch adjustment results of multiple note-like units, obtain the adjusted audio sequence.

12. The method according to claim 1, characterized in that, The adjusting the audio to be adjusted based on the adjusted audio sequence to obtain an adjusted audio includes: Perform pitch shifting processing on the audio to be adjusted according to the adjusted audio sequence to obtain the adjusted audio.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 12.

14. A computer program product, including a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Audio adjustment method, computer device and computer program product

    CN114743526A