Audio adjustment method, computer device and program product

By building audio templates and automatically processing audio parameters, the problem of low efficiency in ghost audio adjustment is solved, and the automation and high efficiency of audio adjustment are achieved.

CN115472146BActive Publication Date: 2025-09-12TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211090755.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-07
Publication Date
2025-09-12
Estimated Expiration
2042-09-07

AI Technical Summary

Technical Problem

The existing technology has low efficiency in ghost audio adjustment and mainly relies on manual adjustment of audio parameters, resulting in low efficiency.

Method used

By constructing an audio template containing tone, beat, text and note distribution, identifying the text and beat identifiers in the dry audio, establishing a correspondence between text words and note names, adjusting the audio part according to the pitch value, fusing the dry audio with the template accompaniment, and generating the adjusted audio.

Benefits of technology

Improved the efficiency of ghost audio adjustment, replaced manual adjustment with automated processing, and improved the speed and efficiency of audio adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472146B_ABST
    Figure CN115472146B_ABST
Patent Text Reader

Abstract

The present application relates to an audio adjustment method, computer equipment and computer program product. After selecting an audio template, the corresponding dry audio is recorded, the note names contained between multiple beat identifiers are detected to determine the note name of each beat, and the corresponding relationship between each text word of the audio template and the note name of each beat is established according to the beat information of the multiple note names, and the pitch value of each text word is determined. According to the pitch value, the audio part corresponding to each text word in the dry audio and each text word in the audio template is adjusted to obtain the adjusted dry audio, and the adjusted dry audio and template accompaniment are fused based on the beat information to obtain the adjusted audio. Compared with the traditional method of manually adjusting audio parameters, this solution parses the template through a specific strategy to obtain parameters for audio adjustment, adjusts the audio based on the audio template, and improves the adjustment efficiency when adjusting ghost audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to an audio adjustment method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0002] With the development of computer technology, it is now possible to listen to and process audio through various terminal devices, such as listening to and processing songs. With the development of audio editing technology, ghost audio and video have gradually become popular. Ghost audio refers to the use of sounds in the material to edit, splice, and tune, and combine it with the accompaniment of the song to obtain a complete audio and video work. Therefore, when users need to obtain ghost audio, they need to adjust the audio. The current method of adjusting audio and generating ghost audio is usually to manually adjust various parameters of the audio to obtain an adjusted ghost audio. However, the efficiency of audio adjustment through manual adjustment is not high.

[0003] Therefore, the current audio adjustment method has the defect of low adjustment efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide an audio adjustment method, device, computer equipment, computer-readable storage medium and computer program product that can improve the adjustment efficiency of ghost audio in response to the above technical problems.

[0005] In a first aspect, the present application provides an audio adjustment method, the method comprising:

[0006] Selecting an audio template, recording dry audio corresponding to the audio template and identifying text and text words in the dry audio; the audio template includes a template tone, text and text words of the audio template, and a template note distribution sequence; the template note distribution sequence includes multiple note names and multiple beat identifiers;

[0007] Detecting the note names contained between the multiple beat identifiers, determining the note name of each beat, and obtaining beat information of the multiple note names;

[0008] According to the beat information, establishing a correspondence between each text word in the audio template and the note name of each beat, to obtain the note name corresponding to each text word in the audio template;

[0009] determining a pitch value of each text word according to a note name corresponding to each text word in the audio template, and adjusting an audio portion corresponding to each text word in the dry audio and the audio template according to the pitch value to obtain an adjusted dry audio;

[0010] A template accompaniment corresponding to the audio template is obtained, and a fusion process is performed on the adjusted dry audio and the template accompaniment based on the beat information to obtain an adjusted audio.

[0011] In one embodiment, selecting an audio template and recording dry audio corresponding to the template includes:

[0012] Displaying at least one audio template to be selected;

[0013] receiving a selection instruction for the at least one audio template to be selected, and determining a selected audio template;

[0014] The text corresponding to the audio template is displayed, and the dry audio input by the user based on the text corresponding to the audio template is recorded.

[0015] In one embodiment, the identifying text and text words in the dry audio includes:

[0016] Acquiring fundamental frequency information of the dry audio, and identifying original text corresponding to the dry audio according to the fundamental frequency information;

[0017] Based on the matching result between the original text and the text of the audio template, the original text is modified to obtain the text of the dry audio, and each text word in the text of the dry audio is obtained; each text word in the text of the dry audio is matched with each text word in the text of the audio template.

[0018] In one embodiment, the method further comprises:

[0019] Obtain the original template audio and its corresponding template accompaniment, beat information and text of the original template audio;

[0020] Acquire a plurality of original template notes in the original template audio, and determine a plurality of template notes corresponding to the number of text words in the text of the original template audio from the plurality of original template notes;

[0021] Adding beat identifiers between the multiple template notes according to the multiple template notes and the beat information to obtain a template note distribution sequence;

[0022] According to the text of the original template audio and the template note distribution sequence, an audio template corresponding to the original template audio is generated, and the audio template and the corresponding template accompaniment are stored in a template library.

[0023] In one embodiment, detecting the note names included between the multiple beat identifiers, determining the note name of each beat, and obtaining the beat information of the multiple note names includes:

[0024] cyclically detecting each character in the template note distribution sequence, and if the character is detected to be a numeric character, determining that the character is a note name;

[0025] If it is detected that the character is a beat mark, and the character between the beat mark and the previous beat mark is a note name, the note name between the beat mark and the previous beat mark is regarded as the note name of a beat;

[0026] According to the multiple note names of one beat, beat information of the multiple note names is obtained.

[0027] In one embodiment, the audio template further includes: a template tone; determining a pitch value of each text word according to a note name corresponding to each text word in the audio template, and adjusting an audio portion corresponding to each text word in the dry audio and each text word in the audio template according to the pitch value to obtain adjusted dry audio, including:

[0028] Obtaining a reference pitch value of the template tone, and determining a pitch value of a note name corresponding to each text word in the audio template based on the reference pitch value; and determining a frequency of the note name corresponding to each text word in the audio template based on the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone;

[0029] According to the frequency of the note name corresponding to each text word in the audio template, the audio part of the dry audio corresponding to each text word in the audio template is adjusted to obtain the adjusted dry audio.

[0030] In one embodiment, obtaining a reference pitch value of the template tone and determining the pitch value of the note name corresponding to each text word in the audio template according to the reference pitch value includes:

[0031] According to the template tone, a preset tone pitch mapping table is searched to determine a reference pitch value corresponding to the template tone; the preset tone pitch mapping table includes a correspondence between multiple tones and reference pitch values;

[0032] For each note name corresponding to a text word in the audio template, a preset note table is searched to determine a first semitone interval value between the note name corresponding to the text word and the first note in the octave space where the note name is located; the preset note table is obtained based on multiple note names in the same octave space and the first semitone interval value between each note name and the first note in the octave space;

[0033] The pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value and the first semitone interval value corresponding to the note name.

[0034] In one embodiment, determining the pitch value of the note name corresponding to each text word in the audio template according to the reference pitch value further includes:

[0035] For each text word in the audio template, detecting an identifier of a note name corresponding to the text word;

[0036] If it is detected that the note name corresponding to the text word has a semitone offset identifier, a semitone offset table is searched according to the semitone offset identifier to determine the second semitone interval value corresponding to the semitone offset identifier; the pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value, the first semitone interval value corresponding to the note name, and the second semitone interval value; the semitone offset table includes a correspondence between multiple semitone offset identifiers and second semitone interval values; and / or

[0037] If it is detected that the note name corresponding to the text word has an octave offset identifier, the octave offset table is queried according to the octave offset identifier, the octave offset value corresponding to the octave offset identifier is determined, and the third semitone interval value corresponding to the octave offset value is obtained; the pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value, the first semitone interval value corresponding to the note name corresponding to the text word, and the third semitone interval value; the octave offset table includes a plurality of correspondences between octave offset identifiers and octave offset values.

[0038] In one embodiment, determining the frequency of the note name corresponding to each text word in the audio template according to the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone includes:

[0039] According to the template tone, a preset frequency table is searched to obtain a frequency corresponding to the template tone; the preset frequency table includes a correspondence between a plurality of template tones and frequencies;

[0040] For each note name corresponding to each text word in the audio template, obtaining a difference between a pitch value of the note name corresponding to the text word and a reference pitch value corresponding to the template tone, and determining a frequency adjustment value corresponding to the note name based on the difference;

[0041] The frequency corresponding to the note name corresponding to the text word is determined according to the product of the frequency corresponding to the template tone and the frequency adjustment value.

[0042] In one embodiment, the audio template further includes: a template beat; and the fusion processing of the adjusted dry audio and the template accompaniment to obtain the adjusted audio includes:

[0043] Determining, based on the beat information, a duration of an audio portion corresponding to each text word in the adjusted dry audio, and adjusting the duration of the audio portion corresponding to each text word in the adjusted dry audio according to the duration of the audio portion corresponding to each text word, so that the duration of the audio portion corresponding to each text word in the adjusted dry audio matches the duration of each beat in the template beat, thereby obtaining dry audio with adjusted duration;

[0044] The dry audio with the adjusted duration is mixed with the template accompaniment to obtain the adjusted audio.

[0045] In a second aspect, the present application provides an audio adjustment device, the device comprising:

[0046] A recording module is configured to select an audio template, record dry audio corresponding to the audio template, and identify text and text words in the dry audio; the audio template includes a template tone, the text and text words of the audio template, and a template note distribution sequence; the template note distribution sequence includes multiple note names and multiple beat identifiers;

[0047] a detection module, configured to detect the note names contained between the plurality of beat identifiers, determine the note name of each beat, and obtain beat information of the plurality of note names;

[0048] a corresponding module, configured to establish a corresponding relationship between each text word in the audio template and the note name of each beat according to the beat information, so as to obtain the note name corresponding to each text word in the audio template;

[0049] an adjustment module, configured to determine a pitch value of each text word according to a note name corresponding to each text word in the audio template, and adjust an audio portion corresponding to each text word in the dry audio according to the pitch value to obtain adjusted dry audio;

[0050] A fusion module is used to obtain a template accompaniment corresponding to the audio template, and perform a fusion process on the adjusted dry audio and the template accompaniment based on the beat information to obtain an adjusted audio.

[0051] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0052] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0053] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0054] The above-mentioned audio adjustment method, device, computer equipment, storage medium and computer program product, after selecting an audio template, records the corresponding dry audio and recognizes the text therein, determines the note name of each beat by detecting the note names contained between multiple beat identifiers, establishes a correspondence between each text word of the audio template and the note name of each beat according to the beat information of the multiple note names, determines the pitch value of each text word according to the note name corresponding to each text word, adjusts each text word in the dry audio and the audio part corresponding to each text word in the audio template according to the pitch value, obtains the adjusted dry audio, and fuses the adjusted dry audio and the template accompaniment based on the beat information to obtain the adjusted audio. Compared with the traditional method of manually adjusting audio parameters, this solution constructs an audio template containing pitch, beat, text and note distribution, and parses the template through a specific strategy to obtain parameters for audio adjustment, adjusts the audio based on the audio template, and improves the adjustment efficiency when adjusting ghost audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 1 is a flow chart of an audio adjustment method according to an embodiment;

[0056] Figure 2 A schematic flow chart of a step of obtaining a template note distribution in one embodiment;

[0057] Figure 3 is a schematic diagram of the distribution of template notes in one embodiment;

[0058] Figure 4 is a schematic diagram of pitch distribution after adjustment in one embodiment;

[0059] Figure 5 is a schematic diagram of the frequency distribution after adjustment in one embodiment;

[0060] Figure 6 is a flowchart of an audio adjustment method according to another embodiment;

[0061] Figure 7 is a structural block diagram of an audio adjustment device in one embodiment;

[0062] Figure 8 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0064] In one embodiment, Figure 1 As shown, an audio adjustment method is provided. This embodiment uses the method applied to a terminal as an example for explanation. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through interaction between the terminal and the server, including the following steps:

[0065] Step S202, select an audio template, record the dry audio corresponding to the audio template and identify the text and text words in the dry audio; the audio template includes the template tone, the text and text words of the audio template, and the template note distribution sequence; the template note distribution sequence includes multiple note names and multiple beat identifiers.

[0066] Among them, the audio template can be a template for audio adjustment of human voice audio. The audio template can be constructed in advance, and the audio template includes the text corresponding to the template, and the text includes multiple text words. The above-mentioned audio template includes information such as the template pitch, the text and text words of the audio template, and the template note distribution sequence. Dry audio can be audio that needs to be adjusted, such as audio that needs to be adjusted for ghost audio. Dry audio can be a human voice audio, and dry audio can be voice information recorded by the user. For example, the user can click a related button in the terminal to turn on the ghost recording function. At this time, the terminal can respond to the user's selection and display at least one audio template to be selected. The user can select the corresponding audio template. After the user selects the corresponding template, the terminal can receive a selection instruction for the above-mentioned at least one audio template to be selected, determine the selected audio template, so that the terminal can display the text corresponding to the selected audio template, and the user can record the corresponding dry audio according to the text. The terminal can record the dry audio input by the user based on the text corresponding to the audio template. The terminal can recognize the text contained in the dry audio, and the text contained in the dry audio also includes multiple text words. That is, the terminal can perform speech recognition on the dry audio and identify the text and text words therein. Among them, the dry audio may contain other non-human noises, so the terminal can obtain the text in the dry audio by extracting the fundamental frequency information in the dry audio and identifying the fundamental frequency information. For example, after the terminal obtains the above-mentioned dry audio, it can first obtain the fundamental frequency information in the dry audio and identify the original text corresponding to the dry audio based on the fundamental frequency information. Among them, the original text can be a text that has not been checked for correctness. Since the user may make mistakes or omissions when recording the dry audio, such as the user reverses two words or forgets to say a word, the terminal can perform a correctness check on the original text. For example, the terminal can match the original text with the text corresponding to the audio template to obtain a matching result. The terminal can modify the original text according to the matching result, correct or supplement the erroneous text words and missing text words therein, thereby obtaining the text corresponding to the above-mentioned dry audio and each text word in the text corresponding to the dry audio. That is, each text word in the text corresponding to the dry audio matches each text word in the text corresponding to the audio template.

[0067] Among them, the above-mentioned audio template can be a template for audio adjustment of dry sound audio, and the audio template can be constructed in advance. Each audio template can be regarded as relevant information of a preset song. The template tone in the audio template is the tone of the preset song. The audio template can also include a template beat, which is the beat information of the preset song. The text of the audio template is part or all of the lyrics in the preset song. The template note distribution sequence can be the simplified notation note information of each text word in the text of the above-mentioned audio template, as well as the beat information between these simplified notation notes. The simplified notation note information and beat information of multiple text words can form a note sequence as a template note distribution sequence. For example, in an audio template, its template key can be F, that is, the melody key of the song corresponding to the template is F; the template beat can include BPM (Beat Per Minute) and time signature information, the BPM can be 89, and the time signature can be 4 / 4, that is, 4 beats per measure, and a quarter note is one beat; the text of the audio template can be preset lyrics information, for example, "I really want to go out and play" is used as the text of the audio template; the template note distribution can be the notes and beat information of each text word in the text of the above audio template, and the template note distribution can form a sequence, specifically "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||", where each number in the template note distribution represents a simplified musical notation note, occupying a duration of 1 / 8 note, "|" represents the beat interval, "||" represents the number of measure intervals, and 0 represents an empty beat. The above audio template can be one of the templates in the template library, and the terminal can construct multiple audio templates based on multiple song audios.

[0068] In addition, in some embodiments, the user can also first input human voice audio, and the terminal can then search for a corresponding audio template based on the human voice audio. For example, after the user enters the human voice audio, the terminal can query a template library based on the above-mentioned dry voice audio. The template library stores multiple audio templates, each audio template including the template pitch, template beat, audio template text, and a template note distribution sequence corresponding to the audio template text; when it is detected that the text in the audio template in the template library corresponds to the text of the above-mentioned dry voice audio, the audio template can be output as a query result. Among them, when the user enters the dry voice audio, it can be voice input according to the set audio template text, so that the terminal can obtain the corresponding audio template based on the voice query input by the user according to the set text. For example, in one embodiment, the terminal can perform fundamental frequency detection on the above-mentioned dry voice audio, and perform text recognition on the dry voice audio based on the detected fundamental frequency information to obtain the corresponding dry voice text. The template library is queried based on the dry voice text to obtain the text of the audio template corresponding to the text word and / or text word number of the dry voice text, and the audio template corresponding to the audio template text is used as the audio template corresponding to the dry voice audio. In this embodiment, the user can input voice information into the terminal according to the set text to form dry audio.

[0069] The text in the voice input by the user may be consistent with the number of text characters and the specific pronunciation of each text character in the text of each audio template contained in the template library, or may be a voice consistent with the number of text characters in the text of the audio template. Thus, the terminal can determine whether to query the audio template based on the number of text characters and the specific form of text characters or based on the number of text characters according to the input situation. After the terminal obtains the above-mentioned dry audio input by the user, it can perform fundamental frequency detection on the dry audio. The terminal can perform text recognition on the dry audio based on the fundamental frequency information obtained by fundamental frequency detection to obtain the dry text in the dry audio. The dry text can be the text in the above-mentioned dry audio, the text of the dry audio contains multiple text characters, and the dry text contains the specific form of the text characters and the number of text characters. Thus, the terminal can query the template library based on the dry text to obtain the text of the audio template that corresponds to the number and specific form of the text characters in the dry text, or obtain the text of the audio template that matches the number of text characters in the dry text. Since the text of the audio template exists in the audio template, the terminal can use the audio template corresponding to the text of the audio template obtained by the above query as the audio template corresponding to the dry audio. Specifically, taking the audio template whose text is "I really want to go out and play" as an example, when the terminal queries the template library according to the dry voice text corresponding to the dry voice audio and obtains the audio template corresponding to the text of the audio template "I really want to go out and play", the dry voice text can be a dry voice text that corresponds one-to-one with the number of words and text word form of the five words "I really want to go out and play", that is, the dry voice text is "I really want to go out and play"; the dry voice text can also be a dry voice text that corresponds to the number of words of the five words "I really want to go out and play", that is, the dry voice text contains five words. Among them, the template library can also include multiple audio templates, and the number of words and text word form of the audio template in each audio template can be different.

[0070] Step S204: Detect the note names included between the multiple beat identifiers, determine the note name of each beat, and obtain beat information of the multiple note names.

[0071] Among them, the above-mentioned audio template may include multiple note names and beat identifiers between the multiple note names. After obtaining the dry audio recorded by the user based on the audio template, the terminal can parse the audio template selected by the user, so that the terminal can adjust the user's dry audio based on the various parameters obtained by the analysis. Among them, the adjustment method of the dry audio includes adjusting the beat, so the terminal needs to parse the beat information from the audio template. Since the template note distribution sequence in the audio template includes multiple note names and multiple beat identifiers, the terminal can detect the note names contained in the multiple beat identifiers in the template note distribution sequence, and determine the note name of each beat. Specifically, the note name between two beat identifiers can be used as the note name of one beat, so that the terminal can obtain the beat information of multiple note names.

[0072] Specifically, the template note distribution can form a sequence, specifically "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||", where each number in the template note distribution represents a simplified musical notation note, also known as a note name, occupying a duration of 1 / 8 of a note, "|" represents a beat interval, "||" represents the number of measure intervals, collectively referred to as a beat identifier, and 0 represents a blank beat. When determining the beat information of each note name, the terminal can cyclically detect each character in the template note distribution sequence. If the character is detected as a numeric character, the terminal can determine that the character is a note name; if the terminal detects that the character is a beat identifier, and the character between the beat identifier and the previous beat identifier is a note name, the terminal can regard the note name between the beat identifier and the previous beat identifier as the note name of a beat. After traversing the entire template note distribution sequence, the terminal can obtain the beat information of multiple note names based on the multiple note names of a beat.

[0073] Step S206: Establish a correspondence between each text word in the audio template and the note name of each beat according to the beat information, and obtain the note name corresponding to each text word in the audio template.

[0074] Among them, the beat information can be the beat information of multiple note names in the sequence obtained by the terminal after parsing the above-mentioned template note distribution sequence. When one note name is included in a beat, each text word in the above-mentioned audio template can correspond one-to-one with the note name in the audio template. After the terminal parses the above-mentioned note name and beat information, it can establish a correspondence between each text word in the audio template and the note name of each beat based on the beat information, for example, one text word corresponds to one note name. In this way, the terminal can obtain the note name corresponding to each text word in the audio template. In addition, in some embodiments, the above-mentioned one text word can also correspond to multiple note names. For example, if the terminal parses and obtains that there are two note names in one beat, the terminal can associate multiple note names in one beat with one text word to obtain the note name corresponding to each text word.

[0075] Step S208, determining the pitch value of each text word according to the note name corresponding to each text word in the audio template, adjusting each text word in the dry audio and the audio part corresponding to each text word in the audio template according to the pitch value, to obtain the adjusted dry audio.

[0076] Among them, the terminal can obtain the note name corresponding to each text word by parsing the audio template. The note name can be a digital form of a simplified musical notation note. The terminal can determine the pitch value of each text word based on the note name corresponding to each of the above text words. For example, the above simplified musical notation note can be converted from the note name to the pitch value according to the tone of the audio template. After the terminal obtains the pitch value of each text word, it can adjust the audio part of each text word in the dry audio corresponding to the text word in the audio template according to the pitch value, that is, each text word in the dry audio corresponds to a part of audio, and the terminal can adjust the part of audio according to the pitch value corresponding to the text word, so that the terminal can obtain the adjusted dry audio. Among them, the terminal can convert the pitch value obtained above into frequency, so that the terminal can adjust the audio part corresponding to each text word based on the frequency.

[0077] Specifically, the above-mentioned audio template also includes a template tone. The template tone has a corresponding reference pitch value, which can be obtained by searching a preset tone pitch mapping table. The preset tone pitch mapping table includes a correspondence between multiple tones and reference pitch values. Therefore, the terminal can determine the pitch value of the note name corresponding to each text word in the audio template based on the reference pitch value. In addition, the terminal can also determine the frequency of the note name corresponding to each text word in the audio template based on the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone, thereby realizing the conversion from pitch value to frequency. Therefore, the terminal can adjust the audio part corresponding to each text word in the audio template in the dry audio according to the frequency of the audio name of each text word in the audio template to obtain the adjusted dry audio. Among them, the template tone can be the tonality of the song corresponding to the audio template, which is the keynote of the song. The template note distribution sequence can be a sequence composed of the note and beat information of each text word in the text of the audio template. The terminal can map the template tone to the pitch, and based on the pitch obtained after mapping and the note information corresponding to each text word in the template note distribution sequence, determine the pitch value of each text word in the text of the audio template through calculation, so that the terminal can obtain the pitch value sequence corresponding to the text of the audio template based on the pitch value of each text word in the text of the audio template, that is, when the text of the audio template contains multiple text words, the pitch value sequence corresponding to the text of the audio template includes the pitch values ​​corresponding to each text word.

[0078] After the terminal obtains the pitch value of the text word in the text of the audio template, it can determine the frequency corresponding to the text of the audio template based on the pitch value of the text word in the text of the audio template and the frequency corresponding to the template tone. Specifically, the terminal can obtain the frequency corresponding to each text word by substituting the pitch value of each text word in the text of the audio template into the adjustment formula constructed based on the frequency of the template tone, the pitch value of the template tone and the pitch value of each text word in the text of the audio template. In this way, the terminal can obtain the frequency sequence corresponding to the text of the audio template based on the frequencies of multiple text words, that is, when the text of the audio template contains multiple text words, the frequency sequence of the text of the audio template includes the frequencies corresponding to each text word.

[0079] Among them, the audio part corresponding to the text of the above-mentioned audio template can be the audio part in the dry audio corresponding to each text word in the text of the audio template, that is, each audio part is the pronunciation part of the text word. The terminal can adjust the frequency of the audio part in the dry audio corresponding to each text word based on the frequency corresponding to each text word in the text of the above-mentioned audio template, so that the terminal can obtain the dry audio after frequency adjustment. Among them, the terminal can adjust the audio part corresponding to each text word in the dry audio based on the frequency of the audio part corresponding to each text word, and obtain the above-mentioned adjusted dry audio, that is, the adjusted dry audio can be an audio obtained based on frequency adjustment.

[0080] Step S210: Acquire a template accompaniment corresponding to the audio template, and perform a fusion process on the adjusted dry audio and the template accompaniment based on the beat information to obtain the adjusted audio.

[0081] The dry audio can be user-recorded human voice audio. The dry audio can include the audio portion of each text word in the text of the audio template. The pitch of the audio portion corresponding to each text word in the dry audio can be audio that matches the pitch of each note in the template note distribution sequence. In other words, the dry audio itself can be audio information with a cappella melody.

[0082] The above-mentioned audio template has a corresponding song, and the song has a corresponding template accompaniment. The terminal can then perform a fusion process on the above-mentioned dry audio and the template accompaniment corresponding to the audio template based on the template beat to obtain the adjusted audio. For example, the beat of the above-mentioned template accompaniment can be consistent with the beat of the template. The terminal can adjust the beat of the above-mentioned dry audio to be consistent with the beat of the template, so that the terminal can fuse the dry audio with the template accompaniment with the consistent beat to obtain the adjusted audio.

[0083] In the above audio adjustment method, after selecting an audio template, the corresponding dry audio is recorded and the text therein is identified. The note name of each beat is determined by detecting the note names contained between multiple beat identifiers. A correspondence between each text word of the audio template and the note name of each beat is established based on the beat information of the multiple note names. The pitch value of each text word is determined based on the note name corresponding to each text word. The audio part corresponding to each text word in the dry audio and each text word in the audio template is adjusted based on the pitch value to obtain the adjusted dry audio, and the adjusted dry audio and template accompaniment are fused based on the beat information to obtain the adjusted audio. Compared with the traditional method of manually adjusting audio parameters, this solution constructs an audio template containing pitch, beat, text and note distribution, and parses the template through a specific strategy to obtain parameters for audio adjustment, and adjusts the audio based on the audio template, thereby improving the adjustment efficiency during ghost audio adjustment.

[0084] In one embodiment, it also includes: obtaining the original template audio and its corresponding template accompaniment, beat information and text of the original template audio; obtaining multiple original template notes in the original template audio, and determining multiple template notes corresponding to the number of text words in the text of the original template audio from the multiple original template notes; adding beat identifiers between the multiple template notes according to the multiple template notes and the beat information to obtain a template note distribution sequence; generating an audio template corresponding to the original template audio according to the text of the original template audio and the template note distribution sequence, and storing the audio template and the corresponding template accompaniment in the template library.

[0085] In this embodiment, the terminal can pre-construct an audio template. The audio template is a description of a melody, primarily consisting of a beat and fundamental frequency sequence, and can be in the format of an XML framework. There can be multiple audio templates, and the terminal can store the constructed multiple audio templates in a template library, allowing the terminal to perform audio adjustments based on the multiple audio templates in the template library. When constructing the audio template, the terminal can obtain the original template audio and its corresponding template accompaniment, the beat information corresponding to the original template audio, and the text corresponding to the original template audio. The template accompaniment can be the accompaniment information of the original template audio, the pitch information can be the base tonality of the original template audio, the beat information can be information such as the BPM and tempo of the original template audio, and the text of the original template audio can be the lyrics information of the original template audio. Based on the original template audio and the aforementioned text, the terminal can determine the template musical note corresponding to each text word in the text. For example, the terminal can perform musical note recognition on each audio portion corresponding to the text in the original template audio to identify the template musical note for that audio portion. The template musical note can be a musical notation musical note. The terminal can also determine the distribution information of the template notes based on the template notes corresponding to each text word in the above text and the beat information of the above original template audio. For example, the terminal can determine the beat distribution of the template notes corresponding to each text word based on the beat information, so that the terminal can obtain the template note distribution sequence corresponding to each text word containing beat information.

[0086] After the terminal obtains the above-mentioned template note distribution sequence, it can generate an audio template corresponding to the original template audio according to the beat information corresponding to the original template audio, the text of the audio template and the template note distribution sequence, and the terminal can store the above-mentioned audio template and the corresponding template accompaniment in the template library. In addition, in some embodiments, the audio template can also store the pitch information of the template audio. Specifically, the above-mentioned original template audio can be a song, then the song has a corresponding template beat, the text of the audio template, the template note distribution sequence and the template pitch information. The template pitch can be F key, that is, the melody tonality of the song corresponding to the template is F key; the template beat can include BPM and time signature information, the BPM can be 89, and the time signature can be 4 / 4 beat, that is, 4 beats per measure, and a quarter note is one beat; the text of the audio template can be preset lyrics information, such as "I really want to go out and play" as the text of the audio template; the template note distribution can be each text in the text of the above-mentioned audio template. The note and beat information of the character, the template note distribution can form a sequence, specifically "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||", where each number in the template note distribution represents a simplified notation note, occupying a duration of 1 / 8 note, "|" represents the beat interval, "||" represents the number of measure intervals, 0 represents a blank beat, and when the lyrics content has multiple lines, the above-mentioned audio template text and the template note distribution sequence corresponding to the audio template text can be repeated, and the new lyrics / note distribution information can be filled in. The terminal can then form an audio template based on the above-mentioned template pitch, template beat, audio template text and template note distribution.

[0087] Through this embodiment, the terminal can construct an audio template based on the template pitch, template beat, audio template text, template notes and other information of the original template audio, so that the terminal can parse the audio template and obtain the corresponding template information to adjust the dry sound audio, thereby improving the efficiency of adjusting the ghost audio pitch.

[0088] In one embodiment, a plurality of template notes corresponding to the number of text words in the text of the original template audio are determined from a plurality of original template notes; beat identifiers are added between the plurality of template notes based on the plurality of template notes and beat information to obtain a template note distribution sequence, including: obtaining each text word contained in the text of the audio template; determining the note name corresponding to each text word based on the audio part corresponding to each text word in the original template audio; and adding beat identifiers between each text word based on the beat information to obtain a template note distribution sequence.

[0089] In this embodiment, the terminal can determine the template musical note corresponding to each text word in the text of the audio template based on the original template audio and the text of the audio template corresponding to the original template audio. The text of the audio template can contain at least one text word. The terminal can obtain each text word contained in the text of the audio template and determine the audio portion corresponding to each text word in the original template audio. Thus, the terminal can determine the audio name corresponding to each text word based on the audio portion corresponding to each text word in the original template audio. Specifically, the terminal can perform musical note recognition on the audio portion corresponding to each text word in the original template audio to obtain the musical notation name of each audio portion. The terminal can also determine the beat identifier of each text word based on the beat information of the original template audio. That is, the terminal can allocate each text word in the text of the audio template according to the beat information, determine the text word to appear in each beat of the beat, and add beat identifiers between multiple template musical notes. Thus, the terminal can determine the template musical note distribution sequence corresponding to the text of the audio template based on the note name and beat identifier. Specifically, the text characters contained in the audio template text may be the five characters "I really want to go out and play." The terminal may perform note recognition on the audio portion corresponding to these five characters in the original template audio to obtain the simplified musical notation note names corresponding to each text character, such as the note name "43212." Each number may be a musical notation note, and each number represents the note name of a text character. The terminal may then combine the template beats and the note names corresponding to the text characters to form a template note distribution sequence. For example, the terminal may add a beat identifier "|" where a beat ends and a beat identifier "||" where a measure ends. Specifically, the sequence may be "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||," where each number in the template note distribution represents a musical notation note, occupying a duration of 1 / 8 of a note. Each number within the brackets represents the note name of a text character. "|" represents a beat interval, "||" represents the number of measure intervals, and 0 represents an empty beat.

[0090] Through this embodiment, the terminal can form a corresponding template note distribution sequence based on the note name of the audio part of each text word in the text of the audio template in the original template audio, and the beat information of the original template audio, so that the terminal performs ghost audio adjustment on the dry sound audio based on the template note distribution sequence, thereby improving the audio adjustment efficiency when performing ghost audio adjustment.

[0091] In one embodiment, a reference pitch value of a template tone is obtained, and the pitch value of a note name corresponding to each text word in the audio template is determined according to the reference pitch value, including: querying a preset tone pitch mapping table according to the template tone to determine the reference pitch value corresponding to the template tone; the preset tone pitch mapping table includes a correspondence between multiple tones and reference pitch values; for each text word in the audio template, querying a preset note table to determine the first semitone interval value between the note name corresponding to the text word and the first note in the octave space where the note name is located; the preset note table is obtained based on multiple note names in the same octave space and the first semitone interval value between each note name and the first note in the octave space; and determining the pitch value of the note name corresponding to the text word according to the sum of the reference pitch value and the first semitone interval value corresponding to the note name.

[0092] In this embodiment, the terminal can obtain a template tone from an audio template, and the template tone can be a reference tone of a song. Each tone will correspond to a pitch value, and the terminal can query a preset tone pitch mapping table based on the template tone to determine the reference pitch value corresponding to the template tone. The preset tone pitch mapping table includes a correspondence between multiple tones and reference pitch values. Specifically, the reference pitch value can be regarded as a note number, and the template tone can be F key. The terminal can query the preset tone pitch mapping table based on the template tone, and the preset tone pitch mapping table can be as follows:

[0093] Note Number 60 62 63 65 67 69 71 Tonality C D E F G A B

[0094] As shown in the table above, each tone is associated with a note number, which is the reference pitch value for that tone. For example, if the tone is F, the terminal can determine its corresponding note number is 65 from the table above. Therefore, the reference pitch value for the template tone F is 65.

[0095] The terminal can also determine the semitone intervals of the notes corresponding to each text word in the audio template based on the template note distribution sequence. The semitone interval can be the number of semitones between a note and another reference note. Before determining the semitone intervals corresponding to each text word, the terminal can first extract the note name corresponding to each text word in the audio template from the template note distribution sequence. For example, Figure 2 As shown, Figure 2The figure is a flow chart illustrating the steps for obtaining a template note distribution in one embodiment. The specific form of the template note distribution sequence can be "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||," where each number in the template note distribution sequence represents a note in simplified musical notation, occupying a duration of 1 / 8 of a note. Each number within the brackets represents a note name in a text word. "|" represents a beat interval, "||" represents the number of measure intervals, and 0 represents an empty beat. The terminal then needs to parse the template note distribution sequence and extract the note names of each text word contained therein. The terminal can implement parsing of the template note distribution sequence based on a state machine solution. Specifically, the terminal can determine whether a note has been detected based on the information of each symbol detected in the template note distribution sequence. For example, when the terminal detects the symbol "(," it can proceed to detect the next symbol ")," so that the terminal can interpret the number within the symbol as the note name in a text word. Furthermore, the terminal can also detect the offset information of each note in the template note distribution sequence. For example, the terminal can detect whether each note has an offset identifier. If so, the terminal can add the offset identifier to the extracted corresponding note name. The offset identifier includes octave shifts and semitone shifts. Octave shifts can be represented by + / -, and semitone shifts can be represented by # / b.

[0096] Specifically, if Figure 3 As shown, Figure 3 Schematic diagram of a template note distribution in one embodiment. The template note distribution sequence may be "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||". The terminal may recognize the template note distribution sequence, thereby obtaining the name of each note contained in the template note distribution sequence and the duration of the audio portion corresponding to each note name, and may form the following: Figure 3 The note names and time distribution diagram shown in the figure, where the time can be the time of the dry audio.

[0097] As can be seen from the above template note distribution diagram, the template note distribution sequence may include multiple note names corresponding to text characters. For each text character in the audio template, the terminal may use the note name corresponding to the text character to query a preset note table to determine the first semitone interval value between the note name corresponding to the text character and the first note in the octave space where the note name resides. The preset note table is derived based on multiple note names in the same octave space and the first semitone interval value between each note name and the first note in the octave space. Specifically, the specific form of the preset note table may be as follows:

[0098] note name n 1 2 3 4 5 6 7 semitone interval f(n) 0 2 4 5 7 9 11

[0099] As can be seen from the table above, the note name can be represented by n, i.e., the note name of each text word mentioned above, and the semitone interval can be represented by f(n). The note names in the table above can be names within the same octave space, in which case the first note name in that octave space is 1. The terminal can determine the first semitone interval value corresponding to each note name based on the semitone interval between each note and the first note name, thereby obtaining the above-mentioned preset note table. Specifically, for the above-mentioned template note distribution sequence "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||", the first semitone interval value corresponding to the first note 4 can be 5, indicating that the note is 5 semitones away from the first note in the octave space. Therefore, by querying the above-mentioned preset note table, the terminal can obtain the first semitone interval value corresponding to the note name of each text word. After obtaining the first semitone interval value corresponding to the note name, the terminal can determine the pitch value of the note name corresponding to the text word according to the sum of the reference pitch value and the first semitone interval value corresponding to the note name, thereby obtaining the pitch value corresponding to each text word. Specifically, the calculation formula of the above pitch value can be as follows: note = N base +f(n), where note can be the pitch value of the note name corresponding to each text word, N base It can be the pitch value corresponding to the above template tone, and f(n) can be the first semitone interval value. Taking the note name as 4 in the above template note distribution sequence as an example, its corresponding tone is F. By querying the preset tone pitch mapping table, it can be obtained that the reference pitch corresponding to tone F is 65, then N base If the value of the first semitone interval corresponding to the note name 4 is 5, the terminal can query the preset note table using the note name 4, and obtain the value of the first semitone interval corresponding to the note name 4 as 65. Then, the pitch value corresponding to the note name can be obtained as 70 through the above pitch value calculation formula. Figure 4 As shown, Figure 4 The terminal can perform the above processing on the note name corresponding to each text word in the above audio template to obtain the following pitch distribution diagram: Figure 4 The pitch value distribution diagram shown in FIG. Wherein, each pitch value in the pitch value distribution diagram is Figure 3 The notes in the correspondence.

[0100] In addition, the above-mentioned note name may also include a corresponding identifier. When the terminal detects that the note name has an identifier, it can obtain the pitch value corresponding to the note name in a specific way. For example, in one embodiment, the pitch value of the note name corresponding to each text word in the audio template is determined according to the reference pitch value, and the following is also included: for each text word in the audio template, the identifier of the note name corresponding to the text word is detected; if it is detected that the note name corresponding to the text word has a semitone offset identifier, the semitone offset table is searched according to the semitone offset identifier to determine the second semitone interval value corresponding to the semitone offset identifier; the pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value, the first semitone interval value corresponding to the note name, and the second semitone interval value; the semitone offset The table includes a correspondence between multiple semitone offset identifiers and second semitone interval values; and / or if it is detected that the note name corresponding to the text word has an octave offset identifier, the octave offset table is queried according to the octave offset identifier, the octave offset value corresponding to the octave offset identifier is determined, and the third semitone interval value corresponding to the octave offset value is obtained; the pitch value of the note name corresponding to the text word is determined according to the reference pitch value, the first semitone interval value corresponding to the note name corresponding to the text word, and the sum of the third semitone interval value; the octave offset table includes a correspondence between multiple octave offset identifiers and octave offset values.

[0101] In this embodiment, the above-mentioned audio template may include note names corresponding to multiple text words, and for each text word, the terminal can detect the identifier corresponding to the note name corresponding to the text word. If the terminal detects that the note name corresponding to the text word has a semitone offset identifier, the terminal can query the semitone offset table based on the semitone offset identifier to determine the second semitone interval value corresponding to the semitone offset identifier. Thus, the terminal can determine the pitch value corresponding to the note name corresponding to the text word based on the above-mentioned reference pitch value, the first semitone interval value corresponding to the note name, and the sum of the second semitone interval value. Thus, the terminal can obtain the pitch value corresponding to each text word. Among them, the above-mentioned semitone offset table includes the correspondence between multiple semitone offset identifiers and the second semitone interval value. Specifically, the terminal can use # / b to represent the rise and fall between semitones, that is, the above-mentioned note name can contain a # identifier or a b identifier. The specific form of the semitone offset table can be as follows:

[0102] Text word representation # b (null) shift 1 -1 0

[0103] The terminal can use shift to indicate the second semitone interval value of the semitone offset. The # symbol indicates a semitone increase, and the b symbol indicates a semitone decrease. The pitch calculation formula with the semitone offset symbol can be as follows: note = N base +f(n)+shift. Thus, the terminal can obtain the pitch value of the note with the semitone shift mark based on the above formula.

[0104] If the terminal detects that the note name has an octave offset identifier, the terminal can query the octave offset table according to the octave offset identifier to determine the third semitone interval value corresponding to the octave offset identifier. Thus, the terminal can determine the pitch value corresponding to the note name corresponding to the text word based on the above-mentioned reference pitch value, the first semitone interval value corresponding to the note name corresponding to the text word, and the sum of the third semitone interval value. Thus, the terminal can obtain the pitch value of each text word in the audio template. Among them, the above-mentioned octave offset table includes a plurality of correspondences between octave offset identifiers and third semitone interval values. Specifically, the terminal can use + / - to indicate the rise and fall between octaves, that is, the above-mentioned note name can have a + mark or a - mark. The specific form of the octave offset table can be as follows:

[0105] Text word representation + - (null) oct 1 -1 0

[0106] The terminal can use oct to indicate the octave offset. The + symbol indicates an octave increase, and the - symbol indicates an octave decrease. The pitch calculation formula with the octave offset symbol can be as follows: note = N base +f(n)+12*oct. Since there are 12 semitones in an octave, the terminal can convert the octave offset into a third semitone interval using 12*oct for calculation. Thus, the terminal can obtain the pitch value of the note with the octave offset identifier based on the above formula.

[0107] In addition, in some embodiments, the semitone offset identifier and the octave offset identifier can exist in the same note name at the same time. In this case, the terminal can obtain the pitch value of the note name by the following formula: note=N base +f(n)+shift+12*oct.

[0108] Through the above embodiment, the terminal can determine the semitone interval value between the note name and the reference pitch through the note name and the various identifiers carried by the note name, thereby obtaining the pitch value corresponding to each note name, and adjusting the dry sound audio according to the pitch value, thereby improving the adjustment efficiency of the audio adjustment.

[0109] In one embodiment, the frequency of the note name corresponding to each text word in the audio template is determined based on the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone, including: querying a preset frequency table according to the template tone to obtain the frequency corresponding to the template tone; the preset frequency table includes a plurality of correspondences between template tones and frequencies; for each note name corresponding to each text word in the audio template, obtaining the difference between the pitch value of the note name corresponding to the text word and the reference pitch value corresponding to the template tone, and determining the frequency adjustment value corresponding to the note name according to the difference; determining the frequency corresponding to the note name corresponding to the text word according to the product of the frequency corresponding to the template tone and the frequency adjustment value.

[0110] In this embodiment, the template tone corresponds to corresponding frequency information. The terminal can query a preset frequency table based on the template tone to obtain the frequency corresponding to the template tone. The preset frequency table includes multiple correspondences between template tones and frequencies. The audio template includes multiple text characters corresponding to musical note names. For each text character, the terminal can obtain the difference between the pitch value corresponding to the musical note name and the reference pitch value corresponding to the template tone, and determine the frequency adjustment value corresponding to the musical note name based on this difference. The terminal can then determine the frequency corresponding to the musical note name based on the product of the frequency corresponding to the template tone and the frequency adjustment value, thereby obtaining the frequency of each text character in the audio template. Specifically, taking the template tone of A as an example, a table query shows that the reference pitch value of A is 69, and its corresponding frequency is 440 Hz. The terminal can then calculate the frequency adjustment value by subtracting the pitch value of the musical note name from the reference pitch value and determining the frequency adjustment value based on this difference. The terminal can then determine the frequency corresponding to the musical note name based on the product of 440 and the frequency adjustment value. The frequency calculation formula obtained by the terminal based on the A tone can be as follows:

[0111] Among them, freq represents the frequency of the note name corresponding to each text word in the audio template, and note represents the pitch value of the note name. is the frequency adjustment value mentioned above.

[0112] It should be noted that the frequency and reference pitch values ​​corresponding to the template tones in the above frequency calculation formula may vary depending on the template tones. Taking the above template note distribution sequence "(4,4)|(3,3)|(2,2)|(1,1)||(2,2)|0,0|0,0|0,0||" as an example, the frequency information of each note name obtained by the terminal through the above frequency formula can be as follows: Figure 5 As shown, Figure 5 FIG. 4 is a schematic diagram of the frequency distribution after adjustment in one embodiment.

[0113] Through this embodiment, the terminal can determine the corresponding frequency calculation formula based on the template tone, thereby obtaining the frequency corresponding to the note name corresponding to each text word, and then the terminal can determine the adjustment of the audio part corresponding to each note name based on these frequencies, thereby improving the adjustment efficiency of the audio adjustment.

[0114] In one embodiment, the adjusted dry audio and the template accompaniment are fused to obtain the adjusted audio, including: determining the duration of the audio portion corresponding to each text word in the adjusted dry audio according to the beat information, and adjusting the duration of the audio portion corresponding to each text word in the adjusted dry audio according to the duration of the audio portion corresponding to each text word, so that the duration of the audio portion corresponding to each text word in the adjusted dry audio matches the duration of each beat in the template beat, to obtain dry audio with adjusted duration; and mixing the dry audio with adjusted duration with the template accompaniment to obtain the adjusted audio.

[0115] In this embodiment, the above-mentioned audio template may include a template beat, and the template beat may be a beat corresponding to the beat of the template accompaniment. The terminal may determine the duration corresponding to each text word in the dry audio according to the template beat. The adjusted dry audio may be the audio obtained after the terminal performs pitch adjustment on each text word in the dry audio entered by the above-mentioned user. The terminal may adjust the duration of the audio part corresponding to each text word in the dry audio according to the duration corresponding to each text word, so that the duration of the audio part corresponding to each text word in the dry audio matches the duration of each beat of the template beat. Specifically, the terminal may achieve duration matching by increasing or reducing the duration of the audio part corresponding to each text word. The terminal may mix the above-mentioned dry audio after duration adjustment with the template accompaniment to obtain the adjusted audio. For example, it may be a ghost audio.

[0116] Through this embodiment, the terminal can match and mix the dry audio with the template accompaniment based on the template beat to obtain the adjusted audio, thereby improving the audio adjustment efficiency of the ghost audio.

[0117] In one embodiment, Figure 6 As shown, Figure 6It is a flow chart of the audio adjustment method in another embodiment. In this embodiment, the above-mentioned audio adjustment can be an adjustment of a ghost audio. The following steps are included: the terminal can construct a ghost template, that is, the above-mentioned audio template, by the above-mentioned method. The terminal can obtain the user's dry sound material, such as the dry sound audio recorded by the user. Specifically, the terminal can display the above-mentioned audio templates, and the user can select an audio template from the multiple audio templates displayed by the terminal, and record the human voice audio according to the text corresponding to the audio template, so that the terminal can record the dry sound audio input by the user as a dry sound material. The terminal can segment the dry sound material into lyrics, and perform audio adjustment based on the above-mentioned ghost template for each lyric, thereby obtaining dry sound melody information. The terminal can also obtain the accompaniment material and obtain the beat information therein, and match the dry sound melody information with the beat through the ghost template. The terminal can also mix the matched accompaniment and the user's dry sound melody to obtain a ghost audio, that is, the above-mentioned adjusted audio, and output it.

[0118] Through this embodiment, the terminal constructs an audio template containing pitch, beat, text and note distribution, and parses the audio template through a specific strategy to obtain parameters for adjusting the dry audio, and adjusts the audio based on the audio template, thereby improving the adjustment efficiency when adjusting ghost audio.

[0119] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0120] Based on the same inventive concept, embodiments of the present application also provide an audio adjustment device for implementing the aforementioned audio adjustment method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more of the following embodiments of the audio adjustment device can be found in the aforementioned limitations of the audio adjustment method and will not be further elaborated here.

[0121] In one embodiment, Figure 7 As shown, an audio adjustment device is provided, including: a recording module 500, a detection module 502, a corresponding module 504, an adjustment module 506 and a fusion module 508, wherein:

[0122] The recording module 500 is configured to select an audio template, record dry audio corresponding to the audio template, and recognize text and text characters in the dry audio; the audio template includes a template tone, the text and text characters of the audio template, and a template note distribution sequence; the template note distribution sequence includes multiple note names and multiple beat identifiers;

[0123] A detection module 502 is configured to detect the note names contained between the multiple beat identifiers, determine the note name of each beat, and obtain beat information of the multiple note names;

[0124] A corresponding module 504 is used to establish a corresponding relationship between each text word in the audio template and the note name of each beat according to the beat information, so as to obtain the note name corresponding to each text word in the audio template;

[0125] An adjustment module 506 is configured to determine a pitch value of each text word according to the note name corresponding to each text word in the audio template, and adjust the audio portion corresponding to each text word in the dry audio according to the pitch value to obtain adjusted dry audio.

[0126] The fusion module 508 is used to obtain the template accompaniment corresponding to the audio template, and perform a fusion process on the adjusted dry audio and the template accompaniment based on the beat information to obtain the adjusted audio.

[0127] In one embodiment, the above-mentioned recording module 500 is specifically used to display at least one audio template to be selected; receive a selection instruction for at least one audio template to be selected, determine the selected audio template; display the text corresponding to the audio template, and record the dry audio input by the user based on the text corresponding to the audio template.

[0128] In one embodiment, the above-mentioned recording module 500 is specifically used to obtain the fundamental frequency information of the dry audio, identify the original text corresponding to the dry audio based on the fundamental frequency information; modify the original text based on the matching result between the original text and the text of the audio template to obtain the text of the dry audio, and obtain each text word in the text of the dry audio; match each text word in the text of the dry audio with each text word in the text of the audio template.

[0129] In one embodiment, the above-mentioned device also includes: a construction module, which is used to obtain the original template audio and its corresponding template accompaniment, beat information and text of the original template audio; obtain multiple original template notes in the original template audio, and determine multiple template notes corresponding to the number of text words in the text of the original template audio from the multiple original template notes; add beat identifiers between the multiple template notes according to the multiple template notes and beat information to obtain a template note distribution sequence; generate an audio template corresponding to the original template audio according to the text of the original template audio and the template note distribution sequence, and store the audio template and the corresponding template accompaniment in the template library.

[0130] In one embodiment, the above-mentioned detection module 502 is specifically used to cyclically detect each character in the template note distribution sequence. If the character is detected as a numeric character, the character is determined to be a note name; if the character is detected as a beat identifier, and the character between the beat identifier and the previous beat identifier is a note name, the note name between the beat identifier and the previous beat identifier is used as the note name of a beat; based on multiple note names of one beat, the beat information of multiple note names is obtained.

[0131] In one embodiment, the above-mentioned adjustment module 506 is specifically used to obtain a reference pitch value of the template tone, and determine the pitch value of the note name corresponding to each text word in the audio template according to the reference pitch value; and determine the frequency of the note name corresponding to each text word in the audio template according to the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone; and adjust the audio part of the dry sound audio corresponding to each text word in the audio template according to the frequency of the note name corresponding to each text word in the audio template to obtain the adjusted dry sound audio.

[0132] In one embodiment, the above-mentioned adjustment module 506 is specifically used to query the preset tone pitch mapping table according to the template tone to determine the reference pitch value corresponding to the template tone; the preset tone pitch mapping table includes the correspondence between multiple tones and reference pitch values; for each text word in the audio template, query the preset note table to determine the first semitone interval value between the note name corresponding to the text word and the first note in the octave space where the note name is located; the preset note table is obtained based on multiple note names in the same octave space and the first semitone interval value between each note name and the first note in the octave space; according to the sum of the reference pitch value and the first semitone interval value corresponding to the note name, determine the pitch value of the note name corresponding to the text word.

[0133] In one embodiment, the adjustment module 506 is specifically configured to detect, for each text word in the audio template, an identifier of a note name corresponding to the text word; if it is detected that the note name corresponding to the text word has a semitone offset identifier, query a semitone offset table based on the semitone offset identifier to determine a second semitone interval value corresponding to the semitone offset identifier; determine the pitch value of the note name corresponding to the text word based on a reference pitch value, a sum of a first semitone interval value corresponding to the note name, and a second semitone interval value; the semitone offset table includes a correspondence between multiple semitone offset identifiers and second semitone interval values; and / or if it is detected that the note name corresponding to the text word has an octave offset identifier, query an octave offset table based on the octave offset identifier to determine an octave offset value corresponding to the octave offset identifier, and obtain a third semitone interval value corresponding to the octave offset value; determine the pitch value of the note name corresponding to the text word based on a sum of the reference pitch value, the first semitone interval value corresponding to the note name corresponding to the text word, and the third semitone interval value; the octave offset table includes a correspondence between multiple octave offset identifiers and octave offset values.

[0134] In one embodiment, the above-mentioned adjustment module 506 is specifically used to query a preset frequency table according to the template tone to obtain the frequency corresponding to the template tone; the preset frequency table includes a plurality of correspondences between template tones and frequencies; for each note name corresponding to each text word in the audio template, obtain the difference between the pitch value of the note name corresponding to the text word and the reference pitch value corresponding to the template tone, and determine the frequency adjustment value corresponding to the note name according to the difference; determine the frequency corresponding to the note name corresponding to the text word according to the product of the frequency corresponding to the template tone and the frequency adjustment value.

[0135] In one embodiment, the above-mentioned fusion module 508 is specifically used to determine the duration of the audio part corresponding to each text word in the adjusted dry audio according to the template beat, and adjust the duration of the audio part corresponding to each text word in the adjusted dry audio according to the duration of the audio part corresponding to each text word, so that the duration of the audio part corresponding to each text word in the adjusted dry audio matches the duration of each beat in the template beat to obtain dry audio with adjusted duration; the dry audio with adjusted duration is mixed with the template accompaniment to obtain adjusted audio.

[0136] Each module in the aforementioned audio adjustment device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0137] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an audio adjustment method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0138] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0139] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned audio adjustment method when executing the computer program.

[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned audio adjustment method is implemented.

[0141] In one embodiment, a computer program product is provided, comprising a computer program, which implements the above-mentioned audio adjustment method when executed by a processor.

[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0143] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0144] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An audio adjustment method, characterized in that: The method comprises: Selecting an audio template, recording dry audio corresponding to the audio template and identifying text and text words in the dry audio; the audio template includes a template tone, text and text words of the audio template, and a template note distribution sequence; the template note distribution sequence includes multiple note names and multiple beat identifiers; Detecting the note names contained between the multiple beat identifiers, determining the note name of each beat, and obtaining beat information of the multiple note names; According to the beat information, establishing a correspondence between each text word in the audio template and the note name of each beat, to obtain the note name corresponding to each text word in the audio template; Obtaining a reference pitch value of the template tone, determining a pitch value of a note name corresponding to each text word in the audio template based on the reference pitch value, determining a frequency of the note name corresponding to each text word based on the pitch value of the note name corresponding to each text word and the frequency corresponding to the template tone, and adjusting an audio portion corresponding to each text word in the dry audio based on the frequency of the note name corresponding to each text word to obtain adjusted dry audio; A template accompaniment corresponding to the audio template is obtained, and a fusion process is performed on the adjusted dry audio and the template accompaniment based on the beat information to obtain an adjusted audio.

2. The method according to claim 1, characterized in that The step of selecting an audio template and recording dry audio corresponding to the template includes: Displaying at least one audio template to be selected; receiving a selection instruction for the at least one audio template to be selected, and determining a selected audio template; The text corresponding to the audio template is displayed, and the dry audio input by the user based on the text corresponding to the audio template is recorded.

3. The method according to claim 1, characterized in that The identifying of text and text words in the dry audio includes: Acquiring fundamental frequency information of the dry audio, and identifying original text corresponding to the dry audio according to the fundamental frequency information; Based on the matching result between the original text and the text of the audio template, the original text is modified to obtain the text of the dry audio, and each text word in the text of the dry audio is obtained; each text word in the text of the dry audio is matched with each text word in the text of the audio template.

4. The method according to claim 1, wherein The method further comprises: Obtain the original template audio and its corresponding template accompaniment, beat information and text of the original template audio; Acquire a plurality of original template notes in the original template audio, and determine a plurality of template notes corresponding to the number of text words in the text of the original template audio from the plurality of original template notes; Adding beat identifiers between the multiple template notes according to the multiple template notes and the beat information to obtain a template note distribution sequence; According to the text of the original template audio and the template note distribution sequence, an audio template corresponding to the original template audio is generated, and the audio template and the corresponding template accompaniment are stored in a template library.

5. The method according to claim 1, wherein The detecting the note names included between the multiple beat identifiers, determining the note name of each beat, and obtaining the beat information of the multiple note names includes: cyclically detecting each character in the template note distribution sequence, and if the character is detected to be a numeric character, determining that the character is a note name; If it is detected that the character is a beat mark, and the character between the beat mark and the previous beat mark is a note name, the note name between the beat mark and the previous beat mark is regarded as the note name of a beat; According to the multiple note names of one beat, beat information of the multiple note names is obtained.

6. The method according to claim 1, characterized in that The step of obtaining a reference pitch value of the template tone and determining the pitch value of the note name corresponding to each text word in the audio template according to the reference pitch value includes: According to the template tone, a preset tone pitch mapping table is searched to determine a reference pitch value corresponding to the template tone; the preset tone pitch mapping table includes a correspondence between multiple tones and reference pitch values; For each text word in the audio template, query a preset note table to determine a first semitone interval value between the note name corresponding to the text word and the first note in the octave space where the note name is located; the preset note table is obtained based on multiple note names in the same octave space and the first semitone interval value between each note name and the first note in the octave space; The pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value and the first semitone interval value corresponding to the note name.

7. The method according to claim 6, characterized in that The step of determining the pitch value of the note name corresponding to each text word in the audio template according to the reference pitch value further includes: For each text word in the audio template, detecting an identifier of a note name corresponding to the text word; If it is detected that the note name corresponding to the text word has a semitone offset identifier, a semitone offset table is searched according to the semitone offset identifier to determine the second semitone interval value corresponding to the semitone offset identifier; the pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value, the first semitone interval value corresponding to the note name, and the second semitone interval value; the semitone offset table includes a correspondence between multiple semitone offset identifiers and second semitone interval values; and / or If it is detected that the note name corresponding to the text word has an octave offset identifier, the octave offset table is queried according to the octave offset identifier, the octave offset value corresponding to the octave offset identifier is determined, and the third semitone interval value corresponding to the octave offset value is obtained; the pitch value of the note name corresponding to the text word is determined according to the sum of the reference pitch value, the first semitone interval value corresponding to the note name corresponding to the text word, and the third semitone interval value; the octave offset table includes a plurality of correspondences between octave offset identifiers and octave offset values.

8. The method according to claim 6, characterized in that The step of determining the frequency of the note name corresponding to each text word in the audio template according to the pitch value of the note name corresponding to each text word in the audio template and the frequency corresponding to the template tone includes: According to the template tone, a preset frequency table is searched to obtain a frequency corresponding to the template tone; the preset frequency table includes a correspondence between a plurality of template tones and frequencies; For each note name corresponding to each text word in the audio template, obtaining a difference between a pitch value of the note name corresponding to the text word and a reference pitch value corresponding to the template tone, and determining a frequency adjustment value corresponding to the note name based on the difference; The frequency corresponding to the note name corresponding to the text word is determined according to the product of the frequency corresponding to the template tone and the frequency adjustment value.

9. The method according to claim 1, characterized in that The audio template further includes: a template beat, and the fusion processing of the adjusted dry audio and the template accompaniment to obtain the adjusted audio includes: Determining, according to the template beat, the duration of the audio portion corresponding to each text word in the adjusted dry audio, and adjusting the duration of the audio portion corresponding to each text word in the adjusted dry audio according to the duration of the audio portion corresponding to each text word, so that the duration of the audio portion corresponding to each text word in the adjusted dry audio matches the duration of each beat in the template beat, thereby obtaining dry audio with adjusted duration; The dry audio with the adjusted duration is mixed with the template accompaniment to obtain the adjusted audio.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Audio generation method and device, computer equipment and storage medium

    CN114333744A

  • Audio adjustment method and computer equipment

    CN114582306A