Audio adjustment method, computer device and computer program product
By obtaining the singing skill information and accuracy of the user's singing audio and setting a personalized adjustment strategy, the problem of insufficient adjustment effect of the fixed method in the existing technology is solved, and a better audio adjustment effect is achieved.
Patent Information
- Application Number
- CN202210333509.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-03-31
AI Technical Summary
In the prior art, when adjusting the user's singing audio in a fixed manner, the adjustment effect is insufficient and cannot adapt to the singing levels of different users.
By obtaining the audio to be adjusted and its corresponding standard audio, the singing skill information and singing accuracy are determined, different adjustment strategies are set according to the singing level, and personalized adjustments are made to the singing skill information and non-singing skill information parts.
It realizes personalized audio adjustment according to the user's singing level, improves the effect of audio adjustment, and adapts to the singing level of different users.
Smart Images

Figure CN114743526B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio adjustment method, apparatus, computer equipment, storage medium, and computer program product. Background Art
[0002] With the advancement of computer technology, users can now record singing audio using mobile devices such as mobile phones. Since each user's singing skills vary, the mobile device can adjust the recorded singing audio to achieve an effect that matches the original singer's singing. Currently, adjusting users' singing audio is typically done in a fixed manner. However, this fixed method suffers from insufficient adjustment effects. Summary of the Invention
[0003] Based on this, it is necessary to provide an audio adjustment method, device, computer equipment, computer-readable storage medium and computer program product that can improve the audio adjustment effect in response to the above technical problems.
[0004] In a first aspect, the present application provides an audio adjustment method, the method comprising:
[0005] Acquire the audio to be adjusted and its corresponding standard audio, determine singing skill information in the audio to be adjusted, and determine the singing accuracy of the audio to be adjusted based on the standard audio;
[0006] Obtaining a singing level corresponding to the audio to be adjusted according to the singing accuracy;
[0007] Based on the singing level, different adjustment strategies are determined for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information, and the audio to be adjusted is adjusted based on the adjustment strategy to obtain the adjusted audio.
[0008] In one embodiment, determining the singing skill information in the audio to be adjusted includes:
[0009] Performing fundamental frequency detection on the audio to be adjusted to obtain a fundamental frequency sequence of the audio to be adjusted;
[0010] According to the fundamental frequency sequence, singing skill information in the audio to be adjusted is identified.
[0011] In one embodiment, identifying singing skill information in the audio to be adjusted based on the fundamental frequency sequence includes:
[0012] Identifying at least one singing skill information in the fundamental frequency sequence;
[0013] The skill type of the at least one singing skill information and the time information and the number of times the at least one singing skill information appears in the fundamental frequency sequence are obtained to obtain the singing skill information in the audio to be adjusted.
[0014] In one embodiment, determining the singing accuracy of the audio to be adjusted based on the standard melody information of the standard audio includes:
[0015] Acquiring standard melody information of the standard audio;
[0016] The singing accuracy of the audio to be adjusted is determined according to the matching degree between the standard melody information and the fundamental frequency sequence of the audio to be adjusted.
[0017] In one embodiment, determining the singing accuracy of the audio to be adjusted based on the matching degree between the standard melody information and the fundamental frequency sequence of the audio to be adjusted includes:
[0018] For each line of lyrics in the standard audio, obtaining a standard melody template sequence corresponding to the line of lyrics in the standard melody information, and determining a cosine similarity between the standard melody template sequence corresponding to the line of lyrics and a fundamental frequency sequence corresponding to the line of lyrics in the audio to be adjusted;
[0019] The singing accuracy of the audio to be adjusted is determined based on an average value of cosine similarities corresponding to multiple lyrics.
[0020] In one embodiment, obtaining a singing level corresponding to the audio to be adjusted based on the singing accuracy, and determining different adjustment strategies for a target audio portion corresponding to the singing skill information and a non-target audio portion not containing the singing skill information in the audio to be adjusted based on the singing level, includes:
[0021] If the singing accuracy is greater than or equal to a first value, determining that the singing level is a first level; and determining, based on the first level, a first adjustment strategy for a target audio portion corresponding to the singing skill information and a non-target audio portion not containing the singing skill information in the audio to be adjusted;
[0022] If the singing accuracy is less than the first value and greater than or equal to the second value, determining that the singing level is the second level; and determining, based on the second level, a second adjustment strategy for the target audio portion corresponding to the singing skill information and the non-target audio portion not containing the singing skill information in the audio to be adjusted;
[0023] If the singing accuracy is less than the second value, the singing level is determined to be the third level; based on the third level, a third adjustment strategy is determined for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information; wherein, the first value is greater than the second value; the degree of adjustment of the target audio part under the first adjustment strategy, the second adjustment strategy, and the third adjustment strategy increases successively.
[0024] In one embodiment, the singing skill information includes at least two of vibrato information, glissando information, transition information, and transition tone information;
[0025] The adjusting the audio to be adjusted based on the adjustment strategy includes:
[0026] If the adjustment strategy is the first adjustment strategy, pitch shifting is performed on target audio parts corresponding to various singing technique information contained in the audio to be adjusted, and pitch shifting and amplitude compression are performed on non-target audio parts based on the first adjustment strategy to fit the standard audio;
[0027] If the adjustment strategy is the second adjustment strategy, performing one of pitch shifting and amplitude compression processing on the first singing skill information contained in the audio to be adjusted and / or performing pitch shifting and amplitude compression processing on the second singing skill information contained in the audio to be adjusted based on the second adjustment strategy, and performing pitch shifting and amplitude compression processing on the non-target audio portion to fit the standard audio;
[0028] If the adjustment strategy is the third adjustment strategy, based on the third adjustment strategy, pitch shifting and amplitude compression processing are performed on the target audio part and non-target audio part corresponding to the various singing skill information contained in the audio to be adjusted to fit the standard audio.
[0029] In one embodiment, adjusting the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio includes:
[0030] Acquire a plurality of note-like units in a fundamental frequency sequence corresponding to the audio to be adjusted;
[0031] Based on the singing range corresponding to the sequence of the multiple note-like units, the adjusted audio is obtained according to the adjustment strategy.
[0032] In one embodiment, adjusting the audio to be adjusted according to the adjustment strategy based on the singing range corresponding to the sequence of the plurality of note-like units to obtain the adjusted audio includes:
[0033] If the sequence containing the note-like units is a target sequence, adjusting the audio to be adjusted according to the adjustment strategy within the octave interval of the target sequence to obtain an adjusted audio; wherein the target sequence is a sequence of adjacent sentences in the fundamental frequency sequence and in the same octave interval;
[0034] or,
[0035] If the sequence in which the note-like units are located is a sequence other than the target sequence, the audio to be adjusted is adjusted according to the adjustment strategy within the octave interval corresponding to the standard melody information to obtain the adjusted audio.
[0036] In one embodiment, obtaining the audio to be adjusted includes:
[0037] Obtaining original audio, and determining a sound quality score of the original audio based on noise information of the original audio;
[0038] If the sound quality score is less than a preset score threshold, noise reduction processing is performed on the original audio to obtain the audio to be adjusted.
[0039] In one embodiment, after obtaining the audio to be adjusted and its corresponding standard audio, the method further includes:
[0040] Obtaining a fundamental frequency sequence corresponding to the audio to be adjusted, and obtaining fundamental frequency parameters in the fundamental frequency sequence; the fundamental frequency parameters include at least one of the following: a range, an average pitch, and a pitch fluctuation sequence;
[0041] Inputting the fundamental frequency parameter into a preset classification model, and determining the audio type of the audio to be adjusted according to an output result of the preset classification model, wherein the audio type includes reading audio or singing audio;
[0042] If the audio to be adjusted is singing audio, the steps of determining singing skill information in the audio to be adjusted and determining the singing accuracy of the audio to be adjusted based on the standard audio are performed.
[0043] In one embodiment, after determining the audio type of the audio to be adjusted according to the output result of the preset classification model, the method further includes:
[0044] If the audio to be adjusted is a reading audio, audio synthesis is performed based on the standard audio and the human voice tonality information and human voice pitch information in the audio to be adjusted to obtain the adjusted audio.
[0045] In a second aspect, the present application provides an audio adjustment device, the device comprising:
[0046] A first acquisition module is used to acquire the audio to be adjusted and its corresponding standard audio, determine the singing skill information in the audio to be adjusted, and determine the singing accuracy of the audio to be adjusted based on the standard audio;
[0047] A second acquisition module is used to obtain the singing level corresponding to the audio to be adjusted according to the singing accuracy;
[0048] An adjustment module is used to determine different adjustment strategies for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information based on the singing level, and adjust the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio.
[0049] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0050] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0051] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0052] The above-mentioned audio adjustment method, apparatus, computer equipment, storage medium and computer program product obtain the audio to be adjusted and its corresponding standard audio, determine the singing skill information in the audio to be adjusted, and determine the singing accuracy of the audio to be adjusted based on the standard melody information of the standard audio, obtain the singing level of the audio to be adjusted based on the singing accuracy, determine the adjustment strategy for the target audio part corresponding to the singing skill information and the non-target audio part that does not contain the singing skill information in the audio to be adjusted based on the singing level, and adjust the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio. Compared with the traditional method of adjusting audio based on a fixed method, this solution determines the singing level of the audio to be adjusted input by the user, and implements different audio adjustment strategies based on the user's singing level, thereby achieving audio adjustment that adapts to the user's level and improving the adjustment effect of the audio adjustment. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A diagram showing an application environment of an audio adjustment method according to an embodiment;
[0054] Figure 2 1 is a flow chart of an audio adjustment method according to an embodiment;
[0055] Figure 3A schematic diagram of an interface for an audio adjustment step in one embodiment;
[0056] Figure 4 is a schematic diagram of a baseband sequence in one embodiment;
[0057] Figure 5 A schematic diagram of an adjustment strategy in one embodiment;
[0058] Figure 6 A schematic diagram of a step of determining a singing range in one embodiment;
[0059] Figure 7 A flowchart of the step of obtaining an audio type in one embodiment;
[0060] Figure 8 1 is a flow chart of an audio synthesis step in one embodiment;
[0061] Figure 9 is a flowchart of an audio synthesis step in another embodiment;
[0062] Figure 10 is a flowchart of an audio adjustment method according to another embodiment;
[0063] Figure 11 is a structural block diagram of an audio adjustment device in one embodiment;
[0064] Figure 12 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0066] The audio adjustment method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the audio to be adjusted input by the user, and send the audio to be adjusted to the server 104. The server 104 can analyze the user's singing level based on the obtained audio to be adjusted, and adjust the corresponding strategy for the audio to be adjusted based on the user's singing level. In addition, the server 104 can also send the adjusted audio to the terminal 102 for playback. Among them, the terminal 102 can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.
[0067] In one embodiment, Figure 2 As shown, an audio adjustment method is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the server in the example:
[0068] Step S202, obtaining the audio to be adjusted and its corresponding standard audio, determining the singing skill information in the audio to be adjusted, and determining the singing accuracy of the audio to be adjusted based on the standard melody information of the standard audio.
[0069] The audio to be adjusted may be audio input by the user, such as the audio of a song recorded by the user while singing. The standard audio may be the standard audio corresponding to the audio to be adjusted. Taking the adjustment of the user's singing audio as an example, the server 104 may obtain the original audio corresponding to the user's singing audio as the standard audio. For example, if the terminal 102 is a mobile phone, the mobile phone may capture the song recording produced by the user singing and send the song recording to the server 104. The server 104 may obtain the ID of the original song corresponding to the song recording. After receiving the audio adjustment task, the server 104 confirms that the materials required for the task, including the song recording, the melody of the original song, and the lyrics file, are complete, and then adjusts the user's song recording.
[0070] In the process of obtaining the audio to be adjusted, the server 104 may obtain the audio to be adjusted through certain processing. For example, in one embodiment, obtaining the audio to be adjusted includes: obtaining the original audio, determining the sound quality score of the original audio based on noise information of the original audio; and if the sound quality score is less than a preset score threshold, performing noise reduction processing on the original audio to obtain the audio to be adjusted.
[0071] In this embodiment, the audio obtained by the server 104 can be the original audio input by the user. Since the recording environments of users are different, the quality of the recorded audio will also be different, and this quality can be reflected in the noise level contained in the original audio. Then, the server 104 can obtain the noise information in the original audio and determine the sound quality score of the original audio based on the noise information in the original audio. The server 104 can also determine whether noise reduction processing needs to be performed on the original audio based on the sound quality score. The server 104 can detect whether the above sound quality score is less than a preset score threshold. If so, the server 104 can determine that there is more noise information in the original audio. At this time, the server 104 can perform noise reduction processing on the original audio to obtain the待调整音频 (to-be-adjusted audio). Specifically, the above original audio can be the user's dry voice input by the user. The server 104 can evaluate the sound quality of the user's dry voice and estimate a sound quality score X, where 0 < X < 100. When X is less than a certain threshold, for example, less than 50 points, the server 104 needs to perform sound quality processing on the user's dry voice to ensure that there are no excessive interference components in the dry voice input before pitch correction. The server 104 can perform noise reduction on the user's dry voice through methods such as OMLSA (optimally-modified log-spectral amplitude) and spleeter (source separation algorithm). Among them, OMLSA is a noise reduction algorithm mainly for steady-state noise and is a classic single-channel audio noise reduction algorithm. After the server 104 performs audio noise reduction through the above methods, it can obtain the to-be-adjusted audio for audio adjustment.
[0072] Among them, for singing, the singing levels of each user are different. Therefore, a unified pitch correction template cannot be used to perform pitch correction on the user's audio, and targeted pitch correction needs to be performed according to the singing level of each user. After obtaining the to-be-adjusted audio of the user, the server 104 can determine the singing skill information that appears in the to-be-adjusted audio, and the server 104 can also obtain the standard melody information of the standard audio corresponding to the to-be-adjusted audio. Thus, the server 104 can determine the singing accuracy of the to-be-adjusted audio according to the standard melody information of the standard audio. That is, the server 104 can perform MIR (Music Information Retriveal) analysis on the to-be-adjusted audio to obtain the above singing skill information and singing accuracy information.
[0073] The singing technique information includes at least one of vibrato information, glissando information, vibrato information, and transition information. Glissando refers to a singing or playing method that slides upward or downward from one note to another; vibrato refers to the fluctuation of the pitch of the voice within a certain range and at a certain frequency during singing, making the notes sound fuller; and vibrato refers to legato transitions, singing three to five notes in a short period of time, which can make the sound rich and varied. In addition, the above-mentioned singing technique information may also include other techniques, such as true voice, falsetto, strong voice, weak voice, breathy voice, pharyngeal voice, mute voice, angry voice, choked voice, crying voice, etc. Different vocal techniques can play a different role in enhancing the emotions of the singing effect.
[0074] Server 104 can determine the user's singing level based on the singing skill information and singing accuracy obtained above, and thus make targeted audio adjustments. For example, for users with good singing skills, the user's original singing skills or singing style must be retained after tuning; for users with poor singing skills, more reference template information is required to ensure that the user's pitch can be significantly improved after tuning.
[0075] Step S204: Obtain the singing level corresponding to the audio to be adjusted according to the singing accuracy.
[0076] The singing accuracy can be determined by server 104 based on the matching degree obtained by comparing the standard melody information with the audio to be adjusted. For example, if the audio to be adjusted is singing audio, the singing accuracy can represent the user's pitch matching degree for the entire song. Based on the singing accuracy obtained above, server 104 can obtain a singing grade corresponding to the audio to be adjusted. The singing grade can represent the singing proficiency of the user's input audio to be adjusted. A higher singing grade indicates a higher singing proficiency of the user. In other words, the singing grade is a reflection of singing proficiency.
[0077] Step S206, based on the singing level, determine the adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information, and adjust the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio.
[0078] Among them, the singing level may include multiple levels, for example, it may include three levels, such as the first level, the second level and the third level. The above three levels can represent different singing levels of the audio to be adjusted input by the user, and the singing levels represented by the first level, the second level and the third level decrease in sequence. For example, the first level may indicate that the singing level of the user's audio to be adjusted is relatively high, the second level may indicate that the singing level of the user's audio to be adjusted is medium, and the third level may indicate that the singing level of the user's audio to be adjusted is relatively poor. For different singing levels, the server 104 may determine different adjustment strategies. In one implementation, the adjustment strategy may indicate the degree of adjustment of the audio. The higher the singing level represented by the singing level, the smaller the adjustment degree of the adjustment strategy. On the contrary, the lower the singing level represented by the singing level, the greater the adjustment degree of the adjustment strategy.
[0079] The server 104 can determine the adjustment strategy for the target audio portion (the audio portion corresponding to the singing skill information) in the audio to be adjusted based on the singing level. Thus, the server 104 can adjust the target audio portion in the above-mentioned audio to be adjusted based on different adjustment strategies to obtain the adjusted audio. Among them, the above-mentioned singing skill information can include multiple types, and the above-mentioned adjustment strategy can include an adjustment method for each type of singing skill information. For example, the audio adjustment can include pitch shifting and amplitude compression methods. The server 104 can determine whether the target audio portion of each type of singing skill information contained in the audio to be adjusted needs to perform pitch shifting and amplitude compression based on the above-mentioned singing level. In addition, the lower the above-mentioned singing level, the more parameters the server 104 adjusts.
[0080] Specifically, if Figure 3 As shown, Figure 3 The following is a schematic diagram of an interface for the audio adjustment step in an embodiment. For example, the terminal 102 is a mobile phone, and the audio to be adjusted is a user's singing audio. The user can record the singing audio in the mobile phone. The server 104 can obtain the singing audio sent by the mobile phone, identify the singing level of the audio based on the singing audio, and return a recommended adjustment intensity to the mobile phone, including options such as light, medium, and deep. Figure 3 The recommended adjustment intensity can be light, wherein light, medium, and deep adjustment can correspond to the first, second, and third levels, respectively. The light, medium, and deep adjustment options correspond to different adjustment strategies, that is, the lower the user's singing level, the more powerful the adjustment option recommended by server 104 to ensure the adjustment effect of the audio after tuning. If the user is not satisfied with the audio adjustment option recommended by server 104, he or she can also select the corresponding adjustment option on his or her own, so that server 104 can adjust the singing audio based on the adjustment strategy corresponding to the adjustment option selected by the user to obtain the adjusted audio.
[0081] Specifically, the server 104 may perform audio adjustment based on the baseband sequence of the audio to be adjusted. After adjusting the baseband sequence based on the corresponding adjustment strategy, the server 104 may obtain an adjusted frequency shift sequence. The server 104 may then perform pitch shifting on the audio to be adjusted input by the user based on the frequency shift sequence to obtain the adjusted audio. Pitch shifting refers to a technique that can change the speed or pitch of the audio.
[0082] In the above-mentioned audio adjustment method, the singing skill information in the audio to be adjusted is determined by obtaining the audio to be adjusted and its corresponding standard audio, and the singing accuracy of the audio to be adjusted is determined based on the standard melody information of the standard audio. The singing level of the audio to be adjusted is obtained based on the singing accuracy. Based on the singing level, an adjustment strategy for the target audio portion corresponding to the singing skill information and the non-target audio portion that does not contain the singing skill information in the audio to be adjusted is determined, and the audio to be adjusted is adjusted based on the adjustment strategy to obtain the adjusted audio. Compared with the traditional method of adjusting audio based on a fixed method, this solution determines the singing level of the audio to be adjusted input by the user, and implements different audio adjustment strategies based on the user's singing level, thereby achieving audio adjustment that adapts to the user's level and improving the adjustment effect of the audio adjustment.
[0083] In one embodiment, determining the singing skill information in the audio to be adjusted includes: performing fundamental frequency detection on the audio to be adjusted to obtain a fundamental frequency sequence corresponding to the audio to be adjusted; and identifying the singing skill information in the audio to be adjusted based on the fundamental frequency sequence.
[0084] In this embodiment, after the server 104 obtains the audio to be adjusted input by the user, it can determine the singing skill information in the audio to be adjusted. The server 104 can identify the singing skill information in the audio to be adjusted based on the baseband sequence. The server 104 can perform baseband detection on the audio to be adjusted to obtain the baseband sequence corresponding to the audio to be adjusted, and the server 104 can identify the singing skill information therein based on the baseband sequence. For example, Figure 4 As shown, Figure 4 The server 104 can detect the fundamental frequency of the input audio. The fundamental frequency is generated by the vibration of the vocal cords. Generally, voiced sounds have a fundamental frequency. The fundamental frequency sequence obtained by the server 104 after detecting the fundamental frequency of the audio to be adjusted can be as follows: Figure 4 As shown in the curve in, the server 104 can also smooth the above-mentioned fundamental frequency sequence through HMM (Hidden Markov Model) to calculate the pitch envelope NLU (Note-LikeUnit) with short-term stability. The note-like unit can be as follows Figure 4As shown by the horizontal line in , the server 104 can adjust the fundamental frequency sequence based on the aforementioned note-like units, thereby adjusting the audio to be adjusted. Before adjusting the fundamental frequency sequence of the audio to be adjusted, the server 104 can identify the singing technique information therein based on the fundamental frequency sequence, so that the server 104 can determine the singing technique information that needs to be adjusted when adjusting the audio.
[0085] For example, in one embodiment, identifying singing skill information in the audio to be adjusted based on the fundamental frequency sequence includes: identifying at least one singing skill information in the fundamental frequency sequence; obtaining the skill type of at least one singing skill information and the time information and number of times it appears in the fundamental frequency sequence, to obtain the singing skill information in the audio to be adjusted.
[0086] In this embodiment, the audio to be adjusted may be a singing audio, and the singing skill information includes vibrato information, glissando information, transposition information, and transition sound information. The server 104 identifies the singing skill information from the above-mentioned baseband sequence. The singing skill information may include multiple types, and the user's corresponding audio to be adjusted may include one or more types, or may not include any singing skill information. When the above-mentioned audio to be adjusted contains singing skill information, the server 104 can identify at least one type of singing skill information in its corresponding baseband sequence, and obtain the skill type of the at least one identified singing skill information, the time information of each type of singing skill information appearing in the above-mentioned baseband sequence, and the number of times it appears, to obtain the singing skill information corresponding to the singing skill contained in the baseband sequence.
[0087] Specifically, taking the example of singing audio as the audio to be adjusted input by the user, after server 104 performs fundamental frequency detection on the singing audio, server 104 may record the fundamental frequency sequence of the audio to be adjusted as F0. Server 104 may then use a vibrato detection algorithm, a portamento detection algorithm, and a transposition detection algorithm in the fundamental frequency sequence to calculate the time and frequency information of the user using these techniques in the audio to be adjusted. Server 104 can then combine the type of singing technique and the time and frequency information of its appearance in the fundamental frequency sequence to form a type of singing technique information. Based on multiple types of singing technique information, server 104 can obtain the singing technique information corresponding to the audio to be adjusted. Specifically, server 104 may detect different singing technique information in the fundamental frequency sequence using different detection methods. For example, the server 104 can determine the vibrato information in the fundamental frequency sequence by calculating the periodicity of the fundamental frequency sequence's fluctuations at a fixed pitch; the server 104 can determine the glissando information in the fundamental frequency sequence by using a 3-state HMM model for modeling, and determine the pitch change trend through modeling; the server 104 can also obtain the standard melody template information corresponding to the standard audio, and compare the number of notes in the standard melody template information with the number of notes corresponding to the fundamental frequency sequence, thereby determining the transition information in the fundamental frequency sequence; the server 104 can also obtain the transition sound information in the fundamental frequency sequence by identifying the jumps between pitches in the fundamental frequency sequence. Among them, the transition sound can represent the connection between words during the singing process, that is, the jump between pitches is not achieved overnight, and there will be a slow transition process.
[0088] The singing technique information in the fundamental frequency sequence identified by the server 104 may include at least one of vibrato information, portamento information, transposition information, and transitional tone information. When performing audio adjustment, the server 104 may determine, based on the adjustment strategy, what adjustments to make to the identified singing technique information to ensure the adjustment effect.
[0089] Through this embodiment, the server 104 can obtain the fundamental frequency sequence corresponding to the audio to be adjusted through fundamental frequency detection, and identify the singing skill information contained therein based on the fundamental frequency sequence, so that the server 104 can determine what adjustments need to be made to the above-mentioned singing skill information based on the adjustment strategy determined during the audio adjustment, thereby improving the adjustment effect of the audio adjustment.
[0090] In one embodiment, the singing accuracy of the audio to be adjusted is determined based on the standard melody information of the standard audio, including: obtaining the standard melody information corresponding to the standard audio; and determining the singing accuracy of the audio to be adjusted based on the matching degree between the standard melody information and the fundamental frequency sequence corresponding to the audio to be adjusted.
[0091] In this embodiment, server 104 may obtain standard audio corresponding to the audio to be adjusted. For example, if the audio to be adjusted is a user singing, the standard audio may be the original singing audio of the song corresponding to the singing audio. Server 104 may obtain standard melody information from the standard audio and determine the singing accuracy of the audio to be adjusted based on the degree of match between the standard melody information and the fundamental frequency sequence corresponding to the audio to be adjusted.
[0092] Among them, the server 104 can convert the standard melody information into sequence information and then match it with the fundamental frequency sequence. For example, in one embodiment, the singing accuracy of the audio to be adjusted is determined based on the matching degree between the standard melody information and the fundamental frequency sequence corresponding to the audio to be adjusted, including: for each line of lyrics in the standard audio, obtaining the standard melody template sequence corresponding to the line of lyrics in the standard melody information, determining the cosine similarity between the standard melody template sequence corresponding to the line of lyrics and the fundamental frequency sequence corresponding to the line of lyrics in the audio to be adjusted; based on the average value of the cosine similarities corresponding to multiple lines of lyrics, determining the singing accuracy of the audio to be adjusted. In this embodiment, the above-mentioned standard audio can be an original song audio, then the standard audio can include multiple lines of lyrics, and the standard audio includes a standard lyrics template and a standard melody template. Then, the server 104 can convert the standard melody template into a sequence form to form a standard melody template sequence, wherein the standard melody template can include a partial sequence corresponding to each line of lyrics. For each line of lyrics in the standard audio, server 104 can obtain the standard melody template sequence corresponding to the line of lyrics in the standard melody information. Server 104 can then determine the cosine similarity between the standard melody template sequence corresponding to the line of lyrics and the fundamental frequency sequence corresponding to the line of lyrics in the audio to be adjusted. That is, server 104 can perform a sequence-based comparison between the standard audio and the audio to be adjusted to obtain the corresponding cosine similarity. Specifically, server 104 can obtain the aforementioned cosine similarity corresponding to each line of lyrics. Server 104 can also obtain the average cosine similarity of multiple lines of lyrics. Thus, server 104 can determine the singing accuracy of the audio to be adjusted based on the average cosine similarity of the multiple lines of lyrics.
[0093] Specifically, taking the singing audio input by the user as an example, each of the above lyrics can correspond to a single sentence score x, then the server 104 can obtain the single sentence score x by calculating the cosine similarity between the template pitch sequence of the standard melody template information and the user's singing pitch sequence sentence by sentence. The server 104 can accumulate the scores x of each single sentence and obtain the comprehensive pitch score y by calculating the average value. The calculation formula of the single sentence score is as follows: x = 100*cosθ; cosθ = (A*B) / (‖A‖*‖B‖), where A is the standard pitch sequence of the standard melody template information, and B is the singing pitch sequence of the user's audio to be adjusted, that is, the above-mentioned fundamental frequency sequence. The calculation formula of the comprehensive pitch score y can be as follows: Where N is the number of lyrics in the standard audio, or the number of all single sentences with scores x, x i is the sentence score of the i-th sentence, where i is greater than or equal to 1 and less than or equal to N. Thus, the server 104 can determine the singing accuracy of the user's audio to be adjusted based on the comprehensive pitch score.
[0094] Through the above embodiment, the server 104 can match the fundamental frequency sequence of the standard audio with the audio to be adjusted, and determine the singing accuracy based on the cosine similarity, so that the server 104 can determine the singing level of the user's audio to be adjusted based on the singing accuracy, and make corresponding audio adjustments based on the singing level, thereby improving the adjustment effect of the audio adjustment.
[0095] In one embodiment, according to the singing accuracy, the singing level corresponding to the audio to be adjusted is obtained, and based on the singing level, the adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part not containing the singing skill information is determined, including: if the singing accuracy is greater than or equal to the first value, the singing level is determined to be the first level; according to the first level, the first adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part not containing the singing skill information is determined; if the singing accuracy is less than the first value and greater than or equal to the second value, the singing level is determined to be the second level; according to the second level, the second adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part not containing the singing skill information is determined; if the singing accuracy is less than the second value, the singing level is determined to be the third level; according to the third level, the third adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part not containing the singing skill information is determined; wherein, the first value is greater than the second value; the degree of adjustment of the target audio part under the first adjustment strategy, the second adjustment strategy, and the third adjustment strategy increases successively. Of course, the number of singing levels and adjustment strategies is not limited to three, and can also be other values set according to actual conditions.
[0096] In this embodiment, the server 104 can grade the singing level of the audio to be adjusted input by the user, for example, determine the singing level of the user's audio to be adjusted through the singing accuracy determined above. The server 104 can compare the singing accuracy with a first value and a second value respectively, wherein the first value is greater than the second value. If the server 104 detects that the singing accuracy is greater than or equal to the first value, the server 104 can determine that the singing level of the user's audio to be adjusted is the first level, so that the server 104 can determine that the adjustment strategy for the target audio portion corresponding to the singing skill information in the audio to be adjusted and the non-target audio portion that does not contain the singing skill information is the first adjustment strategy when the singing level is determined to be the first level. If the server 104 detects that the singing accuracy is less than the first value and greater than or equal to the second value, the server 104 can determine that the singing level is the second level, so that the server 104 can determine that the adjustment strategy for the target audio portion corresponding to the singing skill information in the audio to be adjusted and the non-target audio portion that does not contain the singing skill information is the second adjustment strategy when the singing level is determined to be the second level. If the server 104 detects that the singing accuracy is less than the second value, the server 104 can determine that the singing level is the third level, so that the server 104 can determine that the adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information is the third adjustment strategy when the singing level is determined to be the third level. Moreover, under the above-mentioned first adjustment strategy, second adjustment strategy and third adjustment strategy, the adjustment degree of the target audio part of the singing skill information in the audio to be adjusted by the server 104 increases successively. That is, the user singing level represented by the above-mentioned first level, second level and third level decreases successively, and the lower the user singing level, the greater the adjustment degree of the target audio part corresponding to the singing skill information and the non-target audio part that does not contain the singing skill information by the server 104.
[0097] Specifically, the singing level can be divided into three levels: the first level can be professional, the second level can be semi-professional, and the third level can be amateur. The server 104 can then combine the calculated singing accuracy scores and singing skills to classify the singing level of the user's recorded audio to be adjusted. The server 104 can record the professional level as 0, the semi-professional level as 1, and the amateur level as 2. The calculation formula for the singing level of the user's audio to be adjusted can be as follows:
[0098] As can be seen from the formula, the server 104 can determine that the user's audio to be adjusted is of different singing levels based on the different singing accuracies calculated above.
[0099] Through this embodiment, the server 104 can determine the singing level of the audio to be adjusted input by the user based on the singing accuracy, so that the server 104 can determine different adjustment strategies for the audio to be adjusted based on different singing levels, thereby improving the adjustment effect of the audio adjustment.
[0100] In one embodiment, based on the singing level, an adjustment strategy is determined for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information, including: obtaining a fundamental frequency sequence corresponding to the audio to be adjusted, and obtaining multiple note-like units in the fundamental frequency sequence; according to the singing level, an adjustment strategy is determined for the target audio part corresponding to the singing skill information in the multiple note-like units and the non-target audio part that does not contain the singing skill information.
[0101] In this embodiment, the server 104 may convert the audio to be adjusted into the following format before performing audio adjustment: Figure 4 The fundamental frequency sequence shown in FIG. 1 may include multiple note-like units NLU in the fundamental frequency sequence. After the server 104 performs fundamental frequency detection on the audio to be adjusted and obtains the fundamental frequency sequence of the audio to be adjusted, it may obtain multiple note-like units in the fundamental frequency sequence, such as Figure 4 As shown by the horizontal line in . Each note-like unit may correspond to a line of lyrics in the standard audio. Based on the singing level determined above, the server 104 may determine an adjustment strategy for the target audio portion corresponding to the singing skill information in the multiple note-like units in the fundamental frequency sequence and the non-target audio portion that does not contain the singing skill information. That is, the server 104 may adjust the singing skill information by taking the note-like unit as the adjustment object.
[0102] The server 104 performs two adjustments on the waveform corresponding to the singing technique information in the note-like unit, including: shifting the NLU (Note-Like Unit) of the user's fundamental frequency, and adjusting the dynamic range of the fundamental frequency jitter within the NLU. Pitch shifting refers to the operation of returning the pitch value that deviates from the template to the standard value through the overall up-down and down-tuning operation; dynamic range adjustment refers to controlling the jitter amplitude of the fundamental frequency sequence within a single NLU, for example Figure 4 Rectangle 800 and rectangle 802 in the image. Among them, the degree of jitter in rectangle 800 is relatively large, while the amplitude of the fundamental frequency jitter in rectangle 802 is relatively stable. It should be noted that if the fundamental frequency sequence has almost no jitter in the current NLU, it will produce a mechanical feeling in the listening experience. Therefore, excessive jitter or fixed pitch are not desirable. Server 104 can determine what adjustments need to be made to the waveform of the singing technique information in the note-like unit based on the above-determined adjustment strategy.
[0103] For example, in one embodiment, the audio to be adjusted is adjusted based on the adjustment strategy, including: if the adjustment strategy is the first adjustment strategy, pitch shifting the target audio portion corresponding to various singing skill information contained in multiple note-like units based on the first adjustment strategy, and pitch shifting and amplitude compression processing are performed on the non-target audio portion to fit the standard audio, for example, to fit the standard melody template sequence corresponding to the standard audio information; if the adjustment strategy is the second adjustment strategy, pitch shifting and amplitude compression processing are performed on the first singing skill information contained in multiple note-like units and / or one of pitch shifting and amplitude compression processing is performed on the second singing skill information contained in multiple note-like units in the audio to be adjusted, and pitch shifting and amplitude compression processing are performed on the non-target audio portion to fit the standard audio, for example, to fit the standard melody template sequence corresponding to the standard melody information; if the adjustment strategy is the third adjustment strategy, pitch shifting and amplitude compression processing are performed on both the target audio portion and the non-target audio portion corresponding to various singing skill information contained in multiple note-like units based on the third adjustment strategy to fit the standard melody template sequence corresponding to the standard melody information.
[0104] In this embodiment, the singing technique information may include at least one of vibrato information, portamento information, vibrato information, and transitional sound information. Server 104 may first determine an adjustment strategy for each type of singing technique information based on the user's singing level. The singing level includes a first level, a second level, and a third level, corresponding to the first adjustment strategy, the second adjustment strategy, and the third adjustment strategy, respectively. Furthermore, the degree of adjustment for the singing technique information increases under the first adjustment strategy, the second adjustment strategy, and the third adjustment strategy, respectively. The note-like units in the fundamental frequency sequence may be marked with the time and number of occurrences of the singing technique information. If server 104 determines that the adjustment strategy is the first adjustment strategy, indicating that the user's singing level for the audio to be adjusted is the first level, which is professional, server 104 may reduce the degree of adjustment for the singing technique information. Based on the first adjustment strategy, server 104 may pitch-shift the target audio portion corresponding to the various singing technique information contained in the multiple note-like units in the fundamental frequency sequence. For example, pitch-shift the waveform portion containing the singing technique information in the note-like units, so that this portion of the sequence conforms to the standard melody template sequence corresponding to the standard melody information.
[0105] When the server 104 determines that the adjustment strategy is the second adjustment strategy, it means that the singing level of the user's audio to be adjusted is the second level, which is a semi-professional level. The degree of adjustment of the singing skill information of the audio to be adjusted by the server 104 can be increased compared to the professional level. Based on the second adjustment strategy, the server 104 can perform pitch shifting and amplitude compression processing on the target audio part corresponding to the vibrato information contained in the multiple note-like units in the above-mentioned fundamental frequency sequence, perform pitch shifting processing on the target audio part corresponding to the glissando information and transition sound information contained in the multiple note-like units, and perform amplitude compression processing on the transposition sound information contained in the multiple note-like units. Among them, the above-mentioned vibrato information can be the first singing skill information, and the above-mentioned glissando information, transition sound information and transposition sound information can be the second singing skill information; and for the non-target audio part that does not contain singing skill information in the note-like unit, the server 104 can perform pitch shifting and amplitude compression processing on these non-target audio parts in any adjustment strategy. Therefore, after making corresponding adjustments to the above-mentioned multiple singing skill information, the server 104 can make the adjusted fundamental frequency sequence fit the standard melody template sequence corresponding to the standard melody information. Among them, the server 104 can only perform corresponding processing on the singing skill information in the note-like unit after detecting the corresponding singing skill information.
[0106] When server 104 determines that the adjustment strategy is the third adjustment strategy, it indicates that the singing level of the user's audio to be adjusted is the third level, which is an amateur level. Server 104 may adjust the singing skill information of the audio to be adjusted to a greater degree than that of a semi-professional level. Based on the third adjustment strategy, server 104 may perform pitch shifting and amplitude compression on both the target audio portions and non-target audio portions corresponding to the various singing skills contained in the multiple note-like units in the fundamental frequency sequence, so that the adjusted fundamental frequency sequence conforms to the standard melody template sequence corresponding to the standard melody information.
[0107] Taking the above-mentioned audio to be adjusted as the singing audio input by the user as an example, the standard audio can be the original singing audio corresponding to the singing audio, and the above-mentioned amplitude compression processing can be a dynamic range adjustment, for example Figure 4 The rectangle 800 and rectangle 802 in the figure are shown in FIG. Among them, the jitter degree of rectangle 800 is relatively large, while the fundamental frequency jitter amplitude in rectangle 802 is relatively stable. Then, the server 104 can perform amplitude compression processing on rectangle 800 to improve the effect of audio adjustment. The selection of the adjustment strategy based on the singing level can be as follows: Figure 5 As shown, Figure 5Schematic diagram of an adjustment strategy in one embodiment. The use of singing techniques is closely related to the user's singing style. If the user's pitch is completely adjusted to the template pitch during the tuning process, it will destroy the user's original singing style and have a counterproductive effect. Therefore, when processing semi-professional and professional audio to be adjusted, server 104 will try its best to protect the shape of the pitch sequence of common singing techniques such as vibrato, glissando, and transposition, while ensuring pitch accuracy. In addition, for professional and semi-professional audio, the transition sound information of the transition parts between word heads and words needs to be preserved as much as possible without destroying their original shape. It should be noted that in addition to adjusting the target audio parts corresponding to the singing technique information in the above-mentioned note-like units, server 104 can also adjust other ordinary notes that do not contain singing technique information, such as the non-target audio parts in the above-mentioned note-like units. For the ordinary notes in the note-like units, server 104 can perform pitch shifting and dynamic range adjustment, so that the adjusted sequence fits the corresponding part of the standard melody template sequence.
[0108] Through the above embodiment, the server 104 can make corresponding adjustments to the singing skill information in the note-like units based on multiple note-like units in the fundamental frequency sequence as adjustment units and based on the adjustment strategy corresponding to the singing level, so that the adjusted audio not only conforms to the standard melody information but also can retain the user's own singing characteristics as much as possible, thereby improving the adjustment effect of the audio adjustment.
[0109] In one embodiment, the audio to be adjusted is adjusted based on an adjustment strategy to obtain the adjusted audio, including: determining the singing range corresponding to the sequence in which multiple note-like units are located; adjusting the audio to be adjusted according to the adjustment strategy based on the singing range corresponding to the sequence in which multiple note-like units are located to obtain the adjusted audio.
[0110] In this embodiment, the vocal range refers to the range between the lowest and highest notes that can be sung. A full vocal range covers the entire range from the lowest to the highest notes. Generally speaking, singers with a wide vocal range are better able to perform songs with a wide vocal range. Server 104 performs MIR analysis on the audio to be adjusted to identify singing technique information and vocal range characteristics. Server 104 can classify the vocal range of the audio to be adjusted into multiple categories, such as bass, baritone, tenor, alto, mezzo-soprano, and soprano. The vocal range encompasses multiple octaves, and notes within different octaves may have the same name but different pitches. Therefore, the vocal range of the user represented by the audio to be adjusted and the vocal range of the standard audio song are often not fully matched. For example, if the audio to be adjusted is a user singing a song in a female key, a male singer will typically lower the pitch by one octave, but some sentences may still be sung in the female key. If audio adjustments are performed entirely within a fixed vocal range, the adjustments will be too harsh. Therefore, the server 104 can evaluate the octave distribution of the audio to be adjusted based on the sentence, thereby obtaining the singing range of the audio to be adjusted.
[0111] For example, the note-like unit may include a corresponding sequence, and each note-like unit may correspond to each line of lyrics. The server 104 may first determine the singing range corresponding to the sequence of the plurality of note-like units, that is, determine the singing range to be Figure 4 The singing range corresponding to each line of lyrics can be obtained by looking at the range of the sequence spanned by the horizontal line in the sequence. Since the fundamental frequency sequence includes multiple lines of lyrics, the server 104 can obtain the singing range corresponding to the sequence of multiple note-like units. Among them, the singing range corresponding to the sequence of each note-like unit can be determined based on whether there are other adjacent sequences of note-like units in the same singing range in the sequence of the note-like unit. The server 104 can adjust the audio to be adjusted using the above adjustment strategy in the octave range corresponding to the singing range corresponding to the obtained multiple note-like units, and obtain the adjusted audio.
[0112] The singing range of the sequence where the above-mentioned note-like units are located can be determined by detecting the singing range of the note-like units of the adjacent sentences. For example, in one embodiment, determining the singing range corresponding to the sequence where multiple note-like units are located includes: obtaining a target sequence in the fundamental frequency sequence that is an adjacent sentence and is in the same octave interval, using the octave interval of the target sequence as the singing range of the target sequence during audio adjustment, and using the octave interval corresponding to the standard melody information as the singing range of other sequences other than the target sequence in the fundamental frequency sequence during audio adjustment; determining the sequence type of the sequence where the note-like units are located, the sequence type including the target sequence or other sequences, and determining the singing range corresponding to the sequence where the note-like units are located according to the sequence type.
[0113] In this embodiment, server 104 can obtain sequences within the same octave interval within the fundamental frequency sequence. Server 104 may then obtain sequences corresponding to multiple octave intervals. Each sequence within the same octave interval may contain one or more note-like units, meaning each sequence within the same octave interval may include one or more lyrics. Server 104 can detect whether these sequences within the same octave interval contain adjacent lyrics, for example, if two or more adjacent lyrics are within the same octave interval. If so, server 104 can use the sequence containing two or more adjacent lyrics within the same octave interval as the target sequence. Server 104 can use the octave interval of the target sequence as the singing range of the target sequence during audio adjustment. For other sequences within the fundamental frequency sequence other than the target sequence, such as sequences containing only a single sentence within a specific octave interval, server 104 can use the octave interval corresponding to the standard melody information of the standard audio as the singing range of these other sequences during audio adjustment. The server 104 can determine whether the sequence type of the sequence where the note-like unit is located belongs to the above-mentioned target sequence or the above-mentioned other sequences, so that the server 104 can determine the singing range of the sequence where the note-like unit is located based on the sequence type.
[0114] After the server 104 determines the singing range of the sequence where the note-like units are located, the server 104 can adjust the audio to be adjusted based on the adjustment strategy corresponding to the note-like units in the singing range corresponding to each sequence where the note-like units are located to obtain the adjusted audio. For example, in one embodiment, based on the singing ranges corresponding to the sequences where the note-like units are located, the audio to be adjusted is adjusted according to the adjustment strategy to obtain the adjusted audio, including: if the sequence where the note-like units are located is a target sequence, within the octave interval of the target sequence, adjusting the audio to be adjusted according to the adjustment strategy to obtain the adjusted audio; or if the sequence where the note-like units are located is a sequence other than the target sequence, within the octave interval corresponding to the standard melody information, adjusting the audio to be adjusted according to the adjustment strategy to obtain the adjusted audio.
[0115] In this embodiment, the server 104 can determine whether the above-mentioned note-like unit belongs to the target sequence or another sequence outside the target sequence. If the server 104 determines that the note-like unit belongs to the target sequence, the server 104 can adjust the audio to be adjusted within the octave interval of the singing range corresponding to the target sequence based on the adjustment strategy corresponding to the singing level of the audio to be adjusted, including adjusting the sequence corresponding to the note-like unit based on the above-mentioned adjustment strategy, and performing pitch shifting processing on the audio to be adjusted based on the obtained frequency shift sequence, etc., to obtain the adjusted audio. If the server 104 determines that the note-like unit belongs to another sequence outside the target sequence, the server 104 can adjust the audio to be adjusted within the octave interval corresponding to the standard melody information using the adjustment strategy corresponding to the singing level of the audio to be adjusted, including adjusting the sequence corresponding to the note-like unit based on the above-mentioned adjustment strategy, and performing pitch shifting processing on the audio to be adjusted based on the obtained frequency shift sequence, etc., to obtain the adjusted audio.
[0116] Specifically, if Figure 6 As shown, Figure 6 The figure is a schematic diagram of the steps for determining the singing range in one embodiment. Each note-like unit in the above-mentioned fundamental frequency sequence can represent a line of lyrics. When there are multiple consecutive sentences in the same octave interval, these adjacent sentences can form a sentence group. For example, if and only if the number of adjacent sentences in the same octave interval is greater than or equal to 2, it is counted as a sentence group. The server 104 can perform tuning according to the octave melody of the audio in the sentence group. For audio that is not in a sentence group, the server 104 can perform tuning according to the melody information of the standard pitch template corresponding to the standard audio. This avoids repeated octave changes in the tuned audio.
[0117] Through the above embodiment, the server 104 can determine the singing range corresponding to each sequence of note-like units based on the sequence type of the sequence in which the note-like units are located, and can use the corresponding adjustment strategy to adjust the audio to be adjusted in the singing range corresponding to the sequence of each note-like unit, thereby avoiding repeated octave changes in the audio after tuning, ensuring relative octave consistency, and improving the adjustment effect of the audio adjustment.
[0118] In one embodiment, after obtaining the audio to be adjusted and its corresponding standard audio, it also includes: obtaining the fundamental frequency sequence corresponding to the audio to be adjusted, and obtaining the fundamental frequency parameters in the fundamental frequency sequence; the fundamental frequency parameters include at least one of the following: range, average pitch and pitch fluctuation sequence; inputting the fundamental frequency parameters into a preset classification model, and determining the audio type of the audio to be adjusted according to the output result of the preset classification model, the audio type includes reading audio or singing audio; if the audio to be adjusted is singing audio, executing the steps of determining the singing skill information in the audio to be adjusted, and determining the singing accuracy of the audio to be adjusted according to the standard melody information of the standard audio.
[0119] In this embodiment, the audio to be adjusted input by the user can be multiple types of audio, for example, it can be a reading audio, or it can be a singing audio, etc. The server 104 can identify the audio type of the audio to be adjusted input by the user. The server 104 can first perform fundamental frequency detection on the audio to be adjusted, obtain a corresponding fundamental frequency sequence, and obtain fundamental frequency parameters in the fundamental frequency sequence, including at least one of a range, an average pitch, and a pitch fluctuation sequence. The server 104 can input the above-mentioned fundamental frequency parameters into a preset classification model, and based on the output result of the preset classification model, determine the audio type of the audio to be adjusted, that is, determine whether the audio to be adjusted is a reading audio or a singing audio. When the server 104 determines that the audio to be adjusted is a singing audio, the server 104 can execute the steps of determining the singing skill information in the audio to be adjusted, and determining the singing accuracy of the audio to be adjusted based on the standard melody information of the standard audio, thereby adjusting the audio to be adjusted based on the singing accuracy and singing skill information in the audio to be adjusted.
[0120] The process of classifying the audio to be adjusted by the server 104 can be as follows: Figure 7 As shown, Figure 7 The flowchart of the step of obtaining the audio type in one embodiment is shown in FIG. The fundamental frequency is generated by the vocal cords, and generally voiced sounds have a fundamental frequency. The server 104 can detect the fundamental frequency of the audio to be adjusted, and obtain the following: Figure 4 The fundamental frequency sequence may include multiple note-like units. The server 104 may analyze the note-like units to obtain the maximum fundamental frequency value F0 in the audio to be adjusted. max and the minimum fundamental frequency F0 min , range F0 range =F0 max -F0 min , average pitch And fundamental frequency related features such as the first-order difference sequence of pitch fluctuation. Among them, the first-order difference sequence of pitch fluctuation can be a pitch fluctuation sequence, which can better represent the severity of fundamental frequency jitter by reducing irregular fluctuations between data. Its first-order difference formula is as follows: ΔF0 t =F0 t+1 -F0 t The server 104 may select some features, such as the above-mentioned F0 range 、F0 mean or ΔF0 t For example, taking the Gaussian Mixture Model (GMM) as an example, the server 104 may input the above ΔF0 t Sequence, the classification probability is obtained by the maximum likelihood principle, and its function can be shown as follows: Among them, y is the input feature sequence, is a parameter of the GMM. Thus, the server 104 can determine whether the audio to be adjusted is a spoken audio or a sung audio based on the classification probability. For example, when the classification probability is greater than a set probability threshold, the server 104 can determine that the audio to be adjusted is a sung audio; otherwise, the audio to be adjusted is a spoken audio.
[0121] When the audio to be adjusted is a spoken audio, the server 104 may use different audio adjustment methods to adjust the audio to be adjusted. For example, in one embodiment, after determining the audio type of the audio to be adjusted based on the output of a preset classification model, the method further includes: if the audio to be adjusted is a spoken audio, performing audio synthesis based on the standard audio and the human voice tone information and human voice pitch information in the audio to be adjusted to obtain the adjusted audio.
[0122] In this embodiment, when the server 104 determines that the audio to be adjusted is a reading audio, it can perform audio synthesis based on the standard audio corresponding to the audio to be adjusted, the human voice tone information and the human voice pitch information in the audio to be adjusted to obtain the adjusted audio. The above audio synthesis can be a singing synthesis solution, and singing synthesis can be implemented in many ways. For example, Figure 8 As shown, Figure 8 The figure is a flowchart of the audio synthesis step in one embodiment. The server 104 can perform feature extraction on the input human voice, extract the pitch information F0, the spectral envelope representing the human voice timbre, and the non-periodic information Aperiodicity. The server 104 can also extract the score pitch and score duration from the standard audio corresponding to the human voice. The server 104 can decouple the pitch and timbre, and use the fundamental frequency adjustment technology to use the score pitch and score duration to modify the pitch information of the input human voice separately, so that the pitch sequence is close to the standard template pitch, and then send the modified pitch features, user timbre and other non-periodic information to the vocoder part for resynthesis, thereby achieving the purpose of pitch correction of the original user dry voice.
[0123] In addition, vocal synthesis can also be Figure 9 As shown, Figure 9 The flow chart of the audio synthesis step in another embodiment is shown. The server 104 can train a speaker identification network speaker encoder to extract the speaker's timbre speaker embedding when synthesizing the singing voice, wherein the speaker identification network can exist with the following example. Figure 9 In the acoustic model shown, the server 104 can thus control the timbre of the synthesized output based on the user's timbre.
[0124] Through the above embodiment, when the server 104 detects that the audio to be adjusted is a spoken audio, it can implement audio adjustment of the audio to be adjusted by audio synthesis, thereby improving the adjustment effect of different types of audio to be adjusted.
[0125] In one embodiment, Figure 10 As shown, Figure 10 It is a flowchart of the audio adjustment method in another embodiment. It includes the following processes: the server 104 can perform sound quality detection on the user's dry voice input by the user, detect the noise, echo, popping, etc. in the recorded dry voice, and if the sound quality score is low, the server 104 can perform further noise suppression; the server 104 can also identify the type of the user's dry voice to be adjusted, including singing and reading, etc. The server 104 can adopt different sound adjustment schemes for different types of audio. For the audio dry voice of the reading type, the server 104 can achieve sound adjustment through audio synthesis, for example, by extracting the user's timbre, and synthesizing it with the template melody and template lyrics of the standard audio to obtain synthesized audio. Because the pitch changes, stress, and vocal habits of the reading dry voice are quite different from those of the singing dry voice, ordinary sound adjustment schemes cannot achieve a good sound adjustment effect. Therefore, the server 104 adopts the method of audio synthesis to improve the adjustment effect.
[0126] When the dry voice is singing audio, the server 104 can perform MIR analysis on the singing audio to identify the singing techniques such as vibrato, vibrato, glissando, etc., and analyze the range of the user's dry voice and the range of the user's voice to determine the singing range of the user's audio to be adjusted. The server 104 can also perform vocal detection on the audio to be adjusted to detect the standard degree of the user's pronunciation.
[0127] Before adjusting the audio, the server 104 can perform fundamental frequency detection on the audio to be adjusted to obtain a corresponding fundamental frequency sequence and a word-by-word timestamp. The server 104 can use the fundamental frequency sequence as a carrier and utilize the above-mentioned singing accuracy to determine an adjustment strategy for the audio to be adjusted. The adjustment strategy represents the degree of tuning of the singing skill information in the audio to be adjusted, so that the server 104 can adjust the fundamental frequency sequence of the audio to be adjusted based on the corresponding adjustment strategy to obtain a frequency shift sequence. The server 104 can perform pitch shifting processing on the audio to be adjusted based on the frequency shift sequence to obtain the final tuned audio.
[0128] Through the above embodiment, the server 104 determines the singing level of the audio to be adjusted input by the user and implements different audio adjustment strategies based on the user's singing level, thereby achieving audio adjustment that adapts to the user's level and improving the adjustment effect of the audio adjustment. In addition, when the server 104 detects that the audio to be adjusted is a spoken audio, it can also implement audio adjustment of the audio to be adjusted through audio synthesis, thereby improving the adjustment effect of different types of audio to be adjusted.
[0129] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0130] Based on the same inventive concept, embodiments of the present application also provide an audio adjustment device for implementing the aforementioned audio adjustment method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more of the following embodiments of the audio adjustment device can be found in the aforementioned limitations of the audio adjustment method and will not be further elaborated here.
[0131] In one embodiment, Figure 11 As shown, an audio adjustment device is provided, comprising: a first acquisition module 500, a second acquisition module 502 and an adjustment module 504, wherein:
[0132] The first acquisition module 500 is used to obtain the audio to be adjusted and its corresponding standard audio, determine the singing skill information in the audio to be adjusted, and determine the singing accuracy of the audio to be adjusted based on the standard audio.
[0133] The second acquisition module 502 is used to obtain the singing level corresponding to the audio to be adjusted according to the singing accuracy.
[0134] The adjustment module 504 is used to determine different adjustment strategies for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information based on the singing level, and adjust the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio.
[0135] In one embodiment, the first acquisition module 500 is specifically configured to perform fundamental frequency detection on the audio to be adjusted to obtain a fundamental frequency sequence of the audio to be adjusted; and identify singing technique information in the audio to be adjusted based on the fundamental frequency sequence.
[0136] In one embodiment, the above-mentioned first acquisition module 500 is specifically used to identify at least one singing skill information in the fundamental frequency sequence; obtain the skill type of at least one singing skill information and the time information and number of times the at least one singing skill information appears in the fundamental frequency sequence, and obtain the singing skill information in the audio to be adjusted.
[0137] In one embodiment, the first acquisition module 500 is specifically used to obtain standard melody information of the standard audio; and determine the singing accuracy of the audio to be adjusted based on the matching degree between the standard melody information and the fundamental frequency sequence corresponding to the audio to be adjusted.
[0138] In one embodiment, the first acquisition module 500 is specifically used to obtain, for each line of lyrics in the standard audio, a standard melody template sequence corresponding to the line of lyrics in the standard melody information, and determine the cosine similarity between the standard melody template sequence corresponding to the line of lyrics and the fundamental frequency sequence corresponding to the line of lyrics in the audio to be adjusted; and determine the singing accuracy of the audio to be adjusted based on the average value of the cosine similarities corresponding to multiple lines of lyrics.
[0139] In one embodiment, the above-mentioned second acquisition module 502 is specifically used to determine that the singing level is the first level if the singing accuracy is greater than or equal to the first value; determine the first adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information based on the first level; if the singing accuracy is less than the first value and greater than or equal to the second value, determine that the singing level is the second level; determine the second adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information based on the second level; if the singing accuracy is less than the second value, determine that the singing level is the third level; determine the third adjustment strategy for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information based on the third level; wherein, the first value is greater than the second value; the degree of adjustment of the target audio part under the first adjustment strategy, the second adjustment strategy, and the third adjustment strategy increases successively.
[0140] In one embodiment, the above-mentioned adjustment module 504 is specifically used to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and obtain multiple note-like units in the fundamental frequency sequence; according to the singing level, determine the adjustment strategy for the target audio part corresponding to the singing skill information in the multiple note-like units and the non-target audio part that does not contain the singing skill information.
[0141] In one embodiment, the above-mentioned adjustment module 504 is specifically used to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and obtain multiple note-like units in the fundamental frequency sequence; determine the singing range corresponding to the sequence where the multiple note-like units are located; based on the singing range corresponding to the sequence where the multiple note-like units are located, adjust the audio to be adjusted according to the adjustment strategy to obtain the adjusted audio.
[0142] In one embodiment, the above-mentioned adjustment module 504 is specifically used to identify the target sequence in the fundamental frequency sequence that is an adjacent sentence and is in the same octave interval, and use the octave interval of the target sequence as the singing range of the target sequence during audio adjustment, and use the octave interval corresponding to the standard melody information as the singing range of other sequences other than the target sequence in the fundamental frequency sequence during audio adjustment; determine the sequence type of the sequence where the note-like unit is located, the sequence type includes the target sequence or other sequences, and determine the singing range corresponding to the sequence where the note-like unit is located according to the sequence type.
[0143] In one embodiment, the above-mentioned adjustment module 504 is specifically used to adjust the audio to be adjusted according to the adjustment strategy within the octave range of the singing range corresponding to the target sequence if the sequence where the note-like units are located is the target sequence, so as to obtain the adjusted audio; or, if the sequence where the note-like units are located is other sequences other than the target sequence, adjust the audio to be adjusted according to the adjustment strategy within the octave range corresponding to the standard melody information to obtain the adjusted audio.
[0144] In one embodiment, the above-mentioned adjustment module 504 is specifically used to, if the adjustment strategy is the first adjustment strategy, perform pitch shifting on the target audio parts corresponding to the various singing skill information contained in multiple note-like units based on the first adjustment strategy, and perform pitch shifting and amplitude compression on the non-target audio parts to fit the standard audio, such as a standard melody template sequence corresponding to the standard melody information of the standard audio; if the adjustment strategy is the second adjustment strategy, perform one of the following processings, namely, pitch shifting and amplitude compression on the first singing skill information contained in the audio to be adjusted and / or pitch shifting and amplitude compression on the second singing skill information contained in the audio to be adjusted, and perform pitch shifting and amplitude compression on the non-target audio parts to fit the standard audio; if the adjustment strategy is the third adjustment strategy, perform pitch shifting and amplitude compression on both the target audio parts and the non-target audio parts corresponding to the various singing skill information contained in multiple note-like units based on the third adjustment strategy to fit the standard audio.
[0145] In one embodiment, the above-mentioned first acquisition module 500 is specifically used to obtain the original audio and determine the sound quality score of the original audio based on the noise information of the original audio; if the sound quality score is less than a preset score threshold, the original audio is subjected to noise reduction processing to obtain the audio to be adjusted, and the corresponding standard audio is obtained based on the audio to be adjusted.
[0146] In one embodiment, the above-mentioned device also includes: a classification module, which is used to obtain the fundamental frequency sequence corresponding to the audio to be adjusted, and obtain the fundamental frequency parameters in the fundamental frequency sequence; the fundamental frequency parameters include at least one of the following: range, average pitch and pitch fluctuation sequence; the fundamental frequency parameters are input into a preset classification model, and the audio type of the audio to be adjusted is determined according to the output result of the preset classification model, and the audio type includes reading audio or singing audio; if the audio to be adjusted is singing audio, the steps of determining the singing skill information in the audio to be adjusted and determining the singing accuracy of the audio to be adjusted according to the standard melody information of the standard audio are executed.
[0147] In one embodiment, the above-mentioned device further includes: a reading adjustment module, which is used to synthesize the audio according to the standard audio, the human voice tone information and the human voice pitch information in the audio to be adjusted to obtain the adjusted audio if the audio to be adjusted is a reading audio.
[0148] Each module in the aforementioned audio adjustment device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0149] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 12 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store audio data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an audio adjustment method is implemented.
[0150] Those skilled in the art will understand that Figure 12The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0151] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned audio adjustment method when executing the computer program.
[0152] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned audio adjustment method is implemented.
[0153] In one embodiment, a computer program product is provided, comprising a computer program, which implements the above-mentioned audio adjustment method when executed by a processor.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0155] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0156] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0157] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An audio adjustment method, characterized in that: The method comprises: Acquire the audio to be adjusted and its corresponding standard audio, determine singing skill information in the audio to be adjusted, and determine the singing accuracy of the audio to be adjusted based on the standard audio; According to the singing accuracy, obtaining a singing level corresponding to the audio to be adjusted; wherein, when the singing accuracy is greater than or equal to a first value, it is a first level; when it is less than the first value or equal to a second value, it is a second level; when it is less than the second value, it is a third level; the first value is greater than the second value; Based on the singing level, different adjustment strategies are determined for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information, and the audio to be adjusted is adjusted based on the adjustment strategy to obtain the adjusted audio; wherein, the adjustment degrees of the target audio part for each adjustment strategy determined based on the first level, the second level and the third level respectively increase successively.
2. The method according to claim 1, characterized in that The step of determining the singing skill information in the audio to be adjusted includes: Performing fundamental frequency detection on the audio to be adjusted to obtain a fundamental frequency sequence of the audio to be adjusted; According to the fundamental frequency sequence, singing skill information in the audio to be adjusted is identified.
3. The method according to claim 2, characterized in that The step of identifying singing skill information in the audio to be adjusted according to the fundamental frequency sequence includes: Identifying at least one singing skill information in the fundamental frequency sequence; The skill type of the at least one singing skill information and the time information and the number of times the at least one singing skill information appears in the fundamental frequency sequence are obtained to obtain the singing skill information in the audio to be adjusted.
4. The method according to claim 2, characterized in that Determining the singing accuracy of the audio to be adjusted based on the standard audio includes: Acquiring standard melody information of the standard audio; The singing accuracy of the audio to be adjusted is determined according to the matching degree between the standard melody information and the fundamental frequency sequence of the audio to be adjusted.
5. The method according to claim 4, characterized in that The determining the singing accuracy of the audio to be adjusted according to the matching degree between the standard melody information and the fundamental frequency sequence of the audio to be adjusted includes: For each line of lyrics in the standard audio, obtaining a standard melody template sequence corresponding to the line of lyrics in the standard melody information, and determining a cosine similarity between the standard melody template sequence corresponding to the line of lyrics and a fundamental frequency sequence corresponding to the line of lyrics in the audio to be adjusted; The singing accuracy of the audio to be adjusted is determined based on an average value of cosine similarities corresponding to multiple lyrics.
6. The method according to claim 1, characterized in that The step of obtaining a singing level corresponding to the audio to be adjusted based on the singing accuracy, and determining different adjustment strategies for a target audio portion corresponding to the singing skill information and a non-target audio portion not containing the singing skill information in the audio to be adjusted based on the singing level, includes: If the singing accuracy is greater than or equal to a first value, determining that the singing level is a first level; and determining, based on the first level, a first adjustment strategy for a target audio portion corresponding to the singing skill information and a non-target audio portion not containing the singing skill information in the audio to be adjusted; If the singing accuracy is less than the first value and greater than or equal to the second value, determining that the singing level is the second level; and determining, based on the second level, a second adjustment strategy for the target audio portion corresponding to the singing skill information and the non-target audio portion not containing the singing skill information in the audio to be adjusted; If the singing accuracy is less than the second value, the singing level is determined to be the third level; based on the third level, a third adjustment strategy is determined for the target audio part corresponding to the singing skill information in the audio to be adjusted and the non-target audio part that does not contain the singing skill information.
7. The method according to claim 6, characterized in that The singing skill information includes at least two of vibrato information, glissando information, transition tone information and transition tone information; The adjusting the audio to be adjusted based on the adjustment strategy includes: If the adjustment strategy is the first adjustment strategy, pitch shifting is performed on target audio parts corresponding to various singing technique information contained in the audio to be adjusted, and pitch shifting and amplitude compression are performed on non-target audio parts based on the first adjustment strategy to fit the standard audio; If the adjustment strategy is the second adjustment strategy, performing one of pitch shifting and amplitude compression processing on the first singing skill information contained in the audio to be adjusted and / or performing pitch shifting and amplitude compression processing on the second singing skill information contained in the audio to be adjusted based on the second adjustment strategy, and performing pitch shifting and amplitude compression processing on the non-target audio portion to fit the standard audio; If the adjustment strategy is the third adjustment strategy, based on the third adjustment strategy, pitch shifting and amplitude compression processing are performed on the target audio part and non-target audio part corresponding to the various singing skill information contained in the audio to be adjusted to fit the standard audio.
8. The method according to claim 1, characterized in that The adjusting the audio to be adjusted based on the adjustment strategy to obtain the adjusted audio includes: Acquire a plurality of note-like units in a fundamental frequency sequence corresponding to the audio to be adjusted; Based on the singing range corresponding to the sequence of the multiple note-like units, the audio to be adjusted is adjusted according to the adjustment strategy to obtain the adjusted audio.
9. The method according to claim 8, characterized in that The step of adjusting the audio to be adjusted based on the singing range corresponding to the sequence of the plurality of note-like units according to the adjustment strategy to obtain the adjusted audio includes: If the sequence containing the note-like units is a target sequence, adjusting the audio to be adjusted according to the adjustment strategy within the octave interval of the target sequence to obtain adjusted audio; wherein the target sequence is a sequence of adjacent sentences in the fundamental frequency sequence and in the same octave interval; or, If the sequence in which the note-like units are located is a sequence other than the target sequence, the audio to be adjusted is adjusted according to the adjustment strategy within the octave interval corresponding to the standard melody information to obtain the adjusted audio.
10. The method according to claim 1, characterized in that The step of obtaining the audio to be adjusted includes: Obtaining original audio, and determining a sound quality score of the original audio based on noise information of the original audio; If the sound quality score is less than a preset score threshold, noise reduction processing is performed on the original audio to obtain the audio to be adjusted.
11. The method according to claim 1, characterized in that After obtaining the audio to be adjusted and its corresponding standard audio, the method further includes: Obtaining a fundamental frequency sequence corresponding to the audio to be adjusted, and obtaining fundamental frequency parameters in the fundamental frequency sequence; the fundamental frequency parameters include at least one of the following: a range, an average pitch, and a pitch fluctuation sequence; Inputting the fundamental frequency parameter into a preset classification model, and determining the audio type of the audio to be adjusted according to an output result of the preset classification model, wherein the audio type includes reading audio or singing audio; If the audio to be adjusted is singing audio, the steps of determining singing skill information in the audio to be adjusted and determining the singing accuracy of the audio to be adjusted based on the standard audio are performed.
12. The method according to claim 11, characterized in that After determining the audio type of the audio to be adjusted according to the output result of the preset classification model, the method further includes: If the audio to be adjusted is a reading audio, audio synthesis is performed based on the standard audio and the human voice tonality information and human voice pitch information in the audio to be adjusted to obtain the adjusted audio.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Audio data processing method and device
CN106649643A
Audio output apparatus, synchronization adjustment system, and synchronization adjustment method
JP2006217005A