Audio processing method and device, electronic equipment and storage medium

By extracting the pitch characteristics of the target original audio and predicting the user's singing performance level and adjusting the volume ratio, the problem of unnatural audio after sound editing in the karaoke application is solved, and a more natural and improved singing experience is achieved.

CN119943012AActive Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510146236.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-06
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

After the existing K-song application is used to revise the singing audio and the real vocals, the sound editing effect is not natural.

Method used

By obtaining the target original singing audio segments of the target song, extracting its pitch characteristic information, using the target singing rating model to predict the user's singing performance level, and adjusting the volume ratio according to the level to match the proportions of the singing recording, original singing and accompaniment volume.

Benefits of technology

The natural effect of singing audio is achieved, which avoids unnatural problems after sound modification, and improves the singing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943012A_ABST
    Figure CN119943012A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses an audio processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a target original singing audio segment corresponding to a target song segment in a target song; extracting pitch feature information of the target original singing audio segment; according to the pitch feature information of the target original singing audio segment, through a target singing rating model suitable for the target object, predicting a target singing performance level of the target object for the target song segment; according to the target singing performance level of the target object for the target song segment, determining a target volume proportion suitable for the target object when singing the target song segment; and in response to the singing of the target object to start singing the target song segment, adjusting the volume ratio to the target volume ratio. According to the method, the singing playing effect can be ensured, and the unnatural problem caused by sound correction can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to an audio processing method, device, electronic device and storage medium. Background Art

[0002] In the related art, in karaoke applications, in order to ensure the singing and playback effect, if the user's singing effect is not ideal, the karaoke application will adjust the audio of the user's singing. However, although this method improves the singing and playback effect to a certain extent, the audio after tuning is obviously different from the user's real voice, and the sound after tuning is likely to become unnatural or abrupt. Summary of the invention

[0003] In view of the above problems, the embodiments of the present application propose an audio processing method, device, electronic device and storage medium to improve the above problems.

[0004] According to one aspect of an embodiment of the present application, an audio processing method is provided, including: obtaining a target original singing audio segment corresponding to a target song segment in a target song; extracting pitch feature information of the target original singing audio segment; predicting a target singing performance level of the target object for the target song segment through a target singing rating model applicable to the target object based on the pitch feature information of the target original singing audio segment; determining a target volume ratio applicable to when the target object sings the target song segment based on the target singing performance level of the target object for the target song segment; the target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume; in response to the target object starting to sing the target song segment, adjusting the volume ratio to the target volume ratio.

[0005] According to one aspect of an embodiment of the present application, an audio processing device is provided, including: an acquisition module for acquiring a target original singing audio segment corresponding to a target song segment in a target song; an extraction module for extracting pitch feature information of the target original singing audio segment; a prediction module for predicting a target singing performance level of a target object for the target song segment through a target singing rating model applicable to the target object based on the pitch feature information of the target original singing audio segment; a target volume ratio determination module for determining a target volume ratio applicable to when the target object sings the target song segment based on the target singing performance level of the target object for the target song segment; the target volume ratio refers to the ratio among the singing recording volume, the original singing volume and the accompaniment volume; a ratio adjustment module for adjusting the volume ratio to the target volume ratio in response to the target object starting to sing the target song segment.

[0006] In some embodiments, the audio processing device also includes: a first acquisition module, used to acquire sample singing audio of multiple sample songs sung by the target object and sample original singing audio of each of the sample songs; a segmentation module, used to segment the sample singing audio of the sample songs and the sample original singing audio of the sample songs according to the multiple sample song segments of the sample songs, and obtain multiple first audio segments in the sample singing audio and multiple second audio segments in the sample original singing audio; a pitch feature extraction module, used to extract the first pitch feature information of each of the first audio segments and the second pitch feature information of each of the second audio segments; a singing performance level determination module, used to determine the sample singing performance level of the target object for each of the sample song segments based on the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment corresponding to the same sample song segment; a training module, used to train the target singing rating model based on the second pitch feature information of the multiple second audio segments and the sample singing performance level of the target object for each of the sample song segments.

[0007] In some embodiments, the second pitch feature information includes the feature values ​​of the corresponding second audio segment under multiple pitch features; the training module includes: a first statistical unit, used to count the prior probability of the target object for each singing performance level according to the sample singing performance level of the target object for multiple sample song segments; a second statistical unit, used to count the conditional probability of each singing performance level for the target object under the feature value of each pitch feature according to the feature values ​​of multiple second audio segments under multiple pitch features and the sample singing performance level of the target object for multiple sample song segments; a model determination unit, used to determine the target singing rating model according to the prior probability of the target object for each singing performance level and the conditional probability of each singing performance level under the feature value of each pitch feature.

[0008] In some embodiments, the pitch feature information of the target original singing audio segment includes the feature values ​​of the target original singing audio segment under multiple pitch features; the prediction module is used to: the target singing rating model is processed according to the following process to predict the target singing performance level of the target object for the target song segment: according to the prior probability of the target object for each singing performance level and the conditional probability of the feature values ​​of each singing performance level under the multiple pitch features corresponding to the target original singing audio segment, a posterior probability calculation is performed to determine that the target singing performance level of the target object for the target song segment is the posterior probability of each singing performance level; the singing performance level with the largest posterior probability is used as the target singing performance level of the target object for the target song segment.

[0009] In other embodiments, the training module includes: a prediction unit, used to predict the singing performance rating of the target singing rating model based on the second pitch feature information of each second audio segment, and output the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment; a prediction loss calculation unit, used to calculate the prediction loss based on the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment, and the sample singing performance level of the target object for each sample song segment; a parameter adjustment unit, used to adjust the parameters of the target singing rating model according to the prediction loss until the training end condition is reached.

[0010] In some embodiments, the first pitch feature information includes feature values ​​of the first audio segment under multiple pitch features; the second pitch feature information includes feature values ​​of the second audio segment under multiple pitch features; the singing performance level determination module includes: a feature value deviation calculation unit, used to determine the feature value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under each of the pitch features according to the feature values ​​of the first audio segment under multiple pitch features and the feature values ​​of the second audio segment under multiple pitch features; a sample singing performance level determination unit, used to determine the sample singing performance level of the target object for each of the sample song segments according to the feature value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under multiple pitch features.

[0011] In some embodiments, the target volume ratio determination module includes: a default volume ratio acquisition unit, used to obtain a default volume ratio; an adjustment strategy determination unit, used to determine a target volume ratio adjustment strategy that matches the target singing performance level of the target object for the target song segment; and a volume ratio adjustment unit, used to adjust the default volume ratio according to the target volume ratio adjustment strategy to obtain a target volume ratio applicable to the target object when singing the target song segment.

[0012] In some embodiments, the target volume ratio adjustment strategy is used to adjust the proportion of one of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio, or to adjust the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

[0013] In some embodiments, the adjustment strategy determination unit is used to: if the target singing performance level of the target object for the target song segment is higher than the first singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: maintain the default volume ratio unchanged; if the target singing performance level of the target object for the target song segment is not higher than the first singing performance level, and is higher than the second singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, maintain the proportion of the original singing volume unchanged, and increase the proportion of the accompaniment volume and reduce the proportion of the singing recording volume; if the target singing performance level of the target object for the target song segment is not higher than the first singing performance level, The second singing performance level is higher than the third singing performance level, and the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, increase the proportion of the original singing volume, and reduce the proportion of the accompaniment volume and reduce the proportion of the singing recording volume; if the target singing performance level of the target object for the target song segment is not higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, maintain the proportion of the singing recording volume unchanged, reduce the proportion of the accompaniment volume and increase the proportion of the original singing volume; wherein, the first singing performance level > the second singing performance level > the third singing performance level, and the singing performance level is positively correlated with the singing effect.

[0014] In other embodiments, the target volume ratio determination module is used to: based on the correspondence between the singing performance level and the volume ratio, use the volume ratio corresponding to the target singing performance level of the target object for the target song segment as the target volume ratio applicable to the target object when singing the target song segment; wherein, in the correspondence between the singing performance level and the volume ratio, the higher the singing performance level, the lower the proportion of the original singing volume in the corresponding volume ratio; the singing performance level is positively correlated with the singing effect.

[0015] In some embodiments, the pitch feature information includes feature values ​​of the target original singing audio segment under multiple pitch features; the target original singing audio segment includes multiple target audio frames; the extraction module includes: a fundamental frequency detection unit, used to perform fundamental frequency detection on each target audio frame in the target original singing audio segment, and determine the fundamental frequency of each target audio frame; a pitch conversion unit, used to perform pitch conversion on the fundamental frequency of each target audio frame to obtain the pitch value of each target audio frame; a feature value determination unit, used to determine the feature value of the target original singing audio segment under multiple pitch features according to the pitch values ​​of multiple target audio frames in the target original singing audio segment.

[0016] In some embodiments, the plurality of pitch features include at least two of the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the preceding silent pause.

[0017] In some embodiments, the audio processing device also includes: an audio synthesis module, which is used to perform audio synthesis on the original audio segment corresponding to the target song segment, the corresponding accompaniment audio segment, and the singing audio segment collected during the target object singing the target song segment according to the target volume ratio to obtain a synthesized audio segment corresponding to the target song segment; and a playback module, which is used to play the synthesized audio segment.

[0018] According to one aspect of an embodiment of the present application, an electronic device is provided, including: a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the above audio processing method is implemented.

[0019] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the above audio processing method is implemented.

[0020] According to one aspect of an embodiment of the present application, a computer program product is provided, including computer instructions, which implement the above audio processing method when executed by a processor.

[0021] In the present application, the target singing rating model applicable to the target object is used to predict the target singing performance level of the target object for the target song segment according to the pitch feature information of the target original singing audio segment; and according to the target singing performance level of the target object for the target song segment, the target volume ratio applicable to the target object when singing the target song segment is determined; and when the target object starts singing the target song segment, the volume ratio is adjusted to the target volume ratio. Through the scheme of the present application, the target singing performance level when the target object sings the target song segment is predicted in advance, and the volume ratio is automatically adjusted to the target volume ratio that matches the target singing performance level, and the target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume. By adjusting the ratio between the singing recording volume, the original singing volume and the accompaniment volume, the shortcomings of the target object when singing are covered up. Through the method of the present application, there is no need to repair the singing audio of the target object, and there will be no problem of unnatural singing effect caused by excessive or improper repair. Moreover, after adjusting the volume ratio, the singing playback effect of the subsequent singing audio can be guaranteed, and the auditory experience can be guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1A It is a schematic diagram of an application scenario of the present application according to an embodiment of the present application.

[0024] Figure 1B This is an interactive schematic diagram of singing in a client according to an embodiment of the present application.

[0025] Figure 2 is a flowchart of an audio processing method according to an embodiment of the present application.

[0026] Figure 3 is a flow chart of step 240 according to an embodiment of the present application.

[0027] Figure 4 It is a schematic diagram of an automatic adjustment setting interface of a proportion in a client according to an embodiment of the present application.

[0028] Figure 5 It is a flowchart of a training target singing rating model according to one embodiment of the present application.

[0029] Figure 6 It is a schematic diagram of step 540 according to an embodiment of the present application.

[0030] Figure 7 It is a flowchart of step 550 according to an embodiment of the present application.

[0031] Figure 8 It is a flowchart of step 230 according to an embodiment of the present application.

[0032] Fig. 9 is a flow chart of step 550 according to another embodiment of the present application.

[0033] Fig.10 is a flowchart of an audio processing method according to another embodiment of the present application.

[0034] Fig.11 is a block diagram of an audio processing device according to an embodiment of the present application.

[0035] Fig.12 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION

[0036] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0037] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0038] In the following description, the terms "first\second" and the like are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] The "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the following description, it involves "some embodiments or some embodiments", which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0040] Figure 1A It is a schematic diagram of an application scenario of the present application according to an embodiment of the present application. As shown in FIG1 , the application scenario includes a terminal 110 and a music database 120. The music database 120 can store audio data of songs. The audio data of a song can include the original audio of the song and the accompaniment audio of the song. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a car terminal, a smart TV, a smart speaker, an extended display device, a wearable device, and the like.

[0041] like Figure 1AAs shown, the terminal 110 can run the client of the application, which can be a karaoke application, or other entertainment applications integrated with the karaoke function (such as music applications, live broadcast applications, video applications), or other social applications integrated with the karaoke function (such as instant messaging applications, content sharing applications). In other words, the application can be any application integrated with the karaoke function, and the karaoke function can be the main function of the application or an additional function of the application in addition to the main function. The client can be a desktop client, a mobile client, or a mini-program client. The client can obtain the audio data of the target song from the music database 120, and then obtain the target original singing audio segment corresponding to the target song segment from the audio data of the target song, and then extract the pitch feature of the target original singing audio segment to obtain the pitch feature information of the target original singing audio segment; then, according to the pitch feature information of the target original singing audio segment, the singing performance level is predicted through the target singing rating model applicable to the target object, and the target singing performance level of the target object for the target song segment is obtained; and according to the target singing performance level of the target object for the target song segment, the target volume ratio applicable to the target object when singing the target song segment is determined; the target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume; in response to the target object starting to sing the target song segment, the volume ratio is adjusted to the target volume ratio.

[0042] Figure 1B FIG. 1 is a schematic diagram of an interactive singing session in a client according to an embodiment of the present application. Figure 1B As shown, you can click Figure 1B The "Live Broadcast" control in the public account page shown in (1) shows Figure 1B In the live content selection window shown in (2), if you click Figure 1B The "Live" option in the live content selection window shown in (2) can display Figure 1B The karaoke page shown in (3) above will be displayed; if you click Figure 1B The "Voice Karaoke Room" control in the Karaoke entry page shown in (3) can display Figure 1B If you click the "Karaoke Room" option in the karaoke mode selection window, you can display the karaoke mode selection window. Figure 1B The karaoke page shown in (5) allows you to perform karaoke live with one or more other users.

[0043] When multiple users participate in a live karaoke session, different users are located in different clients. For each client, the target volume ratio applicable to the user of each client when singing the corresponding song segment can be determined according to the method of the present application, and when the user on the client side starts to sing the corresponding song segment, the corresponding volume ratio can be adjusted to the determined applicable target volume ratio.

[0044] The method of the present application is not limited to being executed by a terminal, but may also be executed by a server corresponding to the application, or implemented by the interaction between the terminal where the client is located and the server, which is not specifically limited here.

[0045] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0046] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0047] Figure 2 1 is a flowchart of an audio processing method according to an embodiment of the present application. The method can be executed by an electronic device with processing capabilities, such as a terminal, a server, etc., which is not specifically limited here. Figure 2 As shown, the method at least includes steps 210 to 250, which are described in detail as follows:

[0048] Step 210, obtaining a target original singer audio segment corresponding to a target song segment in a target song.

[0049] The target song may be a song that the target object is currently singing or a song to be sung. The target song segment in the target song refers to any song segment in the target song, for example, it may be a song segment to be sung by the target object in the target song. In some embodiments, the target song may be divided into a plurality of song segments of equal duration according to the duration of the target song, and the durations of different song segments are equal, for example, 0 seconds to 5 seconds of the target song are regarded as the first song segment, 5 seconds to 10 seconds are regarded as the second song segment, and so on.

[0050] In other embodiments, song data of a target song can be obtained from a song database. In addition to basic information of the target song (such as lyrics, composer, singer, arranger, song title, lyrics, etc.), the song data also includes a target original singer audio of the target song and a target accompaniment audio of the target song. The target original singer audio can be segmented according to the target original singer audio of the target song and the pitch values ​​of each original singer audio frame in the target original singer audio to obtain a plurality of original singer audio segments, and the time period of an original singer audio segment in the target song is taken as a song segment. Among them, an original singer audio segment includes a plurality of continuous original singer audio frames, and the pitch value difference (i.e., the absolute value of the pitch difference) of different original singer audio frames in an original singer audio segment does not exceed the first threshold value. For example, if the original audio frames in the target original audio include audio frames 1 to 100 from first to last, the pitch value difference between any two audio frames in audio frame 1 to audio frame 50 does not exceed the first threshold, and the pitch value difference between audio frame 50 and audio frame 51 exceeds the first threshold, thus, audio frames 1 to 50 are divided into an original audio segment, the timestamp of audio frame 1 in the target original audio is time point A, and the timestamp of audio frame 50 in the target original audio is time point B, thus, the section from time point A to time point B in the target song is taken as a song segment.

[0051] Step 220: extracting pitch feature information of the target original singer audio segment.

[0052] In some embodiments, the pitch feature information of the target original singing audio segment may be a pitch value sequence formed by the pitch values ​​of all the original singing audio frames in the target original singing audio segment.

[0053] In other embodiments, the pitch feature information may include feature values ​​of the target original audio segment under one pitch feature or multiple pitch features. The pitch feature may be the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the preceding silent pause.

[0054] It is understandable that if the pitch feature information of the target original singer audio segment includes the feature values ​​of the target original singer audio segment under multiple pitch features, correspondingly, the multiple pitch features include at least two of the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the previous silent pause. It is understandable that the more pitch features involved in the pitch feature information, the more comprehensive the pitch characteristics of the target original singer audio segment reflected by the pitch feature information, which can ensure that the subsequent predicted singing performance is more accurate.

[0055] The characteristic value of an original audio segment at the highest pitch value may be the value of the highest pitch value in the original audio segment, or may be the first pitch level mapped to the value of the highest pitch value in the original audio segment determined according to the mapping relationship between the set value of the highest pitch value and the first pitch level. The mapping relationship between the value of the highest pitch value and the second pitch level may be: the larger the value of the highest pitch value, the higher the first pitch level. For example, four first pitch levels can be set, namely, first pitch level I, first pitch level II, first pitch level III and first pitch level IV. When the numerical value of the highest pitch value is less than 65, the first pitch level mapped by the numerical value of the highest pitch value is the first pitch level I; when the numerical value of the highest pitch value ∈ [65, 70], the first pitch level mapped by the numerical value of the highest pitch value is the first pitch level II; when the numerical value of the highest pitch value ∈ [71, 75], the first pitch level mapped by the numerical value of the highest pitch value is the first pitch level III; when the numerical value of the highest pitch value is greater than 75, the first pitch level mapped by the numerical value of the highest pitch value is the first pitch level IV.

[0056] Similarly, the characteristic value of an original audio segment at the lowest pitch value can be the value of the lowest pitch value in the original audio segment, or it can be the second pitch level mapped to the value of the lowest pitch value in the original audio segment determined based on the mapping relationship between the value of the lowest pitch value and the second pitch level. The mapping relationship between the value of the lowest pitch value and the second pitch level can be: the larger the value of the lowest pitch value, the higher the second pitch level. For example, four second pitch levels can be set, namely second pitch level I, second pitch level II, second pitch level III and second pitch level IV. When the value of the lowest pitch value is less than 50, the second pitch level mapped by the value of the lowest pitch value is the second pitch level I; when the value of the lowest pitch value ∈[50, 55], the second pitch level mapped by the value of the lowest pitch value is the second pitch level II; when the value of the lowest pitch value ∈[56, 60], the second pitch level mapped by the value of the lowest pitch value is the second pitch level III; when the value of the lowest pitch value is greater than 60, the second pitch level mapped by the value of the lowest pitch value is the second pitch level IV.

[0057] The maximum stable duration of the pitch of an original singing audio segment refers to the maximum number of consecutive frames of the original singing audio frames whose pitch value difference does not exceed the second threshold in an original singing audio segment. The second threshold is less than the first threshold, for example, the second threshold is 1. For example, assuming that the second threshold is 1, if an original singing audio segment includes a total of 50 original singing audio frames, the absolute value of the pitch difference between any two adjacent original singing audio frames in original singing audio frame 1 to original singing audio frame 10 does not exceed 1, the absolute value of the pitch difference between any two adjacent original singing audio frames in original singing audio frame 10 to original singing audio frame 15 exceeds 1, the absolute value of the pitch difference between any two adjacent original singing audio frames in original singing audio frame 15 to original singing audio frame 35 does not exceed 1, and the absolute value of the pitch difference between any two adjacent original singing audio frames in original singing audio frame 35 to original singing audio frame 50 exceeds 1. Therefore, the maximum stable duration of the pitch corresponding to the original singing audio segment is 21.

[0058] Similarly, the characteristic value of an original singing audio segment under the maximum stable duration of the pitch can be the value of the maximum stable duration corresponding to the original singing audio segment, or it can be the stability level mapped to the value of the maximum stable duration in the original singing audio segment determined according to the mapping relationship between the set maximum stable duration value and the stability level. The mapping relationship between the maximum stable duration value and the stability level can be: the larger the maximum stable duration value, the higher the stability level. For example, four stability levels are set, namely stability level 1, stability level 2, stability level 3 and stability level 4. When the value of the maximum stable duration in an original singing audio segment is less than two original singing audio frames, the stability level mapped by the value of the maximum stable duration of the pitch is stability level 1; when the value of the maximum stable duration in an original singing audio segment is between 3 and 5 original singing audio frames, the stability level mapped by the value of the maximum stable duration of the pitch is stability level 2; when the value of the maximum stable duration in an original singing audio segment is between 6 and 10 original singing audio frames, the stability level mapped by the value of the maximum stable duration of the pitch is stability level 3; when the value of the maximum stable duration in an original singing audio segment is greater than 10 original singing audio frames, the stability level mapped by the value of the maximum stable duration of the pitch is stability level 4.

[0059] The pitch variance value of an original singing audio segment refers to the variance of the pitch values ​​of all original singing audio frames in the original singing audio segment. Similarly, the characteristic value of an original singing audio segment under the pitch variance value can be the numerical value of the pitch variance value corresponding to the original singing audio segment, or it can be the variance level mapped to the numerical value of the pitch variance value in the original singing audio segment determined according to the mapping relationship between the numerical value of the pitch variance value and the variance level. The mapping relationship between the numerical value of the pitch variance value and the variance level can be: the larger the numerical value of the pitch variance value, the higher the variance level. For example, four variance levels are set, namely variance level 1, variance level 2, variance level 3 and variance level 4. In an original audio segment, the value of the pitch variance value does not exceed 3, and the variance level mapped by the value of the pitch variance value is variance level 1; when the value of the pitch variance value in an original audio segment is ∈(3, 5], the variance level mapped by the value of the pitch variance value is variance level 2; when the value of the pitch variance value in an original audio segment is ∈(5, 10], the variance level mapped by the value of the pitch variance value is variance level 3; when the value of the pitch variance value in an original audio segment is greater than 10, the variance level mapped by the value of the pitch variance value is variance level 4.

[0060] The preceding silent pause duration of an original singing audio segment refers to the number of consecutive silent original singing audio frames before the first original singing audio frame in the original singing audio segment. Among them, whether an original singing audio frame is a silent original singing audio frame can be determined according to the amplitude of the original singing audio frame. If the amplitude of an original singing audio frame is less than the amplitude threshold (the amplitude threshold is a non-negative number), it can be determined that the original singing audio frame is a silent original singing audio frame. On the contrary, if the amplitude of an original singing audio frame is not less than the amplitude threshold, it can be determined that the original singing audio frame is not a silent original singing audio frame.

[0061] Similarly, the characteristic value of the preceding silent pause duration of an original singing audio segment may be the numerical value of the preceding silent pause duration corresponding to the original singing audio segment, or may be the pause level mapped to the numerical value of the preceding silent pause duration determined according to the mapping relationship between the numerical value of the preceding silent pause duration and the pause level. The mapping relationship between the numerical value of the preceding silent pause duration and the pause level may be: the larger the numerical value of the preceding silent pause duration, the higher the pause level.

[0062] For example, four pause levels can be set, namely pause level 1, pause level 2, pause level 3 and pause level 4. In an original audio segment, when the value of the preceding silent pause duration is less than 2, the pause level mapped by the value of the preceding silent pause duration of the pitch is pause level 1; when the value of the preceding silent pause duration in an original audio segment is ∈[3, 8], the pause level mapped by the value of the preceding silent pause duration of the pitch is pause level 2; when the value of the preceding silent pause duration in an original audio segment is ∈[9, 15], the pause level mapped by the value of the preceding silent pause duration of the pitch is pause level 3; when the value of the preceding silent pause duration in an original audio segment is greater than 15, the pause level mapped by the value of the preceding silent pause duration of the pitch is pause level 4.

[0063] It is worth mentioning that the mapping relationship between the numerical value of the highest pitch value and the first pitch level, the mapping relationship between the numerical value of the lowest pitch value and the second pitch level, the mapping relationship between the numerical value of the maximum stable duration and the stability level, the mapping relationship between the numerical value of the pitch variance value and the variance level, and the mapping relationship between the numerical value of the previous silent pause duration and the pause level are merely illustrative examples and cannot be regarded as limiting the scope of use of this application.

[0064] In some embodiments, the pitch feature information of the target original singer audio segment includes feature values ​​of the target original singer audio segment under multiple pitch features. Step 220 includes the following ①-③:

[0065] ① Perform fundamental frequency detection on each target audio frame in the target original audio segment to determine the fundamental frequency of each target audio frame.

[0066] In the present application, the original audio frame in the target original audio segment is called the target audio frame. The target original audio segment can be divided into multiple continuous target audio frames according to equal time lengths. For example, the duration of each target audio frame can be 100ms.

[0067] The fundamental frequency of a target audio frame is equal to the inverse of the fundamental pitch period of the target audio frame. The fundamental pitch period of a target audio frame refers to the longest repetition period in the target audio frame (equivalent to a short segment of the sound signal). In some embodiments, the fundamental pitch period of the target audio frame can be calculated by an autocorrelation method or a Fourier transform method.

[0068] In a specific embodiment, considering the presence of various noise interferences in the sound signal, the Fourier transform method may be greatly affected by noise in a low signal-to-noise ratio environment, and the noise may cause the spectrum peaks in the spectrum to become blurred, thereby affecting the accurate estimation of the pitch period. Therefore, in order to reduce the accuracy of the pitch period calculation affected by noise, the autocorrelation method may be used to calculate the pitch period of the target audio frame.

[0069] The autocorrelation method is to find the fundamental period by calculating the similarity between the sound signal and its delayed version (i.e., the short-time autocorrelation function). For a discrete sound signal x[n], its autocorrelation function Rxx[k] is defined as:

[0070] Rxx[k] = Σx[n] * x[nk]; (Formula 1)

[0071] Here, k is the delay, which can also be understood as the number of sampling points between x[nk] and x[n].

[0072] Afterwards, the point k=0 is removed from the autocorrelation function, and the delay k corresponding to the maximum value of the autocorrelation function is found. The determined k is the pitch period.

[0073] In addition, considering that the target original audio segment (or the target original audio) is sampled according to a preset sampling rate, that is, the target audio frame includes multiple sampling points sampled according to the sampling rate, assuming that the sampling rate of the target audio frame is fs, according to the fundamental frequency period k determined by the above formula 1, the fundamental frequency F of the target audio frame can be determined according to the following formula 2:

[0074] F = fs / k; (Formula 2)

[0075] ② Perform pitch conversion on the fundamental frequency of each target audio frame to obtain the pitch value of each target audio frame.

[0076] The pitch value T of the target audio frame may be determined according to the fundamental frequency of the target audio frame according to the following formula 3:

[0077] T = 12 × log2( F / F1) + 69; (Formula 3)

[0078] Wherein, F1 is a reference base frequency, for example, F1 may be 440.

[0079] ③ According to the pitch values ​​of multiple target audio frames in the target original singing audio segment, determine the feature value of the target original singing audio segment under multiple pitch features.

[0080] After determining the pitch value of each target audio frame in the target original singer audio segment, the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the pitch features in the preceding silent pause duration can be determined accordingly, and then the characteristic value of the target original singer audio segment under each pitch feature can be determined.

[0081] Step 230, based on the pitch feature information of the target original singer's audio segment, a target singing performance level of the target subject for the target song segment is predicted by a target singing rating model applicable to the target subject.

[0082] The singing rating model suitable for performing singing rating on the target object is called the target singing rating model. The target singing rating model can be obtained by training the sample singing audio of the target object singing multiple sample songs and the sample original singing audio of each sample song. The sample singing audio of the sample song refers to the audio recorded by the target object during the singing of the sample song. The singing rating model can be a Bayesian classification model, an SVM (Support Vector Machine) model, or a deep learning network model, which is not specifically limited here.

[0083] Since the target singing rating model can be trained by the target object singing the sample singing audio of multiple sample songs and the sample original singing audio of each sample song, during the training process, the target singing rating model can learn the target object's singing performance for different song segments with different pitch features. The target singing performance level of the target object for the target song segment is the predicted singing effect performance when the target object sings the target song segment, that is, the target singing performance level reflects the singing effect when the target object sings the target song segment.

[0084] Step 240, based on the target singing performance level of the target object for the target song segment, determine the target volume ratio applicable to the target object when singing the target song segment; the target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume.

[0085] The singing recording volume refers to the volume of the audio recorded when the subject sings. The original singing volume refers to the playback volume set for the original singing audio. The accompaniment volume refers to the playback volume set for the accompaniment audio.

[0086] It can be understood that if the target singing performance level of the target object for the target song segment indicates that the target object's singing effect when singing the target song segment is not good, if only the audio recorded by the target object during the singing of the target song segment is played, the listening effect will be poor for the listener; if the target singing performance level indicates that the target object's singing effect when singing the target song segment is good, if only the audio recorded by the target object during the singing of the target song segment is played, the listening effect will be good and the listening experience will be better for the listener.

[0087] Therefore, in the present application, in order to ensure the auditory effect during the playback of the singing audio recorded for the object, when it is predicted that the singing effect of the object is not good, the volume ratio can be adjusted in advance to achieve the purpose of adjusting the proportion of the volume of different audio components in the three audio components in the synthesized audio (i.e., singing audio component, original singing audio component and accompaniment audio component). In this way, when the synthesized audio is played, there is a difference in the volume of the three audio components in the synthesized audio. For example, if the adjusted volume ratio is: singing recording volume: original singing volume: accompaniment volume = k1: k2: k3, then in the three audio components in the synthesized audio, the singing recording volume: original singing volume: accompaniment volume = k1: k2: k3, assuming that the set playback volume on the playback side is B, then when the synthesized audio is played, theoretically the playback volume of the singing audio component is The playback volume of the original audio component is The playback volume of the accompaniment audio component is By adjusting the volume ratio, when playing the synthesized audio, the listener perceives different volumes of different audio components, or in other words, the proportions of the volumes of different audio components perceived by the listener in the overall playback volume are different.

[0088] Therefore, when the subject's singing effect is not good, the volume ratio can be adjusted in advance so that in the synthesized audio perceived by the listener, the singing recording volume can be lower than the volume of the audio played when the volume ratio is not adjusted, while the accompaniment volume and / or the original singing volume of the original audio can be higher than the volume of the audio played when the volume ratio is not adjusted, so as to ensure that the overall auditory effect is better.

[0089] In the volume ratio, the sum of the proportion of the singing recording volume, the proportion of the original singing volume and the proportion of the accompaniment volume is 1. The target volume ratio refers to the volume ratio applicable to the target object singing the target song segment. Among them, the volume ratio applicable to the target object singing the target song segment refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume when the auditory effect during the recording of the singing audio segment recorded when the target object sings the target song segment is better.

[0090] It can be understood that the proportion of the singing recording volume in the target volume ratio is related to the singing effect indicated by the target singing performance level, that is, when the singing effect indicated by the target singing performance level is poor, the proportion of the original singing volume and the proportion of the accompaniment volume in the target volume ratio can be higher as a whole; when the singing effect indicated by the target singing performance level is better, the proportion of the original singing volume and the proportion of the accompaniment volume in the target volume ratio can be lower as a whole.

[0091] In some embodiments, step 240 includes: according to the correspondence between the singing performance level and the volume ratio, using the volume ratio corresponding to the target singing performance level of the target object for the target song segment as the target volume ratio applicable to the target object singing the target song segment; wherein, in the correspondence between the singing performance level and the volume ratio, the higher the singing performance level, the lower the proportion of the original singing volume in the corresponding volume ratio; the singing performance level is positively correlated with the singing effect.

[0092] The higher the singing performance level, the better the singing effect. When the singing effect is better, the proportion of the original singing volume in the volume ratio is lower, so that the proportion of the original singing volume in the final audio played is smaller, and the proportion of the target object's singing audio volume in the final audio played is higher, which can ensure the target object's sense of participation in the singing. Moreover, due to the good singing effect, the final audio played also gives the listener a higher auditory experience. When the singing effect is worse, the proportion of the original singing volume in the volume ratio is higher, so that the proportion of the original singing volume in the final audio played is higher. In this way, the original singing is used in the final audio played to cover up the shortcomings of the target object in singing, ensuring the overall auditory experience of the listener.

[0093] That is to say, the volume ratio applicable to each singing performance level can be set in advance. In this way, after determining the target singing performance level of the target object for the target song segment, the volume ratio applicable to the target singing performance level can be used as the target volume ratio applicable to the target object when singing the target song segment.

[0094] Table 1 below is a corresponding relationship between singing performance level and volume ratio according to an embodiment of the present application. In Table 1 below, singing performance level I is higher than singing performance level II, singing performance level II is higher than singing performance level III, singing performance level III is higher than singing performance level IV, singing performance level IV is higher than singing performance level V. In other words, the singing effect indicated by singing performance level I is better than the singing effect indicated by singing performance level II, the singing effect indicated by singing performance level II is better than the singing effect indicated by singing performance level III, and so on. The corresponding relationship shown in Table 1 is only an illustrative example and cannot be considered as a limitation on the scope of use of the present application.

[0095] Singing performance level Volume ratio (singing recording volume: accompaniment volume: original singing volume) Singing Performance Level I, II 0.5:0.5:0 Vocal Performance Level III 0.4:0.6:0 Vocal Performance Level IV 0.45:0.4:0.15 Singing Performance Level V 0.5:0.2:0.3

[0096] Table 1

[0097] For example, if the singing performance level is set to 5 in total, namely: excellent, good, average, slightly poor and very poor, on this basis, the singing performance level I above can be excellent, the singing performance level II can be good, the singing performance level III can be average, the singing performance level IV can be slightly poor, and the singing performance level V can be very poor.

[0098] In other embodiments, Figure 3 As shown, step 240 includes steps 310 to 330:

[0099] Step 310, obtaining a default volume ratio.

[0100] The default volume ratio in the client may be the volume ratio set by default in the client. It may also be the volume ratio set by the target object in the interactive interface of the client before singing the song. The default volume ratio may be: singing recording volume: accompaniment volume: original singing volume = 0.5:0.5:0. Of course, it is not limited to this and may be other ratios.

[0101] In the karaoke function provided by the client, since the karaoke function provided by the client is mainly to provide users with lyrics and accompaniment to assist users in singing, the proportion of the original singing volume in the default volume ratio set in the client can be 0.

[0102] Step 320, based on the target singing performance level of the target object for the target song segment, determine a target volume ratio adjustment strategy that matches the target singing performance level.

[0103] The correspondence between the singing performance level and the volume ratio adjustment strategy can be set in advance, with one singing performance level corresponding to one volume ratio adjustment strategy. In this way, after determining the target singing performance level, the volume ratio adjustment strategy corresponding to the target singing performance level can be determined as the target volume ratio adjustment strategy that matches the target singing performance level.

[0104] One singing performance level corresponds to one volume ratio adjustment strategy for adjusting the default volume ratio. One volume ratio adjustment strategy can set the adjustment amount and direction of the proportion of each audio component volume (i.e. singing recording volume, original singing volume, accompaniment volume), and the adjustment direction includes reduction or increase. Among them, when the singing effect indicated by a singing performance level is poor, the volume ratio adjustment strategy corresponding to the singing performance level at least indicates that the proportion of the original singing volume should be increased on the basis of the default volume ratio. Of course, in the volume ratio adjustment strategies corresponding to different singing performance levels, the adjustment amount of the proportion of the same audio component volume may be different.

[0105] In some embodiments, step 320 includes the following 1)-4):

[0106] 1) If the target singing performance level of the target object for the target song segment is higher than the first singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: maintaining the default volume ratio unchanged.

[0107] The singing effect indicated by the first singing performance level is better. When the target singing performance level is not lower than the first singing performance level, it means that the singing effect indicated by the target singing performance level is not worse than the singing effect indicated by the first singing performance level, indicating that the singing effect of the target object singing the target song segment may be better. Therefore, the default volume ratio is maintained and no adjustment is needed.

[0108] 2) If the target singing performance level of the target object for the target song segment is not higher than the first singing performance level, and is higher than the second singing performance level, the target volume ratio adjustment strategy that matches the target singing performance level is: on the basis of the default volume ratio, the proportion of the original singing volume remains unchanged, and the proportion of the accompaniment volume is increased and the proportion of the singing recording volume is reduced.

[0109] The singing effect indicated by the first singing performance level is better than the singing effect indicated by the second singing performance level. When the target singing performance level is not higher than the first singing performance level, and higher than the second singing performance level, the default volume ratio is slightly intervened. On the basis of the default volume ratio, the proportion of the original singing volume is maintained unchanged, and the proportion of the accompaniment volume is increased and the proportion of the singing recording volume is reduced, which is equivalent to increasing the proportion of the accompaniment volume to cover up the shortcomings of the target object's singing audio and ensure the subsequent overall playback effect.

[0110] 3) If the target singing performance level of the target object for the target song segment is not higher than the second singing performance level and higher than the third singing performance level, the target volume ratio adjustment strategy that matches the target singing performance level is: on the basis of the default volume ratio, increase the proportion of the original singing volume, and reduce the proportion of the accompaniment volume and the proportion of the singing recording volume.

[0111] The singing effect indicated by the second singing performance level is better than the singing effect indicated by the third singing performance level. When the target singing performance level is not higher than the second singing performance level and higher than the third singing performance level, it indicates that the singing effect of the target object for the target song segment is better than the singing effect indicated by the third singing performance level, but not more than the singing effect indicated by the second singing performance level. In this case, only increasing the proportion of the accompaniment volume may still not make the overall playback effect good enough. Therefore, the default volume ratio is moderately intervened to increase the proportion of the original singing volume to introduce the original singing audio to ensure the overall playback effect. It is understandable that the introduction of the original singing volume and the accompaniment volume of the same proportion has a greater intervention effect on the introduction of the original singing volume, and its masking effect on covering up the deficiencies in the singing audio is greater.

[0112] 4) If the target singing performance level of the target object for the target song segment is not higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, maintain the proportion of singing recording volume unchanged, reduce the proportion of accompaniment volume and increase the proportion of original singing volume. Among them, the first singing performance level> the second singing performance level>

[0113] The third is the singing performance level. The singing performance level is positively correlated with the singing effect.

[0114] When the target singing performance level is not higher than the third singing performance level, it indicates that the target object's singing effect for the target song segment is not as good as the singing effect indicated by the third singing performance level. Therefore, the default volume ratio is heavily intervened. On the basis of the default volume ratio, the proportion of the singing recording volume is maintained unchanged, the proportion of the accompaniment volume is reduced, and the proportion of the original singing volume is increased.

[0115] In some embodiments, in the case of 4) above, the adjustment amount of the proportion of the original singing volume specified in the corresponding target volume ratio adjustment strategy can be greater than the adjustment amount of the proportion of the original singing volume specified in the corresponding target volume ratio adjustment strategy in 3) above, so that after the default volume ratio is adjusted according to the target volume ratio adjustment strategy corresponding to 4) above, the proportion of the original singing volume is greater than the proportion of the original singing volume after the default volume ratio is adjusted according to the target volume ratio adjustment strategy corresponding to 3) above.

[0116] Step 330, according to the target volume ratio adjustment strategy, the default volume ratio is adjusted to obtain a target volume ratio suitable for the target object to sing the target song segment.

[0117] The target volume ratio adjustment strategy is used to adjust the proportion of one of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio, or to adjust the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

[0118] In some embodiments, the sum of the proportions of the three volume components in the volume ratio can be set not to exceed 1, that is, the sum of the proportions of the three audio components in the volume ratio can be less than 1. In this case, the target volume ratio adjustment strategy can be limited to adjusting the proportion of one of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio. Of course, the limit can also be to adjust the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

[0119] In some embodiments, the volume ratio can be set so that the sum of the proportions of the three audio components is always 1. In this way, the target volume ratio adjustment strategy is limited to adjusting the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

[0120] Table 2 below shows the corresponding relationship between the singing performance level and the volume ratio adjustment strategy according to an embodiment of the present application, and shows the volume ratio before and after the volume ratio is adjusted according to the corresponding volume ratio adjustment strategy. The volume ratios in Table 2 are all singing recording volume: accompaniment volume: original singing volume.

[0121]

[0122] Table 2

[0123] In the above embodiment, the volume ratio adjustment strategies corresponding to the singing performance levels in four ranges are set. In other embodiments, the singing performance levels can also be divided into more detailed ranges to perform more refined ratio adjustments. The proportion adjustment amount of each audio component volume under each volume ratio adjustment strategy in Table 2 above is only an illustrative example and cannot be considered as a limitation on the scope of use of this application.

[0124] In some embodiments, if the singing performance level is set to 5 levels, namely: excellent, good, average, slightly poor and very poor, on this basis, if the first singing performance level is set to average, the second singing performance level is set to slightly poor, and the third singing performance level is set to very poor. Correspondingly, if the target singing performance level is excellent or good, the default volume ratio is maintained unchanged according to the volume ratio adjustment strategy of "maintaining the default volume ratio unchanged", for example, maintaining the singing recording volume: accompaniment volume: original singing volume = 0.5:0.5:0.

[0125] If the target singing performance level is average, according to the volume ratio adjustment strategy of "on the basis of the default volume ratio, maintain the proportion of the original singing volume unchanged, and increase the proportion of the accompaniment volume and reduce the proportion of the singing recording volume", the default volume ratio (singing recording volume: accompaniment volume: original singing volume = 0.5:0.5:0) is adjusted to singing recording volume: accompaniment volume: original singing volume = 0.4:0.6:0.

[0126] If the target singing performance level is slightly poor, according to the volume ratio adjustment strategy of "on the basis of the default volume ratio, increase the proportion of the original singing volume, and reduce the proportion of the accompaniment volume and the proportion of the singing recording volume", the default volume ratio (singing recording volume: accompaniment volume: original singing volume = 0.5:0.5:0) is adjusted to singing recording volume: accompaniment volume: original singing volume = 0.45:0.4:0.15.

[0127] If the target singing performance level is extremely poor, according to the volume ratio adjustment strategy of "on the basis of the default volume ratio, maintain the proportion of singing recording volume unchanged, reduce the proportion of accompaniment volume and increase the proportion of original singing volume", the default volume ratio (singing recording volume: accompaniment volume: original singing volume = 0.5:0.5:0) is adjusted to singing recording volume: accompaniment volume: original singing volume = 0.5:0.2:0.3.

[0128] Through the above-mentioned volume ratio adjustment strategy, the performance deficiency of the user (target object) when singing the target song segment can be masked by using the accompaniment and the original audio, thereby ensuring the auditory effect of the subsequent playback of the synthesized audio. Moreover, it is convenient for the user to adjust his or her own status in time during the masking gaps of the accompaniment and the original audio, which is conducive to improving the singing performance of subsequent song fragments.

[0129] In some embodiments, before step 330, the method further includes: obtaining adjustment indication information from the client; if the adjustment indication information indicates that automatic adjustment of the volume ratio is allowed, adjusting the default volume ratio according to the target volume ratio adjustment strategy to obtain a target volume ratio suitable for the target object to sing the target song segment; if the adjustment indication information indicates that automatic adjustment of the volume ratio is not allowed, there is no need to adjust the default volume ratio.

[0130] The client interface may provide an option for setting whether to allow automatic adjustment of the volume ratio. If the user sets the option to allow automatic adjustment of the volume ratio, adjustment indication information indicating that automatic adjustment of the volume ratio is allowed is generated and stored; if the user does not set the option to allow automatic adjustment of the volume ratio, adjustment indication information indicating that automatic adjustment of the volume ratio is not allowed is saved.

[0131] In some embodiments, the adjustment indication information may also indicate whether the proportion of each audio component volume in the three audio component volumes (singing recording volume, accompaniment volume, original singing volume) is allowed to be automatically adjusted. For ease of description, the audio component volume indicated by the adjustment indication information that is allowed to be automatically adjusted is referred to as the first audio component volume, and the audio component volume indicated that is not allowed to be automatically adjusted is referred to as the second audio component volume. In this case, in step 330, the proportion of the first audio component volume in the default volume ratio can be adjusted according to the target volume ratio adjustment strategy, without adjusting the proportion of the second audio component volume in the default volume ratio, and maintaining the proportion of the second audio component volume in the default volume ratio.

[0132] Figure 4 is a schematic diagram of an automatic adjustment setting interface of a proportion in a client according to an embodiment of the present application, such as Figure 4 As shown in the figure, the "Microphone Volume Automatic Adjustment" option is used to set whether the proportion of the singing recording volume is allowed to be automatically adjusted. If the "Microphone Volume Automatic Adjustment" option is set to a checked state, it means that the proportion of the singing recording volume is allowed to be automatically adjusted; the "Accompaniment Volume Automatic Adjustment" option is used to set whether the proportion of the accompaniment volume is allowed to be automatically adjusted, and the "Original Singer Volume Automatic Adjustment" option is used to set whether the proportion of the original singing volume is allowed to be automatically adjusted. If all three options are set to a checked state, it means that the proportion of the volume of each individual audio component is allowed to be automatically adjusted.

[0133] It is understandable that in Figure 4 In a corresponding embodiment, the sum of the proportions of the three volume components in the volume ratio may be preset to not exceed 1, so that the proportion of the volume of a single volume component may be allowed to be adjusted.

[0134] Step 250, in response to the target object singing and starting to sing the target song segment, adjusting the volume ratio to the target volume ratio.

[0135] Among them, adjusting the volume ratio to the target volume ratio refers to adjusting the volume ratio in the client where the target object is located to the target volume ratio. After adjusting to the target volume ratio, in the process of the target object singing the target song segment, according to the target volume ratio, the recorded target object singing the target song segment, the original singing audio segment corresponding to the target song segment, and the accompaniment audio segment corresponding to the target song segment are synthesized to obtain a synthesized audio segment. That is to say, in the synthesized audio segment: the volume proportion of the singing audio segment is equal to the proportion of the singing recording volume in the target volume ratio, the volume proportion of the original singing audio segment corresponding to the target song segment is equal to the proportion of the original singing volume in the target volume ratio, and the volume proportion of the accompaniment audio segment corresponding to the target song segment is equal to the proportion of the accompaniment volume in the target volume ratio. When the target object starts singing the target song segment, the middle volume ratio is adjusted to the target volume ratio. In this way, in the process of the target object singing the target song segment, the volume recorded for the target object singing the song segment is the product of the set total volume (for example, the set recording volume) and the proportion of the singing recording volume in the target volume ratio. Subsequently, the accompaniment volume determined by the proportion of the accompaniment volume in the target volume ratio is used to play the accompaniment audio segment corresponding to the target song segment, and the original singing volume determined by the proportion of the original singing volume in the target volume ratio is used to play the original singing audio segment corresponding to the target song segment.

[0136] In some embodiments, after step 250, the method further includes: performing audio synthesis on the original audio segment corresponding to the target song segment, the corresponding accompaniment audio segment, and the singing audio segment collected during the target object singing the target song segment according to the target volume ratio to obtain a synthesized audio segment corresponding to the target song segment; and playing the synthesized audio segment.

[0137] The volume of the original singing audio segment, the volume of the accompaniment audio segment, and the volume of the collected singing audio segment corresponding to the target song segment to be synthesized can be determined according to the recording volume on the client side where the target object is located (i.e., the microphone volume set on the client side where the target object is located) and the target volume ratio. Assume that the target volume ratio is: singing recording volume: original singing volume: accompaniment volume = k1: k2: k3; Assume that the set recording volume is V, in some embodiments, the set recording volume can be used to constrain the volume of the singing audio segment, that is, to ensure that the volume of the singing audio segment is V; the volume of the original singing audio segment is The volume of the accompaniment audio segment is In this way, after the three are synthesized, the volume of the audio segment sung in the synthesized audio segment is V; the volume of the original audio segment is The volume of the accompaniment audio segment is

[0138] In other embodiments, the set recording volume can be used to limit the volume of the synthesized audio segment, that is, to ensure that the volume of the synthesized audio segment is V. Correspondingly, it can be determined that: the volume of the singing audio segment is The volume of the original audio segment is The volume of the accompaniment audio segment is After the three are synthesized, the volume of the audio segment sung in the synthesized audio segment is The volume of the original audio segment is The volume of the accompaniment audio segment is

[0139] In the present application, the target singing performance level of the target object for the target song segment is predicted according to the pitch feature information of the target original singing audio segment through the target singing rating model applicable to the target object; and according to the target singing performance level of the target object for the target song segment, the target volume ratio applicable to the target object when singing the target song segment is determined; and when the target object starts singing the target song segment, the volume ratio is adjusted to the target volume ratio. Through the scheme of the present application, the target singing performance level when the target object sings the target song segment is predicted in advance, and the volume ratio is automatically adjusted to the target volume ratio that matches the target singing performance level, and the target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume, and the shortcomings of the target object singing are covered up by adjusting the ratio between the singing recording volume, the original singing volume and the accompaniment volume. Through the method of the present application, there is no need to repair the singing audio of the target object, and there will be no problem of unnatural singing effect caused by excessive or improper repair. Moreover, after adjusting the volume ratio, the singing playback effect of the subsequent singing audio can be guaranteed, and the auditory experience can be guaranteed.

[0140] In some embodiments, Figure 5 As shown, before step 230, the method further includes:

[0141] Step 510, obtaining sample singing audio of the target object singing multiple sample songs and sample original singing audio of each sample song.

[0142] The sample song refers to a song used to train the target singing rating model. The original singing audio of the sample song is called the sample original singing audio. The sample singing audio of the target object singing the sample song is obtained by recording the target object singing the sample song. The sample original singing audio of the sample song can be obtained from a song database.

[0143] Step 520, segment the sample singing audio of the sample song and the sample original singing audio of the sample song according to the multiple sample song segments of the sample song, and obtain multiple first audio segments in the sample singing audio and multiple second audio segments in the sample original singing audio.

[0144] The method of segmenting the sample song is similar to the method of segmenting the target song mentioned above, and will not be repeated here. The singing audio segment corresponding to the sample song segment in the sample singing audio is called the first audio segment, and the original singing audio segment corresponding to the sample song segment in the sample original singing audio is called the second audio segment. It can be understood that for the same sample song segment, the number of first audio segments and the number of second audio segments are equal to the number of sample song segments in the sample song. A sample song segment corresponds to a first audio segment and a second audio segment. For example, if a sample song segment is from the 0th to the 10th s in the sample song, then the first audio segment corresponding to the sample song segment refers to the segment from the 0th to the 10th s in the sample singing audio, and the second audio segment corresponding to the sample song segment refers to the segment from the 0th to the 10th s in the sample original singing audio.

[0145] Step 530: extract first pitch feature information of each first audio segment and second pitch feature information of each second audio segment.

[0146] In some embodiments, the first pitch feature information may be a pitch value sequence formed by the pitch values ​​of all audio frames in the first audio segment, and correspondingly, the second pitch feature information may be a pitch value sequence formed by the pitch values ​​of all audio frames in the second audio segment.

[0147] In other embodiments, the first pitch feature information may include the feature value of the first audio segment under one pitch feature or multiple pitch features. Similarly, the second pitch feature information includes the feature value of the second audio segment under one pitch feature or multiple pitch features, and the pitch feature may be the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the preceding silent pause. The pitch features involved in the first pitch feature information are the same as the pitch features involved in the second pitch feature information. The specific method of extracting the first pitch feature information from the first audio segment and extracting the second pitch feature information from the second audio segment is similar to the method of extracting the pitch feature information from the target original audio segment mentioned above, and will not be repeated here.

[0148] Step 540, determining the sample singing performance level of the target object for each sample song segment based on the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment corresponding to the same sample song segment.

[0149] The target object's sample singing performance level for a sample song segment is used to indicate the target object's singing effect when singing the sample song segment. For the first audio segment and the second audio segment corresponding to the same sample song segment, the higher the similarity between the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment, the closer the target object's singing for the sample song segment is to the original singer, indicating that the target object's singing for the sample song segment indicated by the sample singing performance level is better; conversely, the smaller the similarity between the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment, the more different the target object's singing for the sample song segment is from the original singer, indicating that the target object's singing for the sample song segment indicated by the sample singing performance level is worse.

[0150] In some embodiments, the first pitch feature information of each first audio segment can be vectorized to obtain the first pitch feature vector of each first audio segment, and the second pitch feature information of each second audio segment can be vectorized to obtain the second pitch feature vector of each second audio segment. The first pitch feature vector and the second pitch feature vector have the same dimension, and the same dimension represents the same pitch feature. Afterwards, the first pitch feature vector of the first audio segment and the second pitch feature vector of the second audio segment corresponding to the same sample song segment are calculated for similarity to obtain the pitch feature similarity of the two. Afterwards, the sample singing performance level of the target object for each sample song segment is determined based on the pitch feature similarity.

[0151] In some embodiments, if the first pitch feature information can be a pitch value sequence formed by the pitch values ​​of all audio frames in the first audio segment, and the second pitch feature information can be a pitch value sequence formed by the pitch values ​​of all audio frames in the second audio segment, the two audio frames aligned in the first audio segment and the second audio segment corresponding to the same sample song segment are regarded as an audio frame group, one audio frame in an audio frame group comes from the first audio segment, and the other audio frame comes from the second audio segment. The t-th audio frame of the first audio segment in the same sample song segment and the t-th audio frame in the second audio segment are two audio frames aligned, and t is a positive integer. On this basis, according to the first pitch feature information and the second pitch feature information, the pitch values ​​of the two aligned audio frames in each audio frame group can be subtracted, and the absolute value can be taken to obtain the pitch value difference corresponding to each audio frame group. Afterwards, the pitch value differences corresponding to all audio frame groups determined for the same sample song segment are averaged to obtain the average pitch value difference of the target object for each sample song segment. Subsequently, based on the correspondence between the pitch value difference and the singing performance level, the singing performance level corresponding to the calculated average pitch value difference is determined as the sample singing performance level of the target object for the corresponding sample song segment.

[0152] In other embodiments, the correspondence between the pitch value difference and the singing score can also be calculated. After calculating the pitch value difference corresponding to each audio frame group, the singing score of the target object for each sample song segment is determined. For example, if the pitch value difference does not exceed the first pitch value threshold, the corresponding singing score is plus N1 points (N1 is a positive integer). If the pitch value difference ∈ (first pitch value threshold, second pitch value threshold), the corresponding singing score is deducted N2 points (N2 is a positive integer), and the second pitch value threshold is greater than the first pitch value threshold, for example, the first pitch value threshold is 2, and the second pitch value threshold is 4; if the pitch value difference exceeds the second pitch value threshold, the corresponding singing score is deducted N3 points (N3 is a positive integer). N1, N2 and N3 can be set according to actual needs, for example, N2 is less than N1 and less than N3, for example, N1 is 3, N2 is 1, and N3 is 3. Afterwards, the singing scores corresponding to all the audio frame groups determined for the same sample song segment are accumulated and averaged to obtain the first average singing score of the target object for each sample song segment. Thereafter, according to the corresponding relationship between the singing score and the singing performance level, the singing performance level corresponding to the average singing score of the target object for each sample song segment is used as the sample singing performance level of the target object for the corresponding sample song segment.

[0153] In some embodiments, the first pitch feature information includes feature values ​​of the first audio segment under multiple pitch features; the second pitch feature information includes feature values ​​of the second audio segment under multiple pitch features; Figure 6 As shown, step 540 includes:

[0154] Step 610, determining the characteristic value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under each pitch feature based on the characteristic value of the first audio segment corresponding to the same sample song segment under multiple pitch features and the characteristic value of the second audio segment under multiple pitch features.

[0155] For the first audio segment and the second audio segment corresponding to the same sample song segment, the feature values ​​of the first audio segment and the second audio segment under the same pitch feature are subtracted, and the absolute value is taken to obtain the feature value deviation under the pitch feature.

[0156] Step 620, determining the sample singing performance level of the target object for each sample song segment based on the feature value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under multiple pitch features.

[0157] In some embodiments, the eigenvalue deviations of the first audio segment and the second audio segment corresponding to the same sample song segment under multiple pitch features can be weightedly calculated to obtain a reference feature eigenvalue deviation, and then based on the correspondence between the eigenvalue deviation and the singing performance level, the singing performance level corresponding to the reference feature eigenvalue deviation is used as the sample singing performance level of the target object for the corresponding sample song segment.

[0158] In some embodiments, a correspondence between the eigenvalue deviation and the singing score can be set, and then, based on the eigenvalue deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under each pitch feature, the singing score corresponding to the eigenvalue deviation under each pitch feature is determined; then, the singing scores corresponding to the eigenvalue deviations under multiple pitch features are accumulated and averaged to obtain a second average singing score; thereafter, based on the correspondence between the singing score and the singing performance level, the singing performance level corresponding to the second average singing score is used as the sample singing performance level of the target object for the corresponding sample song segment.

[0159] For example, if there are a total of 5 singing performance levels, namely excellent, good, average, slightly poor and extremely poor, the corresponding relationship between the singing score and the singing performance level can be: 1) If the singing score exceeds the first score threshold, the corresponding singing performance level is excellent; 2) If the singing score ∈ (second score threshold, first score threshold], the corresponding singing performance level is good; 3) If the singing score ∈ (third score threshold, second score threshold], the corresponding singing performance level is average; 4) If the singing score ∈ (fourth score threshold, third score threshold], the corresponding singing performance level is slightly poor; 5) If the singing score does not exceed the fourth score threshold, the corresponding singing performance level is extremely poor. The first score threshold>the second score threshold>the third score threshold>the fourth score threshold, and they are all non-negative numbers, which can be set according to actual needs. For example, the first score threshold is 2.5, the second score threshold is 2.0, the third score threshold is 1.2, and the fourth score threshold is 0.5.

[0160] Step 550, training a target singing rating model based on the second pitch feature information of the plurality of second audio segments and the sample singing performance level of the target object for each sample song segment.

[0161] In some embodiments, the second pitch feature information includes feature values ​​of the corresponding second audio segment under multiple pitch features; the target singing rating model may be a Bayesian classification model. Figure 7 As shown, step 550 includes the following steps 710 to 730:

[0162] Step 710, based on the target object's sample singing performance levels for multiple sample song segments, statistically calculate the target object's prior probability for each singing performance level.

[0163] When the target singing rating model is a Bayesian classification model, the process of training the target singing rating model is mainly to determine the prior probability of the target object for each singing performance level, and the conditional probability of each singing performance level of the target object under the characteristic value of each pitch feature. In this way, after training, the target singing rating model predicts the probability of each singing performance level corresponding to the song segment to be sung by the target object based on the prior probability determined in the training process and the determined conditional probability, combined with the pitch feature information of the original audio segment corresponding to the song segment to be sung by the target object, and then determines the singing performance level that finally corresponds to the target object singing the song segment.

[0164] In some embodiments, among the sample singing performance levels of the target object for multiple sample song segments, the number of sample song segments corresponding to each singing performance level can be counted, and then the probability of occurrence of each singing performance level can be counted.

[0165] For example, if a total of 5 singing performance levels are set, namely singing performance level I, singing performance level II, singing performance level III, singing performance level IV and singing performance level V, among the sample singing performance levels of the target object for 1000 sample song segments, the statistical numbers of sample song segments with singing performance levels of singing performance level I, singing performance level II, singing performance level III, singing performance level IV and singing performance level V are 50, 150, 200, 350 and 250 respectively. Then, the prior probabilities of the target object for singing performance level I, singing performance level II, singing performance level III, singing performance level IV and singing performance level V can be statistically determined to be 0.05, 0.15, 0.2, 0.35 and 0.25 respectively.

[0166] Step 720, based on the feature values ​​of the multiple second audio segments under the multiple pitch features and the sample singing performance levels of the target object for the multiple sample song segments, the conditional probability of each singing performance level under the feature value of each pitch feature is counted for the target object.

[0167] For example, if there are three eigenvalues ​​of pitch feature 1, namely eigenvalue A1, eigenvalue A2 and eigenvalue A3, assuming that when the eigenvalue of pitch feature 1 is eigenvalue A1, it is represented as pitch feature 1_A1. Among the 50 sample song segments with singing performance level I as above, there are three sample song segments with the eigenvalue of pitch feature 1 being eigenvalue A1. Then, the conditional probability that the target object has pitch feature 1 being eigenvalue A1 in singing performance level I is P(pitch feature 1_A1|singing performance level I) = 3 / 50 = 0.06.

[0168] Step 730, determining the target singing rating model according to the prior probability of the target object for each singing performance level, and the conditional probability of each singing performance level under the characteristic value of each pitch feature.

[0169] The prior probability of the target object for each singing performance level, as well as the conditional probability of each singing performance level under the characteristic value of each pitch feature, can be used as parameters of the target singing rating model and as the basis for subsequent singing performance level prediction.

[0170] Through the above training process, the target singing rating model suitable for the target object can be determined in a targeted manner to ensure the accuracy of the singing performance prediction made by the subsequent target singing rating model.

[0171] In some embodiments, if the target singing rating model is a Bayesian classification model, the pitch feature information of the target original singing audio segment includes feature values ​​of the target original singing audio segment under multiple pitch features; Figure 8 As shown, in step 230, the target singing rating model is processed according to the following steps 810-820 to predict the target singing performance level of the target object for each target song segment:

[0172] Step 810, performs a posterior probability calculation based on the prior probability of the target object for each singing performance level and the conditional probability of the characteristic values ​​of each singing performance level under multiple pitch features corresponding to the target original audio segment, to determine that the target singing performance level of the target object for the target song segment is the posterior probability of each singing performance level.

[0173] Step 820, taking the singing performance level with the largest posterior probability as the target singing performance level of the target object for the target song segment.

[0174] For example, if there are 5 pitch features involved in total, namely pitch feature 1, pitch feature 2, pitch feature 3, pitch feature 4 and pitch feature 5, the 5 pitch features can be the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the preceding silent pause duration mentioned above. If in the pitch feature information of the target original audio segment: the feature value of pitch feature 1 is feature value A3 (expressed as pitch feature 1_A3), the feature value of pitch feature 2 is feature value B1 (expressed as pitch feature 2_B1), the feature value of pitch feature 3 is feature value C2 (expressed as pitch feature 3_C2), the feature value of pitch feature 4 is feature value D3 (expressed as pitch feature 4_D3), and the feature value of pitch feature 5 is feature value E4 (expressed as pitch feature 5_E4). If there are a total of 5 singing performance levels, namely singing performance level I, singing performance level II, singing performance level III, singing performance level IV and singing performance level V, the target singing performance level of the target object for the target song segment can be determined according to the following formula to determine the posterior probability of each singing performance level:

[0175] P(singing performance level I|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) = P(singing performance level I)*P(pitch feature 1_A3|singing performance level I)*P(pitch feature 2_B1|singing performance level I)*P(pitch feature 3_C2|singing performance level I)*P(pitch feature 4_D3|singing performance level I)*P(pitch feature 5_E4|singing performance level I); (Formula 4)

[0176] P(singing performance level II|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) = P(singing performance level II)*P(pitch feature 1_A3|singing performance level II)*P(pitch feature 2_B1|singing performance level II)*P(pitch feature 3_C2|singing performance level II)*P(pitch feature 4_D3|singing performance level II)*P(pitch feature 5_E4|singing performance level II); (Formula 5)

[0177] P(singing performance level III|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) = P(singing performance level III)*P(pitch feature 1_A3|singing performance level III)*P(pitch feature 2_B1|singing performance level III)*P(pitch feature 3_C2|singing performance level III)*P(pitch feature 4_D3|singing performance level III)*P(pitch feature 5_E4|singing performance level III); (Formula 6)

[0178] P(singing performance level IV|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) = P(singing performance level IV)*P(pitch feature 1_A3|singing performance level IV)*P(pitch feature 2_B1|singing performance level IV)*P(pitch feature 3_C2|singing performance level IV)*P(pitch feature 4_D3|singing performance level IV)*P(pitch feature 5_E4|singing performance level IV); (Formula 7)

[0179] P(singing performance level V|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) = P(singing performance level V)*P(pitch feature 1_A3|singing performance level V)*P(pitch feature 2_B1|singing performance level V)*P(pitch feature 3_C2|singing performance level V)*P(pitch feature 4_D3|singing performance level V)*P(pitch feature 5_E4|singing performance level V); (Formula 8)

[0180] In the above equation, P(singing performance level I|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) represents the posterior probability that the target singing performance level of the target object for the target song segment is singing performance level I; P(singing performance level I) represents the prior probability that the target object has singing performance level I.

[0181] P(Pitch Feature 1_A3|Singing Performance Level I) represents the conditional probability of singing performance level I for the target object when the characteristic value of pitch feature 1 is characteristic value A3. Similarly, P(Pitch Feature 2_B1|Singing Performance Level I) represents the conditional probability of singing performance level I for the target object when the characteristic value of pitch feature 2 is characteristic value B1. The physical meanings of other characters are similar and will not be explained here one by one.

[0182] After determining the posterior probabilities that the target singing performance level of the target object for the target song segment is singing performance level I to V respectively according to the above formulas 4 to 8, the singing performance level corresponding to the maximum posterior probability is used as the target singing performance level of the target object for the target song segment. For example, if P (singing performance level V|pitch feature 1_A3, pitch feature 2_B1, pitch feature 3_C2, pitch feature 4_D3, pitch feature 5_E4) calculated according to formula 8 is the largest, it can be determined that the target singing performance level of the target object for the target song segment is singing performance level V.

[0183] In other embodiments, if the target singing rating model is a support vector machine model or a deep learning network model, the following can be used: Fig. 9 The training process is shown in Fig. 9 As shown, step 550 includes:

[0184] Step 910: The target singing rating model predicts the singing performance rating based on the second pitch feature information of each second audio segment, and outputs the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment.

[0185] After the second pitch feature information of the second audio segment is input into the target singing rating model, the target singing rating model can extract features of the second pitch feature information, and then classify it according to the extracted features. The predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment is the probability of each singing performance level, and the singing performance level with the largest probability is used as the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment.

[0186] Step 920, calculate the predicted loss based on the predicted singing performance level of the target object for the sample song segments corresponding to each second audio segment, and the sample singing performance level of the target object for each sample song segment.

[0187] The prediction loss can be calculated by combining the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment and the sample singing performance level of the target object for each sample song segment through a loss function. The loss function can be an absolute value loss function, a mean square error loss function, a root mean square error loss function, a cross entropy loss function, etc., which are not specifically limited here.

[0188] Step 930, adjust the parameters of the target singing rating model according to the predicted loss until the training end condition is reached.

[0189] The training end condition may be that the number of iterations of the target singing rating model reaches a threshold, or the loss function converges, which is not specifically limited here. In the above training process, the sample singing performance level of the target object for each sample song segment is used as supervision information, and the target singing rating model is supervised for training, so that the target singing rating model learns the pitch features of the original singing audio segments (second audio segments) of the target object with different singing performance levels, so that after training, the target singing rating model can accurately predict the singing performance level of the target object for the input audio segment based on the pitch feature information of the input audio segment.

[0190] Fig.10 is a flowchart of an audio processing method according to an embodiment of the present application, such as Fig.10 As shown, including:

[0191] Step 1010, recording audio while the target object is singing the sample song, to obtain sample singing audio of the sample song.

[0192] Step 1020, obtaining the sample original audio of the sample song.

[0193] Step 1030, calculate the pitch feature deviation of the sample singing audio relative to the sample original singing audio.

[0194] Among them, the sample singing audio and the sample original singing audio can be segmented according to the multiple sample song segments of the sample song, the sample singing audio is divided into multiple first audio segments, and the sample original singing audio is divided into multiple second audio segments; then, the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment are extracted respectively; and for the first audio segment and the second audio segment that are aligned (that is, the first audio segment and the second audio segment corresponding to the same sample song segment), the pitch feature deviation is calculated according to the corresponding first pitch feature information and the second pitch feature information. Please refer to the above description for the specific calculation method.

[0195] Afterwards, the sample singing performance level of the target object for each sample song segment in the sample song is determined based on the pitch feature deviation between the aligned first audio segment and the second audio segment.

[0196] Step 1040, online training of the target singing rating model.

[0197] That is, the target singing rating model is trained by the target object's sample singing performance level for each sample song segment in the sample song and the second pitch feature information of multiple second audio segments. Please refer to the above description for the specific process of training.

[0198] Step 1050, predicting the target singing performance level of the target subject for the target song segment through the target singing rating model.

[0199] The target original singer audio segment corresponding to the target song clip can be obtained, and the pitch feature information of the target original singer audio segment can be extracted, and the pitch feature information of the target original singer audio segment can be input into the target singing rating model. The target singing rating model predicts the singing performance level according to the pitch feature information of the target original singer audio segment, and outputs the target singing performance level of the target object for the target song clip.

[0200] Step 1060, dynamically adjust the volume ratio according to the target singing performance level.

[0201] Specifically, the target volume ratio applicable to the target object when singing the target song segment can be determined according to the target singing performance level, and in response to the target object starting to sing the target song segment, the volume ratio in the client is adjusted to the target volume ratio.

[0202] In the scheme of the present application, a target singing rating model suitable for a target object is trained in a targeted manner, and the target singing rating model predicts the target singing performance level of the target object for the target song segment based on the pitch feature information of the target original singing audio segment, and adjusts the volume ratio during the target object singing the target song segment in a targeted manner. The target singing performance level can reflect whether the singer (target object) is unable to control the target song segment. When it is determined that the singing effect indicated by the target singing performance level is not good, the volume ratio is adjusted, that is, the ratio between the singing recording volume, the original singing volume and the accompaniment volume is adjusted to mask the shortcomings of the target object in singing the target song segment. Moreover, it is also convenient for the singer to adjust his or her voice state in a timely manner to improve the singing performance and ensure the final auditory effect.

[0203] In the related art, in order to improve the auditory effect of singing, the singing assistance application uses automatic tuning technology to tune the singer's singing recording audio, but this method has two problems: 1. The tuning effect is strongly related to the tuning algorithm, especially in online real-time scene applications. Automatic tuning technology generally does not reach the level of perfectly matching the generated audio features with the singer's vocal characteristics. Therefore, listeners can easily notice that the audio has been tuned and is significantly different from the singer's real voice; 2. After tuning, the sound is likely to become unnatural or abrupt, and the effect after tuning does not meet expectations.

[0204] The solution of this application does not modify the audio recording of the singer's singing, does not change the singer's original voice, avoids the negative impact of unnatural singing caused by excessive or improper tuning, and uses an adaptive volume adjustment method to participate in the whole process of the singer's singing, so that the singer can fully experience the fun of singing, rather than just understanding the objective scoring results of the singing. Allowing singers to sing more freely and happily in this way can enhance the singer's singing experience.

[0205] The following describes an apparatus embodiment of the present application, which can be used to execute the method in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned method embodiment of the present application.

[0206] Fig.11 is a block diagram of an audio processing device according to an embodiment of the present application, such as Fig.11 As shown, the audio processing device includes: an acquisition module 1110, used to acquire a target original singing audio segment corresponding to a target song segment in a target song; an extraction module 1120, used to extract pitch feature information of the target original singing audio segment; a prediction module 1130, used to predict a target singing performance level of a target object for a target song segment through a target singing rating model applicable to the target object based on the pitch feature information of the target original singing audio segment; a target volume ratio determination module 1140, used to determine a target volume ratio applicable to a target object when singing the target song segment based on the target singing performance level of the target object for the target song segment; the target volume ratio refers to the ratio among the singing recording volume, the original singing volume and the accompaniment volume; a ratio adjustment module 1150, used to adjust the volume ratio to the target volume ratio in response to the target object starting to sing the target song segment.

[0207] In some embodiments, the audio processing device also includes: a first acquisition module, used to acquire sample singing audio of multiple sample songs sung by the target object and sample original singing audio of each sample song; a segmentation module, used to segment the sample singing audio of the sample songs and the sample original singing audio of the sample songs according to the multiple sample song segments of the sample songs, and obtain multiple first audio segments in the sample singing audio and multiple second audio segments in the sample original singing audio; a pitch feature extraction module, used to extract the first pitch feature information of each first audio segment and the second pitch feature information of each second audio segment; a singing performance level determination module, used to determine the sample singing performance level of the target object for each sample song segment based on the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment corresponding to the same sample song segment; a training module, used to train the target singing rating model based on the second pitch feature information of the multiple second audio segments and the sample singing performance level of the target object for each sample song segment.

[0208] In some embodiments, the second pitch feature information includes feature values ​​of the corresponding second audio segment under multiple pitch features; the training module includes:

[0209] A first statistical unit is used to count the prior probability of the target object for each singing performance level according to the sample singing performance levels of the target object for multiple sample song segments;

[0210] A second statistical unit is used to count the conditional probabilities of each singing performance level under the characteristic value of each pitch feature for the target object according to the characteristic values ​​of the plurality of second audio segments under the plurality of pitch features and the sample singing performance levels of the target object for the plurality of sample song segments;

[0211] The model determination unit is used to determine the target singing rating model according to the prior probability of the target object for each singing performance level and the conditional probability of each singing performance level under the characteristic value of each pitch feature.

[0212] In some embodiments, the pitch feature information of the target original singer audio segment includes the feature values ​​of the target original singer audio segment under multiple pitch features; the prediction module 1130 is used to: the target singing rating model is processed according to the following process to predict the target singing performance level of the target object for the target song segment: according to the prior probability of the target object for each singing performance level and the conditional probability of the feature values ​​of each singing performance level under multiple pitch features corresponding to the target original singer audio segment, a posterior probability calculation is performed to determine that the target singing performance level of the target object for the target song segment is the posterior probability of each singing performance level; the singing performance level with the largest posterior probability is used as the target singing performance level of the target object for the target song segment.

[0213] In some embodiments, the training module includes:

[0214] A prediction unit, configured to predict the singing performance rating of the target object according to the second pitch feature information of each second audio segment by using the target singing rating model, and output the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment;

[0215] A prediction loss calculation unit, used to calculate the prediction loss according to the predicted singing performance level of the target object for the sample song segments corresponding to each second audio segment and the sample singing performance level of the target object for each sample song segment;

[0216] The parameter adjustment unit is used to adjust the parameters of the target singing rating model according to the prediction loss until the training end condition is reached.

[0217] In some embodiments, the first pitch feature information includes feature values ​​of the first audio segment under multiple pitch features; the second pitch feature information includes feature values ​​of the second audio segment under multiple pitch features;

[0218] Singing performance level determination module, including:

[0219] A feature value deviation calculation unit, for determining feature value deviations of the first audio segment and the second audio segment corresponding to the same sample song segment under each pitch feature according to feature values ​​of the first audio segment under multiple pitch features and feature values ​​of the second audio segment under multiple pitch features;

[0220] The sample singing performance level determination unit is used to determine the sample singing performance level of the target object for each sample song segment based on the feature value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under multiple pitch features.

[0221] In some embodiments, the target volume ratio determination module 1140 includes:

[0222] A default volume ratio obtaining unit, used to obtain a default volume ratio;

[0223] An adjustment strategy determination unit, for determining a target volume ratio adjustment strategy adapted to the target singing performance level according to the target singing performance level of the target object for the target song segment;

[0224] The volume ratio adjustment unit is used to adjust the default volume ratio according to the target volume ratio adjustment strategy to obtain a target volume ratio suitable for the target object to sing the target song segment.

[0225] In some embodiments, the target volume ratio adjustment strategy is used to adjust the proportion of one of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio, or to adjust the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

[0226] In some embodiments, the adjustment strategy determination unit is configured to:

[0227] If the target singing performance level of the target object for the target song segment is higher than the first singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: maintaining the default volume ratio unchanged;

[0228] If the target singing performance level of the target object for the target song segment is not higher than the first singing performance level, and higher than the second singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, the proportion of the original singing volume is maintained unchanged, and the proportion of the accompaniment volume is increased and the proportion of the singing recording volume is reduced;

[0229] If the target singing performance level of the target object for the target song segment is not higher than the second singing performance level and higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, increase the proportion of the original singing volume, and reduce the proportion of the accompaniment volume and the proportion of the singing recording volume;

[0230] If the target singing performance level of the target object for the target song segment is not higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, the proportion of the singing recording volume is maintained unchanged, the proportion of the accompaniment volume is reduced, and the proportion of the original singing volume is increased;

[0231] Among them, the first singing performance level > the second singing performance level > the third singing performance level, and the singing performance level is positively correlated with the singing effect.

[0232] In some other embodiments, the target volume ratio determination module 1140 is used to:

[0233] According to the correspondence between the singing performance level and the volume ratio, the volume ratio corresponding to the target singing performance level of the target object for the target song segment is used as the target volume ratio applicable to the target object singing the target song segment; among them, in the correspondence between the singing performance level and the volume ratio, the higher the singing performance level, the lower the proportion of the original singer's volume in the corresponding volume ratio; the singing performance level is positively correlated with the singing effect.

[0234] In some embodiments, the pitch feature information includes feature values ​​of a target original singing audio segment under multiple pitch features; the target original singing audio segment includes multiple target audio frames; the extraction module 1120 includes:

[0235] A fundamental frequency detection unit, used to detect the fundamental frequency of each target audio frame in the target original audio segment, and determine the fundamental frequency of each target audio frame;

[0236] A pitch conversion unit, used to perform pitch conversion on the fundamental frequency of each target audio frame to obtain a pitch value of each target audio frame;

[0237] The feature value determination unit is used to determine the feature value of the target original singing audio segment under multiple pitch features according to the pitch values ​​of multiple target audio frames in the target original singing audio segment.

[0238] In some embodiments, the plurality of pitch features include at least two of the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the preceding silent pause.

[0239] In some embodiments, the audio processing device also includes: an audio synthesis module, which is used to perform audio synthesis on the original audio segment corresponding to the target song segment, the corresponding accompaniment audio segment, and the singing audio segment collected during the target object singing the target song segment according to the target volume ratio to obtain a synthesized audio segment corresponding to the target song segment; and a playback module, which is used to play the synthesized audio segment.

[0240] Fig.12 The structure diagram of the computer system suitable for implementing the electronic device of the embodiment of the present application is shown. It should be noted that: Fig.12 The computer system 1200 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application. The electronic device can be used to execute the audio processing method provided by the application, and the electronic device can be a terminal such as a smart phone, a tablet computer, a smart TV, etc.

[0241] like Fig.12As shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM) 1203, such as executing the method in the above embodiment. In RAM 1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0242] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 1208 as needed.

[0243] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1209, and / or installed from a removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.

[0244] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0245] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0246] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0247] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.

[0248] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit of the module or unit function.

[0249] According to one aspect of the embodiments of the present application, a computer program product is provided, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in any of the above embodiments.

[0250] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0251] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0252] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0253] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. An audio processing method, characterized in that: include: Obtain the target original audio segment corresponding to the target song segment in the target song; Extracting pitch feature information of the target original singer audio segment; According to the pitch feature information of the target original singer audio segment, the target singing performance level of the target object for the target song segment is predicted by a target singing rating model applicable to the target object; Determining a target volume ratio applicable to when the target subject sings the target song segment according to the target singing performance level of the target subject for the target song segment; The target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume; In response to the target object singing and starting to sing the target song segment, the volume ratio is adjusted to the target volume ratio.

2. The method according to claim 1, characterized in that Before predicting the target singing performance level of the target object for the target song segment based on the pitch feature information of the target original singing audio segment by using a target singing rating model applicable to the target object, the method further includes: Obtaining sample singing audio of the target object singing a plurality of sample songs and sample original singing audio of each of the sample songs; According to the multiple sample song segments of the sample song, the sample singing audio of the sample song and the sample original singing audio of the sample song are segmented to obtain multiple first audio segments in the sample singing audio and multiple second audio segments in the sample original singing audio; Extracting first pitch feature information of each of the first audio segments and second pitch feature information of each of the second audio segments; Determining the sample singing performance level of the target object for each of the sample song segments according to the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment corresponding to the same sample song segment; The target singing rating model is trained based on the second pitch feature information of the plurality of the second audio segments and the sample singing performance level of the target object for each of the sample song segments.

3. The method according to claim 2, characterized in that The second pitch feature information includes a feature value of the corresponding second audio segment under multiple pitch features; The method of training the target singing rating model according to the second pitch feature information of the plurality of the second audio segments and the sample singing performance levels of the target object for the plurality of the sample song segments comprises: According to the sample singing performance levels of the target object for the plurality of the sample song segments, counting the prior probability of the target object for each singing performance level; According to the characteristic values ​​of the plurality of second audio segments under the plurality of pitch features and the sample singing performance levels of the target object for the plurality of sample song segments, the conditional probability of each singing performance level under the characteristic value of each pitch feature is counted for the target object; The target singing rating model is determined according to the prior probability of the target object for each singing performance level and the conditional probability of each singing performance level under the characteristic value of each pitch feature.

4. The method according to claim 3, characterized in that The pitch feature information of the target original singer audio segment includes feature values ​​of the target original singer audio segment under multiple pitch features; The method of predicting the target singing performance level of the target object for the target song segment based on the pitch feature information of the target original singing audio segment by using a target singing rating model applicable to the target object comprises: The target singing rating model is processed according to the following process to predict the target singing performance level of the target object for the target song segment: According to the prior probability of the target object for each singing performance level and the conditional probability of the characteristic values ​​of each singing performance level under the multiple pitch features corresponding to the target original audio segment, a posterior probability calculation is performed to determine that the target singing performance level of the target object for the target song segment is the posterior probability of each singing performance level; The singing performance level with the largest posterior probability is used as the target singing performance level of the target object for the target song segment.

5. The method according to claim 2, characterized in that: The method of training the target singing rating model according to the second pitch feature information of the plurality of the second audio segments and the sample singing performance levels of the target object for the plurality of the sample song segments comprises: The target singing rating model predicts the singing performance rating according to the second pitch feature information of each second audio segment, and outputs the predicted singing performance level of the target object for the sample song segment corresponding to each second audio segment; Calculate the prediction loss based on the predicted singing performance level of the target object for the sample song segments corresponding to each of the second audio segments, and the sample singing performance level of the target object for each of the sample song segments; According to the predicted loss, the parameters of the target singing rating model are adjusted until the training end condition is reached.

6. The method according to claim 2, characterized in that The first pitch feature information includes feature values ​​of the first audio segment under multiple pitch features; the second pitch feature information includes feature values ​​of the second audio segment under multiple pitch features; The step of determining the sample singing performance level of the target object for each of the sample song segments according to the first pitch feature information of the first audio segment and the second pitch feature information of the second audio segment corresponding to the same sample song segment comprises: Determine the characteristic value deviation of the first audio segment and the second audio segment corresponding to the same sample song segment under each of the pitch features according to the characteristic values ​​of the first audio segment corresponding to the same sample song segment under multiple pitch features and the characteristic values ​​of the second audio segment under multiple pitch features; The sample singing performance level of the target object for each of the sample song segments is determined based on the feature value deviations of the first audio segment and the second audio segment corresponding to the same sample song segment under the multiple pitch features.

7. The method according to claim 1, characterized in that The step of determining a target volume ratio applicable to when the target object sings the target song segment according to the target singing performance level of the target object for the target song segment includes: Get the default volume ratio; Determining a target volume ratio adjustment strategy that matches the target singing performance level according to the target object for the target song segment; According to the target volume ratio adjustment strategy, the default volume ratio is adjusted to obtain a target volume ratio suitable for the target object when singing the target song segment.

8. The method according to claim 7, characterized in that The target volume ratio adjustment strategy is used to adjust the proportion of one of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio, or to adjust the proportion of at least two of the singing recording volume, the original singing volume and the accompaniment volume in the volume ratio.

9. The method according to claim 7, characterized in that: The step of determining a target volume ratio adjustment strategy that matches the target singing performance level according to the target object for the target song segment includes: If the target singing performance level of the target object for the target song segment is higher than the first singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: maintaining the default volume ratio unchanged; If the target singing performance level of the target object for the target song segment is not higher than the first singing performance level, and is higher than the second singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, the proportion of the original singing volume is maintained unchanged, and the proportion of the accompaniment volume is increased and the proportion of the singing recording volume is reduced; If the target singing performance level of the target object for the target song segment is not higher than the second singing performance level and higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, the proportion of the original singing volume is increased, and the proportion of the accompaniment volume and the proportion of the singing recording volume are reduced; If the target singing performance level of the target object for the target song segment is not higher than the third singing performance level, the target volume ratio adjustment strategy adapted to the target singing performance level is: on the basis of the default volume ratio, the proportion of the singing recording volume is maintained unchanged, the proportion of the accompaniment volume is reduced and the proportion of the original singing volume is increased; Among them, the first singing performance level > the second singing performance level > the third singing performance level, and the singing performance level is positively correlated with the singing effect.

10. The method according to claim 1, characterized in that The step of determining a target volume ratio applicable to when the target object sings the target song segment according to the target singing performance level of the target object for the target song segment includes: According to the correspondence between the singing performance level and the volume ratio, the volume ratio corresponding to the target singing performance level of the target object for the target song segment is used as the target volume ratio applicable to when the target object sings the target song segment; Among them, in the corresponding relationship between the singing performance level and the volume ratio, the higher the singing performance level, the lower the proportion of the original singer's volume in the corresponding volume ratio; the singing performance level is positively correlated with the singing effect.

11. The method according to any one of claims 1 to 10, characterized in that The pitch feature information includes feature values ​​of the target original singing audio segment under multiple pitch features; the target original singing audio segment includes multiple target audio frames; The step of extracting pitch feature information of the target original singer audio segment comprises: Performing fundamental frequency detection on each target audio frame in the target original audio segment to determine the fundamental frequency of each target audio frame; Performing pitch conversion on the fundamental frequency of each target audio frame to obtain a pitch value of each target audio frame; According to the pitch values ​​of multiple target audio frames in the target original singing audio segment, the feature value of the target original singing audio segment under multiple pitch features is determined.

12. The method according to claim 11, characterized in that The multiple pitch features include at least two of the highest pitch value, the lowest pitch value, the maximum stable duration of the pitch, the pitch variance value, and the duration of the preceding silent pause.

13. The method according to any one of claims 1 to 10, characterized in that After the target object starts singing the target song segment in response to the target object singing, and the volume ratio is adjusted to the target volume ratio, the method further includes: According to the target volume ratio, the original audio segment corresponding to the target song segment, the corresponding accompaniment audio segment, and the singing audio segment collected during the target subject singing the target song segment are synthesized to obtain a synthesized audio segment corresponding to the target song segment; The synthesized audio segment is played.

14. An audio processing device, characterized in that: include: An acquisition module, used to acquire a target original audio segment corresponding to a target song segment in a target song; An extraction module, used to extract pitch feature information of the target original singer audio segment; A prediction module, configured to predict a target singing performance grade of a target object for the target song segment based on the pitch feature information of the target original singing audio segment and through a target singing rating model applicable to the target object; A target volume ratio determination module, for determining a target volume ratio applicable to when the target object sings the target song segment according to the target singing performance level of the target object for the target song segment; The target volume ratio refers to the ratio between the singing recording volume, the original singing volume and the accompaniment volume; The ratio adjustment module is used to adjust the volume ratio to the target volume ratio in response to the target object starting to sing the target song segment.

15. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 13 is implemented.

16. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.

17. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Singing grading method and device and terminal

    CN107507628A

  • Audio processing method and device, electronic equipment and storage medium

    CN112216294A

  • Singing scoring method, computer equipment and storage medium

    CN115658959A

  • Singing intonation scoring method and device, equipment, medium and product

    CN116110431A

  • Audio adjustment method and device, electronic equipment and storage medium

    CN119049433A

Cited By

  • Singing accompaniment data management method and system based on adaptive adjustment

    CN120823821A

  • Adaptive adjustment-based singing accompaniment data management method and system

    CN120823821B