Audio processing method and device, storage medium and electronic equipment
By comparing the matching parameters of the singing audio with the standard audio during the chorus process and assigning weight coefficients for mixing, the problem of poor mixing accuracy caused by the uneven skill levels of multiple singers was solved, resulting in a more harmonious and unified chorus effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-31
AI Technical Summary
The varying vocal abilities of the singers resulted in poor mixing quality for the chorus, and related technologies failed to effectively address the issue of poor mixing accuracy caused by discrepancies in the vocal audio.
By acquiring multiple singing audios and comparing them with pre-set standard singing audios, and assigning weight coefficients to each singing audio based on the matching parameters, a mixing operation is performed to determine the target singing audio, thus achieving automatic mixing for multi-person chorus.
It improves the coordination and consistency between multiple vocal audio recordings, enhances the quality and listening experience of choral singing, and solves the problem of poor mixing accuracy.
Smart Images

Figure CN121768334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to an audio processing method and apparatus, a storage medium, and an electronic device. Background Technology
[0002] Currently, when mixing audio recordings from multiple singers, the varying skill levels of the singers, such as differences in their audio recordings, often result in poor mixing quality for choral performances. Furthermore, current technologies do not consider these differences in skill levels, making it difficult to guarantee a satisfactory mixing result. In summary, current technologies suffer from the technical problem of poor mixing accuracy due to incoordination between the audio recordings of multiple singers.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides an audio processing method and apparatus, storage medium and electronic device to at least solve the technical problem of poor mixing accuracy caused by the incoordination between multiple singing audios.
[0005] According to one aspect of the embodiments of this application, an audio processing method is provided, comprising: acquiring a first set of singing audio, wherein the first set of singing audio includes multiple audio inputs for a target music; determining a set of matching parameters based on the first set of singing audio and a corresponding standard singing audio, wherein the standard singing audio is a singing audio pre-set for the target music, and a first matching parameter in the set of matching parameters is used to indicate the degree of matching between a first singing audio in the first set of singing audio and the standard singing audio; assigning a first set of weighting coefficients to the first set of singing audio based on the set of matching parameters; performing a mixing operation on the first set of singing audio using the first set of weighting coefficients; and determining a target singing audio.
[0006] According to another aspect of the embodiments of this application, an audio processing apparatus is also provided, comprising: an acquisition module, configured to acquire a first set of singing audio, wherein the first set of singing audio includes multiple audio inputs for a target music; a determination module, configured to assign a first set of weight coefficients to the first set of singing audio based on the set of matching parameters, and perform a mixing operation on the first set of singing audio using the first set of weight coefficients to determine a target singing audio; and an execution module, configured to assign a first set of weight coefficients to the first set of singing audio based on the set of matching parameters, and perform a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio.
[0007] Optionally, the device is configured to determine a set of matching parameters based on the first set of singing videos and the corresponding standard singing videos in the following manner: determining a set of pitches corresponding to the first set of singing audios and a standard pitch corresponding to the standard singing audios, wherein the set of pitches and the standard pitches correspond to the same moment of the target music, and the first matching parameters are determined by the set of pitches, the first pitch corresponding to the first singing audios, and the standard pitches; and comparing each pitch in the set of pitches with the standard pitches to determine the set of matching parameters.
[0008] Optionally, the device is used to determine the set of matching parameters by comparing each pitch in the set of pitches with the standard pitch in the following manner: performing pitch conversion operations on the set of pitches and the standard pitch respectively to determine a set of note difference parameters and a set of interval distance parameters, wherein the note difference parameters are used to indicate the pitch difference between a pitch in the set of pitches and the standard pitch, and the interval distance parameters are used to indicate the interval relationship between a pitch in the set of pitches and the standard pitch; determining the set of matching parameters based on a predetermined pitch weighting coefficient, the set of note difference parameters and the set of interval distance parameters, wherein the pitch weighting coefficient is used to indicate the degree of influence of the note difference parameters and the interval distance parameters on the matching parameters respectively.
[0009] Optionally, the device is used to determine the set of matching parameters by comparing each pitch in the set of pitches with the standard pitch in the following manner: comparing each pitch in the set of pitches with the standard pitch to determine a first set of matching parameters, wherein the first set of matching parameters represents the degree of matching between the first singing audio corresponding to the i-th frame and the standard singing audio when the target music is played to the i-th frame, where i is a positive integer; using the first set of matching parameters and a second set of matching parameters to determine the set of matching parameters, wherein the second set of matching parameters represents the degree of matching between the first singing audio in the first set of singing audio corresponding to the in-th frame and the standard singing audio when the target music is played to the in-th frame, where n≤i, and n is a positive integer.
[0010] Optionally, the device is used to determine the set of matching parameters by comparing each pitch in the set of pitches with the standard pitch in the following manner: when the target music is played to the j-th frame, a comparison operation is performed on each pitch in the set of pitches with the standard pitch to determine a set of initial matching parameters corresponding to the j-th frame, wherein the set of initial matching parameters represents the initial matching degree between the first singing audio and the standard singing audio corresponding to the j-th frame when the target music is played to the j-th frame, j > 1, j is a positive integer; the set of matching parameters is determined using the set of initial matching parameters and a set of historical matching parameters, wherein the historical matching parameters correspond to the jm-th frame of the target music, and the set of matching parameters corresponds to the j-th frame of the target music, m < j, m is a positive integer.
[0011] Optionally, the apparatus is configured to determine the set of matching parameters using the set of initial matching parameters and the set of historical matching parameters in the following manner: obtaining a first smoothing coefficient; determining a set of second short-term matching parameters using the set of initial matching parameters, the first smoothing coefficient, and a set of first short-term matching parameters, wherein the set of historical matching parameters includes the set of first short-term matching parameters, the first short-term matching parameters corresponding to the jm-th frame of the target music, the second short-term matching parameters corresponding to the j-th frame of the target music, and the first smoothing coefficient indicating the degree of influence of the first short-term matching parameters and the initial matching parameters on the second short-term matching parameters; obtaining a second smoothing coefficient; and determining a set of first long-term matching parameters using the set of initial matching parameters, the second smoothing coefficient, and a set of first long-term matching parameters. A set of second long-term matching parameters is determined, wherein the set of historical matching parameters includes the set of first long-term matching parameters, the first long-term matching parameters correspond to the jm-th frame of the target music, the second long-term matching parameters correspond to the j-th frame of the target music, and the second smoothing coefficient is used to indicate the degree of influence of the first long-term matching parameters and the initial matching parameters on the second long-term matching parameters, wherein the first smoothing coefficient is less than the second smoothing coefficient; a third smoothing coefficient is obtained, and the set of second short-term matching parameters, the set of second long-term matching parameters, and the third smoothing coefficient are used to determine the set of matching parameters, wherein the third smoothing coefficient is used to indicate the degree of influence of the second short-term matching parameters and the second long-term matching parameters on the matching parameters.
[0012] Optionally, the device is used to determine a set of pitches corresponding to the first set of singing audio and a standard pitch corresponding to the standard singing audio in the following manner: inputting the first set of singing audio and the standard singing audio into a preset autocorrelation function model respectively to determine a set of fundamental frequencies and a standard fundamental frequency, wherein the preset autocorrelation function is used to calculate the similarity between the signal and its own delayed version to determine a set of fundamental frequency periods and a standard fundamental frequency period, and determining the set of fundamental frequencies and the standard fundamental frequency based on the set of fundamental frequency periods, the standard fundamental frequency period and a preset sampling rate; performing a conversion operation on the set of fundamental frequencies and the standard fundamental frequency to determine the set of pitches and the standard pitch.
[0013] Optionally, the device is configured to assign a first set of weight coefficients to the first group of singing audio based on the set of matching parameters in the following manner, and perform a mixing operation on the first group of singing audio using the first set of weight coefficients to determine the target singing audio: determining whether the set of matching parameters meets a preset matching condition; in response to a second matching parameter in the set of matching parameters meeting the preset matching condition, assigning the first set of weight coefficients to the second matching parameter according to the value of the second matching parameter, and performing a mixing operation on the first group of singing audio using the first set of weight coefficients to determine the target singing audio; in response to a third matching parameter in the set of matching parameters not meeting the preset matching condition, deleting the third matching parameter from the set of matching parameters, and assigning the first set of weight coefficients to the remaining fourth matching parameter in the set of matching parameters according to the value of the fourth matching parameter, and performing a mixing operation on the first group of singing audio using the first set of weight coefficients to determine the target singing audio.
[0014] Optionally, the device is configured to assign a first set of weight coefficients to the first group of singing audio based on the set of matching parameters in the following manner, and perform a mixing operation on the first group of singing audio using the first set of weight coefficients to determine the target singing audio: determining the value of each matching parameter in the set of matching parameters and a plurality of pre-set value intervals, wherein different value intervals in the plurality of value intervals correspond to different weight coefficients; assigning the first set of weight coefficients to the first group of singing audio according to the distribution of each matching parameter in the set of matching parameters in the plurality of value intervals, and performing a mixing operation on the first group of singing audio using the first set of weight coefficients to determine the target singing audio.
[0015] Optionally, the device is further configured to: when playing a first segment of the target music, determine the set of matching parameters corresponding to each frame in the first segment, wherein the first segment represents a segment pre-labeled for the target music; assign a first set of weighting coefficients to the first group of vocal audio using the set of matching parameters, perform a mixing operation on the first group of vocal audio using the first set of weighting coefficients, and determine a first target vocal audio; when playing a second segment of the target music on the target application, directly perform a mixing operation on the first group of vocal audio using the first set of weighting coefficients, and determine a second target vocal audio, wherein the first segment is different from the second segment.
[0016] Optionally, the device is configured to acquire the first set of vocal audio in the following manner: acquiring the first set of vocal audio during the playback of the target music on the target application, wherein the first set of vocal audio corresponds to a first set of accounts, and the first set of vocal audio includes audio input by each account in the first set of accounts for the target music; or acquiring the first set of vocal audio during the playback of the target music on the target application, wherein the first set of vocal audio corresponds to a target account, and the first set of vocal audio includes audio input by the target account for the target music at different times.
[0017] Optionally, the device is configured to assign a first set of weight coefficients to the first group of singing audio based on the set of matching parameters in at least one of the following ways: perform a mixing operation on the first group of singing audio using the first set of weight coefficients; after determining the target singing audio, obtain a second group of singing audio corresponding to the first group of accounts when playing similar music on the target application; wherein the second group of singing audio includes audio input by each account in the first group of accounts for the similar music, and the similar music refers to music marked as being of the same type as the target music; assign the first set of weight coefficients to the second group of singing audio based on the set of matching parameters; and perform a mixing operation on the first group of singing audio using the first set of weight coefficients. A mixing operation is performed on the second group of singing audio to determine a third target singing audio. When the target music is played on the target application, a third group of singing audio corresponding to the second group of accounts is obtained, wherein the third group of singing audio includes the audio input by each account in the second group of accounts for the target music. When the first account in the first group of accounts and the second account in the second group of accounts have the same tag, the matching parameters determined for the first account are used as the matching parameters for the second account to determine the second group of weight coefficients corresponding to the second group of accounts. The second group of weight coefficients are used to perform a mixing operation on the third group of singing audio to determine a fourth target singing audio.
[0018] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described audio processing method when it is run.
[0019] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio processing method described above.
[0020] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described audio processing method through the computer program.
[0021] In this embodiment, target music is played on the target application, and a first group of accounts is allowed to sing together. The singing audio of the first group of accounts is obtained and compared with a standard singing audio to determine matching parameters. Then, weight coefficients are assigned to each account based on the matching coefficients to perform a mixing operation on the singing audio to obtain the target singing audio. This realizes automatic mixing of multiple people singing together, making the voices of the participants more harmonious and unified, thereby improving the quality and listening experience of the music chorus. In turn, it achieves the goal of coordinating and consistent singing audio between multiple accounts, realizing the technical effect of improving mixing accuracy, and thus solving the technical problem of poor mixing accuracy caused by the lack of coordination between multiple singing audios. Attached Figure Description
[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 This is a schematic diagram of an application environment for an optional audio processing method according to an embodiment of this application;
[0024] Figure 2 This is a schematic flowchart of an optional audio processing method according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of an optional audio processing method according to an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0027] Figure 5 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0028] Figure 6 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0030] Figure 8 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0031] Figure 9 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0032] Figure 10 This is a schematic diagram of another optional audio processing method according to an embodiment of this application;
[0033] Figure 11 This is a schematic diagram of the structure of an optional audio processing apparatus according to an embodiment of this application;
[0034] Figure 12 This is a schematic diagram of the structure of an optional audio processing product according to an embodiment of this application;
[0035] Figure 13 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] The present application will be described below with reference to embodiments:
[0039] According to one aspect of the embodiments of this application, an audio processing method is provided. Optionally, in this embodiment, the above-described audio processing method can be applied to, for example... Figure 1 The hardware environment shown consists of server 101 and terminal device 103. For example... Figure 1 As shown, server 101 is connected to terminal 103 via a network and can be used to provide services to terminal devices or applications installed on terminal devices. The applications can be video applications, instant messaging applications, browser applications, educational applications, game applications, etc. Database 105 can be set up on the server or independently of the server to provide data storage services for server 101, such as a game data storage server. The network mentioned above can include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks, metropolitan area networks, and wide area networks. The wireless network includes Bluetooth, WIFI, and other networks that enable wireless communication. Terminal device 103 can be a terminal configured with an application, and can include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, handheld computers, MID (Mobile Internet Devices), PADs, desktop computers, smart TVs, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, virtual reality (VR) terminals, augmented reality (AR) terminals, mixed reality (MR) terminals, and other computer devices. The server mentioned above can be a single server, a server cluster composed of multiple servers, or a cloud server.
[0040] Combination Figure 1As shown, the above-mentioned audio processing method can be executed by an electronic device, which can be a terminal device or a server. That is, the embodiments of this application can also be implemented by a terminal device or a server respectively, or by a terminal device and a server together.
[0041] The above is merely an example, and this embodiment does not impose any specific limitations.
[0042] Alternatively, as an optional implementation, such as Figure 2 As shown, the above audio processing method includes:
[0043] For example, the embodiments of this application can be applied to various scenarios such as audio and video, cloud technology, and artificial intelligence.
[0044] S202, Obtain the first set of singing audio, wherein the first set of singing audio includes multiple audio inputs for the target music;
[0045] Optionally, in this application embodiment, the target music may include, but is not limited to, music that is allowed to be played on the target application, and the target application may include, but is not limited to, social applications, music live streaming platforms, etc.
[0046] It should be noted that the aforementioned target music may also include, but is not limited to, songs played using a music player, music videos (MVs) played using a video player, or radio and podcasts.
[0047] In one exemplary embodiment, Figure 3 This is a schematic diagram of an optional audio processing method according to an embodiment of this application. Accounts A and B in the target application both have singing permissions, which may include, but are not limited to, [examples of such methods]. Figure 3 As shown:
[0048] S302, Account A searched for a song named "A" in the target application and found the song A;
[0049] S304. Since only accounts with an account level of A or above can sing song A, account A has reached the A level and is confirmed to be able to sing song A, and sends a real-time duet request message to account B.
[0050] S306, Account B receives a request message for a real-time duet and checks whether it can sing song A;
[0051] S308, Account B's account level has not reached level A, therefore, Account B cannot sing song A. The target application's background automatically prompts Account A that it cannot sing a duet with Account B, and proceeds to step S316.
[0052] S310, Account B confirms that it can sing song A, execute S312;
[0053] S312, when account A or account B clicks the play button, song A will start playing synchronously on the terminal device logged in by account A and the terminal device logged in by account B;
[0054] S314, Account A and Account B begin singing together using the microphones on their respective terminal devices;
[0055] S316, End.
[0056] In one exemplary embodiment, Figure 4 This is a schematic diagram of another optional audio processing method according to an embodiment of this application. The first group of accounts mentioned above... Figure 4 As shown, each singer is used by singers in different regions. Each singer logs in to the target application using an account from the first group of accounts, then confirms joining the chorus, or clicks the message to join the chorus after receiving an invitation from another account in the first group of accounts. Figure 5 This is a schematic diagram of an optional audio processing method according to an embodiment of this application, wherein the target application's chorus interface is displayed as follows. Figure 5 As shown, this includes clicking to play the target music, a start chorus button, a re-record button, a pause button, and a lyrics display area.
[0057] In another exemplary embodiment, Figure 6 This is a schematic diagram of another optional audio processing method according to an embodiment of this application. The chorus interfaces of the first group of accounts can be the same or different. If they are the same, the chorus interface can be as follows: Figure 6 As shown in the chorus interface 602, any account can click to play the target music and start the chorus; or, if one of the accounts in this first group is the chorus organizer, only that account's chorus interface will display buttons for clicking play, stopping playback, and repeating the target music recording. The chorus interface can be as follows: Figure 6 The chorus interface shown in 602 is as described in the example. However, the chorus interfaces of other participating accounts do not display the button to click and play the target music. The chorus interface can be as follows: Figure 6 As shown in 604, only the lyrics area is displayed to achieve efficient management of the recorded chorus.
[0058] S204, determine a set of matching parameters based on the first set of singing videos and the corresponding standard singing videos, wherein the standard singing audio is a singing audio that is pre-set for the target music, and the first matching parameter in the set of matching parameters is used to indicate the degree of matching between the first singing audio in the first set of singing audio and the standard singing audio;
[0059] Optionally, in this embodiment, the aforementioned standard singing audio refers to a pre-set audio sample with high-quality recording and accurate pitch, which can be used as a benchmark for comparing the first set of singing audio to evaluate the pitch accuracy of the first set of singing audio. The aforementioned set of matching parameters refers to parameters used to measure the degree of matching between the first singing audio in the first set of singing audio and the standard singing audio. The aforementioned pitch is used to describe the highness or lowness of the sound in the audio and can be determined by the vibration frequency. The higher the frequency, the higher the pitch; the lower the frequency, the lower the pitch.
[0060] For example, firstly, pitch analysis is performed on the first singing audio and the standard singing audio, then the pitches of the two are compared, the difference between the first singing audio and the standard singing audio is calculated, and the value of the first matching parameter is determined based on the difference between the two. The first matching parameter may be expressed by percentage, fraction or other forms of measurement.
[0061] In one exemplary embodiment, Figure 7 This is a schematic diagram of another optional audio processing method according to an embodiment of this application. The process for determining the first matching parameter may include, but is not limited to, the following: Figure 7 As shown:
[0062] S702, acquire the first set of singing audio and standard singing audio;
[0063] S704, compare the detected pitch in each singing audio in the first group of singing audio with the standard pitch in the standard singing audio to obtain a set of pitch difference values;
[0064] S706 converts this set of pitch differences into a set of note differences and interval distances through pitch conversion: Note difference Tdif = 1 - (abs((T%12) - (Tst%12))) / 12; Interval distance Tdist = 1 - min(4, abs((int)(T / 12) - (int)(Tst / 12))) / 4, where "int" means taking only the integer part, T refers to the detected pitch, Tst refers to the standard pitch, and T = 12 * log2(F / 440) + 69 (F refers to the audio acquisition frequency);
[0065] S708, calculates the relative pitch matching value of each singing audio in the first group of singing audio based on the note difference and interval distance:
[0066] The relative pitch matching value Mcur = 0.5 * (a * Tdif + (1 - a) * Tdist), where a is the pitch weight, which is a constant less than 1, for example, a = 0.75;
[0067] S710, calculate the short-time relative pitch matching value Mst and the long-time relative pitch matching value Mlt for each vocal audio in the first group of vocal audio:
[0068] Mst(i)=b*Mst(ii)+(1-b)*Mcur;
[0069] Mlt(i)=c*Mlt(ii)+(1-c)*Mcur;
[0070] Mf(i)=d*Mst(i)+(1-d)*Mlt(i);
[0071] Where i represents the value corresponding to the current frame, i-1 represents the value corresponding to the previous frame, b, c, and d are smoothing coefficients, which are constants less than 1, for example b = 0.7, c = 0.995, and d = 0.5. The initial frame corresponds to a value of 0, that is, there is no singing audio.
[0072] S712, using the short-time relative pitch matching value Mst and the long-time relative pitch matching value Mlt to determine the pitch matching value Mf:
[0073] Sort the pitch matching values of each singing audio from largest to smallest to obtain the sorted pitch matching values (the above set of matching parameters). The higher the pitch matching value, the higher the singing level, while the lower the ranking, the lower the singing level. One matching value in this set of pitch matching values corresponds to one singing audio in the first set of singing audio.
[0074] S714, Query the matching value corresponding to the first singing audio from this set of pitch matching values as the first matching parameter.
[0075] S206, assign a first set of weight coefficients to the first set of singing audio based on a set of matching parameters, perform a mixing operation on the first set of singing audio using the first set of weight coefficients, and determine the target singing audio.
[0076] For example, by means of Figure 7 After the operation shown, the above set of matching parameters can be obtained. Then, each matching parameter in this set of matching parameters is used to assign a weight coefficient to each singing audio in the first set of singing audio, which is the first set of weight coefficients.
[0077] Optionally, in the embodiments of this application, the values of each weight coefficient in the first group of weight coefficients may be, but are not limited to, the same or different.
[0078] Furthermore, different singing audios with different matching parameter values will be assigned different weight coefficients, and then the above mixing operation will be performed. This mixing operation may include, but is not limited to, adjusting the volume, balancing the pitch and rhythm of each audio, etc., to obtain the final target singing audio.
[0079] In an exemplary embodiment, a matching threshold Thrd is first set. If the matching parameter of the singing audio is less than Thrd, it indicates that the singing audio differs significantly from the standard audio, and the singing audio will negatively affect the final chorus effect. Therefore, the weight coefficient of the singing audio is set to the minimum weight coefficient Wmin. Then, the matching parameters are sorted according to their values. Singing audios that exceed the matching threshold Thrd are assigned corresponding weight coefficients according to their value order. For example, the weight coefficients of the singing audios corresponding to the top 25% of the matching coefficients in the above set of matching parameters are all set to 0.4, the weight coefficients of the singing audios corresponding to the matching coefficients ranked between 25% and 60% are set to 0.2, and the mixing weights of the remaining singers are all set to 0.08.
[0080] It should be noted that the above method for determining the weighting coefficients is only one example, and may also include, but is not limited to, determining the weighting coefficients of each singer's audio based on the similarity of timbre, the rhythmic synchronization of the audio, the volume consistency of the audio, user evaluations or votes, the clarity and noise level of the audio.
[0081] In one exemplary embodiment, timbre analysis technology is used to compare the timbre characteristics of the singing audio with those of a standard audio; the higher the similarity, the higher the weighting coefficient.
[0082] In yet another exemplary embodiment, the weighting coefficient is increased by analyzing whether the rhythm of the singing audio is synchronized with the rhythm of the standard audio. The better the synchronization, the higher the weighting coefficient.
[0083] In another exemplary embodiment, the weighting factor is higher when the volume of the singing audio is compared with the volume of the standard audio. The more consistent the volumes are, the higher the weighting factor.
[0084] In another exemplary embodiment, listeners may be allowed to evaluate or vote on the singing audio in real time, and the weighting coefficient may be adjusted based on the evaluation or voting results. Alternatively, the clarity and background noise level of the singing audio may be analyzed, with higher clarity and lower noise resulting in a higher weighting coefficient.
[0085] Furthermore, a mixing operation is performed on the first set of vocal audio using the first set of weighting coefficients. This mixing operation can be represented as follows: The target singing audio is obtained as described above, where j is the singer (account) number, M is the total number of singers, s(i,j) represents the audio signal in the singing audio of the j-th singer in the i-th frame, and W(j) is the weight coefficient of the j-th singer.
[0086] For example, the embodiments of this application can be applied to real-time chorus scenarios on music social platforms. Specifically, this may include, but is not limited to, processing the singer's audio using the aforementioned audio processing method in each frame during the chorus time of the chorus song (target music), or processing the singer's audio at specific time points (e.g., at the music start time frame, the climax start time frame, and the climax end time frame). Specifically: at the beginning of the chorus song, the singer's audio is processed using the aforementioned audio processing method to ensure that the initial part of the audio is synchronized with the music and that the sound quality meets expectations; at the beginning of the climax of the music, the singer's audio is processed using the aforementioned audio processing method, such as increasing reverb or adjusting the volume, to enhance emotional expression; at the end of the climax, the reverb is gradually reduced or the volume is adjusted to allow the audio to smoothly transition to the next part of the song.
[0087] For example, the aforementioned set of singing audio may include two or more singing audio. For instance, the multiple singing audio may include the aforementioned first singing audio and second singing audio. The first singing audio and the second singing audio may be sung by different accounts, and may be the same music segment or different music segments of the target music. This application does not limit this. That is, the aforementioned set of singing audio may be different accounts singing the same song.
[0088] For example, in a set of singing audios that includes two singing audios, singing audio A (first singing audio) and singing audio B (second singing audio) are compared with standard singing audios respectively. Based on the comparison results of singing audio A and singing audio B, the corresponding weight coefficients are configured for the terminal or account corresponding to singing audio A or singing audio B.
[0089] In one exemplary embodiment, Figure 8 This is a schematic diagram of another optional audio processing method according to an embodiment of this application. In a real-time chorus scenario on a music social platform (the aforementioned target application), the process of processing the singer's audio using the above-mentioned audio processing method may include, but is not limited to, the following: Figure 8 As shown:
[0090] S802, Start;
[0091] S804, Account A and Account B log in to the music social platform. Account A and Account B are the first group of accounts mentioned above.
[0092] S806, Account A invites Account B to sing a duet;
[0093] S808, Account B accepted Account A's invitation to sing a duet and joined the song duet;
[0094] S810, Account A clicks the play button on the music social platform display interface to start playing songs synchronously for Account A and Account B;
[0095] S812, Account A inputs singing audio A through a terminal device, and Account B inputs singing audio B through a terminal device;
[0096] S814, retrieves the pre-set standard vocal audio of the song;
[0097] S816, compare the singing audio A with the standard singing audio, and the singing audio B with the standard singing audio, respectively, to obtain the matching parameter A of the singing audio A and the matching parameter B of the singing audio B;
[0098] S818 compares the values of matching parameter A and matching parameter B;
[0099] S820, if the value of matching parameter A is greater than the value of matching parameter B, then the weight coefficient B set for the singing audio B is lower than the weight coefficient A set for the singing audio A.
[0100] S822, if the value of matching parameter A is less than the value of matching parameter B, then the weight coefficient B set for the singing audio B is higher than the weight coefficient A set for the singing audio A.
[0101] S824, based on weight coefficient A and weight coefficient B, performs a mixing operation on singing audio A and singing audio B to obtain the target singing audio;
[0102] S826, End.
[0103] For example, if the original volume of both vocal audio A and vocal audio B is 1.0, the weighting factor A is 0.7, and the weighting factor B is 0.3, then during the mixing process, the volume of vocal audio A will be adjusted to 0.7, and the volume of vocal audio B will be adjusted to 0.3. Then, these two adjusted audio signals are added together to obtain the final target vocal audio.
[0104] The above operations ensure the quality of the duet between Account A and Account B. By comparing the differences between the two accounts' performances and the standard audio recording, the weights of the two accounts in the mixing are dynamically adjusted to produce the best duet effect.
[0105] In an exemplary embodiment, in an application scenario where dubbing software (the aforementioned target application) is used for dubbing, the process of processing the dubbing audio using the aforementioned audio processing method may include, but is not limited to:
[0106] S802, Start;
[0107] S804, Account A and Account B log in to the dubbing software. Account A and Account B are the first group of accounts mentioned above.
[0108] S806, Account A invited Account B to do the same video dubbing;
[0109] S808, Account B accepted Account A's voiceover invitation and joined the video voiceover;
[0110] S810, Account A clicks the play button in the dubbing software to start playing the video to be dubbed synchronously for Account A and Account B;
[0111] S812, Account A inputs voice-over audio A through a terminal device, and Account B inputs voice-over audio B through a terminal device;
[0112] S814, retrieves the pre-set standard dubbing audio for the dubbing video;
[0113] S816, compare dubbing audio A with standard dubbing audio, dubbing audio B with standard dubbing audio, and obtain matching parameter A for dubbing audio A and matching parameter B for dubbing audio B respectively;
[0114] S818 compares the values of matching parameter A and matching parameter B, and assigns weight coefficient A and weight coefficient B to the dubbing audio A.
[0115] S820 performs a mixing operation on dubbing audio A and dubbing audio B based on weight coefficient A and weight coefficient B to obtain the target dubbing audio;
[0116] S822, End.
[0117] Through the above operations, high-quality video dubbing was achieved. The quality of the dubbing audio was optimized by comparing it with standard dubbing audio and allocating weighting coefficients to improve the accuracy and consistency of the dubbing and ensure that the final dubbing audio meets the expected quality standards.
[0118] In this embodiment, a target music is played on a target application, where the target music is set to allow chorus singing. During the playback of the target music, a first set of singing audio corresponding to a first group of accounts is obtained, wherein the first set of singing audio includes multiple audio inputs for the target music. A comparison operation is performed on the first set of singing audio and a standard singing audio pre-set for the target music to determine a set of matching parameters. The first matching parameter in the set of matching parameters indicates the degree of matching between the first singing audio in the first set of singing audio and the standard singing audio, and the first matching parameter is determined by the first pitch of the first singing audio and the standard pitch of the standard singing audio. Based on the set of matching parameters, a first set of weighting coefficients is assigned to the first set of singing audio, and then... The first group performs a mixing operation on the first group's singing audio, determining the target singing audio method by playing the target music on the target application and allowing the first group's accounts to sing together. The singing audio of the first group's accounts is obtained and compared with the standard singing audio to determine the matching parameters. Then, weight coefficients are assigned to each account based on the matching coefficients to perform a mixing operation on the singing audio, obtaining the target singing audio. This realizes automatic mixing of multiple people singing together, making the participants' voices more harmonious and unified, thereby improving the quality and listening experience of the music chorus. In turn, it achieves the goal of coordinating and consistent singing audio between multiple accounts, realizing the technical effect of improving mixing accuracy, and thus solving the technical problem of poor mixing accuracy caused by the lack of coordination between the singing audio of multiple singers.
[0119] As an optional approach, the above-mentioned determination of a set of matching parameters based on the first set of singing videos and the corresponding standard singing videos includes: determining a set of pitches corresponding to the first set of singing audio and the standard pitch corresponding to the standard singing audio, wherein the set of pitches and the standard pitch correspond to the same moment of the target music, and the first matching parameters are determined by the set of pitches, the first pitch corresponding to the first singing audio, and the standard pitch; and comparing each pitch in the set of pitches with the standard pitch to determine the set of matching parameters.
[0120] For example, after obtaining the first set of singing audio, the above detection operation can be performed on the first set of singing audio and the standard singing audio respectively, as in the embodiments of this application, to obtain the above set of pitches and standard pitches corresponding to the set of singing audio. Then, the set of pitches and standard pitches are compared to obtain the matching parameters of each singing audio. The matching parameters are used to evaluate the singer's skills, pitch accuracy, timbre, etc., and to quantify the similarity or difference between the first set of singing audio and the standard singing audio.
[0121] Optionally, in the embodiments of this application, the above-mentioned set of pitches and standard pitches can be understood as the pitch of a sound being determined by its frequency.
[0122] In an exemplary embodiment, assuming that the pitch of one of the singing audio segments in the first group of singing audio is C4 (261.63Hz) at the 5th second, and the pitch of the standard singing audio is also C4 (261.63Hz) at the 5th second, then, a pitch detection algorithm (autocorrelation method, Fourier transform method, etc.) is used to analyze the two audio segments to determine the pitch of the two audio segments at the 5th second. The pitches of the two audio segments are compared, the frequency difference is calculated, and the frequency difference is found to be 0Hz. At this time, the matching parameter can be set to 1, indicating that the pitch of the singing audio is completely consistent with the standard pitch.
[0123] In yet another exemplary embodiment, assuming the pitch detection algorithm is an autocorrelation method, the pitch can be determined by calculating the similarity between the audio signal and itself at different time delays.
[0124] In another exemplary embodiment, assuming the pitch detection algorithm is a Fourier transform method, the pitch can be determined by converting the audio signal to the frequency domain and then analyzing the spectrum.
[0125] As an optional approach, the above-mentioned comparison of each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine the above-mentioned set of matching parameters includes: performing pitch conversion operations on the above-mentioned set of pitches and the above-mentioned standard pitches respectively to determine a set of note difference parameters and a set of interval distance parameters, wherein the above-mentioned note difference parameters are used to indicate the pitch difference between a pitch in the above-mentioned set of pitches and the above-mentioned standard pitch, and the above-mentioned interval distance parameters are used to indicate the interval relationship between a pitch in the above-mentioned set of pitches and the above-mentioned standard pitch; determining the above-mentioned set of matching parameters based on a predetermined pitch weighting coefficient, the above-mentioned set of note difference parameters and the above-mentioned set of interval distance parameters, wherein the above-mentioned pitch weighting coefficient is used to indicate the degree of influence of the above-mentioned note difference parameters and the above-mentioned interval distance parameters on the above-mentioned matching parameters respectively.
[0126] For example, after obtaining the above set of pitches and standard pitches, as in the embodiments of this application, the pitch conversion operation is first performed on each pitch to obtain specific note difference parameters and interval distance parameters. Then, based on the predetermined pitch weight coefficient, note difference parameters and interval distance parameters, the matching degree between each singing audio and the standard audio is obtained, and the matching parameters are used to represent it, effectively quantifying the difference between the singing audio and the standard audio and improving the accuracy of the matching parameters.
[0127] Optionally, in the embodiments of this application, the above-mentioned note difference parameter refers to the interval between two notes in the singing audio. The interval can be divided into a whole tone (1 semitone), a semitone (0.5 semitones), etc. For example, in C major, C to D is a whole tone (1), C to E is a whole tone plus a semitone (2), and C to G is a perfect fifth (7). The note difference parameter is mainly used to describe the relative distance between notes;
[0128] Optionally, in this embodiment, the aforementioned interval distance parameter refers to the actual pitch difference between two notes in the sung audio. The interval distance parameter focuses on the absolute pitch difference between the notes. For example, C4 (middle C) to D4 is a semitone, and C3 (low C) to D3 is also a semitone. The interval distance parameter helps to analyze the actual pitch relationship between notes, rather than just focusing on their relative distance. The aforementioned pitch conversion operation refers to converting the pitch into the aforementioned set of note difference parameters and the aforementioned set of interval distance parameters, including but not limited to implementation through pitch detection algorithms (e.g., autocorrelation method) and spectral analysis.
[0129] In an exemplary embodiment, suppose the singing audio has a set of pitches: A, B, C, and a standard pitch: D. First, these pitches are converted into note values, assuming the converted values are: A(60), B(62), C(64), D(65). The note difference parameter represents the calculation of the difference between each pitch and the standard pitch D. For example: the difference between A and D: |65-60|=5, the difference between B and D: |65-62|=3, and the difference between C and D: |65-64|=1. The interval distance parameter represents the calculation of the interval relationship between each pitch and the standard pitch D. For example: the interval between A and D is 5 semitones; the interval between B and D is 3 semitones; the interval between C and D is 1 semitone. Next, weighting coefficients are set for the note difference parameter and the interval distance parameter. For example, the note difference weight is 0.7 and the interval distance weight is 0.3. For each pitch, the note difference and interval distance are weighted according to the weighting coefficients. For example, the matching parameter for A is 0.7*5+0.3*5=5.6; the matching parameter for B is 0.7*3+0.3*3=3.2; the matching parameter for C is 0.7*1+0.3*1=1.0, to obtain the above set of matching parameters.
[0130] As an optional approach, the above-mentioned comparison of each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine the above-mentioned set of matching parameters includes: comparing each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine a first set of matching parameters, wherein the above-mentioned first set of matching parameters represents the degree of matching between the first singing audio corresponding to the i-th frame and the above-mentioned standard singing audio when the target music is played to the i-th frame, where i is a positive integer; using the above-mentioned first set of matching parameters and a second set of matching parameters to determine the above-mentioned set of matching parameters, wherein the above-mentioned second set of matching parameters represents the degree of matching between the first singing audio in the above-mentioned first set of singing audio corresponding to the in-th frame and the above-mentioned standard singing audio when the target music is played to the in-th frame, where n≤i, and n is a positive integer.
[0131] For example, after obtaining the above set of pitches and standard pitches, as in the embodiments of this application, the singing audio of each frame in the target music playback process can be compared with the standard singing audio to determine the matching parameters, thereby realizing the pitch detection and correction of the target music.
[0132] Optionally, both the first set of matching parameters and the second set of matching parameters mentioned above may include several matching parameters.
[0133] In an exemplary embodiment, the matching parameter of a singing audio in the first frame is matching parameter A, and the matching parameter in the second frame is matching parameter B. Therefore, the aforementioned set of second matching parameters includes matching parameter A, and the aforementioned set of first matching parameters includes matching parameter B.
[0134] In yet another exemplary embodiment, the matching parameter of a singing audio in the first frame is matching parameter A, and the matching parameter in the eighth frame is matching parameter B. Therefore, the aforementioned set of second matching parameters includes matching parameter A, and the aforementioned set of first matching parameters includes matching parameter B.
[0135] For example, the set of matching parameters may be determined by combining the first set of matching parameters and the second set of matching parameters using a weighted average method. Assuming that the weight of the first set of matching parameters is set to 0.7 and the weight of the second set of matching parameters is set to 0.3, then one of the matching parameters in the set of matching parameters is determined as 0.7 * one matching parameter of the first set of matching parameters + 0.3 * one matching parameter of the second set of matching parameters.
[0136] As an optional approach, the above-mentioned comparison of each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine the above-mentioned set of matching parameters includes: when the target music is played to the j-th frame, performing a comparison operation on each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine a set of initial matching parameters corresponding to the j-th frame, wherein the above-mentioned set of initial matching parameters represents the initial matching degree between the first singing audio and the above-mentioned standard singing audio corresponding to the j-th frame when the target music is played to the j-th frame, j>1, j is a positive integer; using the above-mentioned set of initial matching parameters and a set of historical matching parameters to determine the above-mentioned set of matching parameters, wherein the above-mentioned historical matching parameters correspond to the jm-th frame of the target music, and the above-mentioned set of matching parameters corresponds to the j-th frame of the target music, m<j, m is a positive integer.
[0137] For example, when comparing each pitch in the above set of pitches with the standard pitch, as in the embodiments of this application, the above set of initial matching parameters corresponding to the current j-th frame can be determined first, and then the above set of matching parameters can be determined using the set of initial matching parameters and a set of historical matching parameters corresponding to the historical time period, thereby effectively improving the accuracy of the above set of matching parameters.
[0138] Optionally, in the embodiments of this application, the above-mentioned set of historical matching parameters can be understood as the matching parameters corresponding to the jm frame of the target music, which were generated earlier than the above-mentioned set of initial matching parameters.
[0139] In an exemplary embodiment, assuming the target music is played to the 10th frame (j=10), the pitch of the 10th frame is first analyzed and compared with the corresponding pitch of the standard singing audio to obtain a set of initial matching parameters. Then, a set of historical matching parameters of the 5th frame (j-5=5) is obtained. Combining this set of initial matching parameters and a set of historical matching parameters, a set of matching parameters for the 10th frame is determined.
[0140] As an optional approach, the method of determining the set of matching parameters using the aforementioned set of initial matching parameters and a set of historical matching parameters includes: obtaining a first smoothing coefficient; determining a set of second short-term matching parameters using the aforementioned set of initial matching parameters, the aforementioned first smoothing coefficient, and a set of first short-term matching parameters, wherein the aforementioned set of historical matching parameters includes the aforementioned set of first short-term matching parameters, the aforementioned first short-term matching parameters correspond to the jm-th frame of the target music, the aforementioned second short-term matching parameters correspond to the j-th frame of the target music, and the aforementioned first smoothing coefficient is used to indicate the degree of influence of the aforementioned first short-term matching parameters and the aforementioned initial matching parameters on the aforementioned second short-term matching parameters; obtaining a second smoothing coefficient; and determining a set of second short-term matching parameters using the aforementioned set of initial matching parameters, the aforementioned second smoothing coefficient, and a set of first long-term matching parameters. A second set of long-term matching parameters is determined, wherein the set of historical matching parameters includes the set of first long-term matching parameters, the first long-term matching parameters correspond to the jm-th frame of the target music, the second long-term matching parameters correspond to the j-th frame of the target music, and the second smoothing coefficient is used to indicate the degree of influence of the first long-term matching parameters and the initial matching parameters on the second long-term matching parameters, wherein the first smoothing coefficient is less than the second smoothing coefficient; a third smoothing coefficient is obtained, and the set of matching parameters is determined using the set of second short-term matching parameters, the set of second long-term matching parameters, and the third smoothing coefficient, wherein the third smoothing coefficient is used to indicate the degree of influence of the second short-term matching parameters and the second long-term matching parameters on the matching parameters.
[0141] For example, after obtaining the above set of initial matching parameters and the set of historical matching parameters, the above set of matching parameters can be generated based on the different smoothing coefficients as in the embodiments of this application. The set of matching parameters can be adjusted and optimized by different smoothing coefficients and historical data to ensure the availability and stability of the set of matching parameters.
[0142] In an exemplary embodiment, the aforementioned set of second short-term matching parameters may be expressed as, but is not limited to, as: Mst(j) = b * Mst(jm) + (1-b) * Mcur, where Mst(j) refers to any one of the second short-term matching parameters in the aforementioned set of second short-term matching parameters, b is a first smoothing coefficient, which can be set manually, including but not limited to a constant less than 1, Mst(jm) is a first short-term matching parameter in the aforementioned set of first short-term matching parameters, j is used to indicate the singing audio (or the account in the aforementioned set of accounts), Mcur refers to the relative pitch matching value, Mcur = 0.5 * (a * note difference parameter + (1-a) * interval distance parameter), the value of a can be flexibly set, and this application does not impose any limitations.
[0143] In yet another exemplary embodiment, the aforementioned set of second long-term matching parameters may be expressed as, but is not limited to: Mlt(j) = c*Mlt(jm) + (1-c)*Mcur, where Mlt(j) refers to any one of the second long-term matching parameters in the aforementioned set of second long-term matching parameters, b is a first smoothing coefficient, which can be set manually, including but not limited to a constant less than 1, and Mlt(jm) is a first long-term matching parameter in the aforementioned set of first long-term matching parameters.
[0144] For example, after determining the above set of second short-term matching parameters and the above set of second long-term matching parameters, the above set of matching parameters may be expressed as, but is not limited to: Mf(j)=d*Mst(j)+(1-d)*Mlt(j) to determine the above set of matching parameters, Mf(j) refers to any one of the matching parameters in the above set of matching parameters, where d refers to the above third smoothing parameter, which can be set manually, including but not limited to a constant less than 1.
[0145] As an optional approach, the determination of the set of pitches corresponding to the first set of singing audio and the standard pitch corresponding to the standard singing audio includes: inputting the first set of singing audio and the standard singing audio into a preset autocorrelation function model to determine a set of fundamental frequencies and a standard fundamental frequency, wherein the preset autocorrelation function is used to calculate the similarity between the signal and its own delayed version to determine a set of fundamental frequency periods and a standard fundamental frequency period, and determining the set of fundamental frequencies and the standard fundamental frequency based on the set of fundamental frequency periods, the standard fundamental frequency period, and a preset sampling rate; performing a conversion operation on the set of fundamental frequencies and the standard fundamental frequency to determine the set of pitches and the standard pitch.
[0146] For example, in the process of detecting the first set of singing audio, as in the embodiments of this application, the preset autocorrelation function model can be used to determine the set of fundamental frequencies and the standard fundamental frequency. Then, the set of pitches and the standard pitch can be determined using the set of fundamental frequencies and the standard fundamental frequency, thereby realizing pitch detection and matching. By comparing the difference between the set of fundamental frequencies and the standard fundamental frequency, the pitch in the singing audio can be accurately determined, thereby realizing the evaluation and analysis of the singing audio.
[0147] Optionally, in the embodiments of this application, one of the fundamental frequencies in the above-mentioned set of fundamental frequencies refers to the frequency of one of the singing audios in the above-mentioned first set of singing audios. Similarly, the above-mentioned standard fundamental frequency refers to the frequency of the above-mentioned standard singing audio. The above-mentioned fundamental period is the longest repetition period in a periodic signal. The fundamental frequency is the reciprocal of the fundamental period. The fundamental period includes, but is not limited to, the autocorrelation method and the Fourier transform method. The above-mentioned preset autocorrelation function model can be represented by the above-mentioned preset autocorrelation function. The preset autocorrelation function can be expressed as: R(k)=Σ(x(n)*x(nk))), where k is the delay, x(n) is the audio signal of the singing audio, and n is the time index in the audio signal x(n), representing each time point in the discrete time sequence of the signal.
[0148] Furthermore, first remove the audio signal with k=0 from the preset autocorrelation function, find the delay k corresponding to the maximum remaining output value in the preset autocorrelation function, and the delay k is the pitch period, while the fundamental frequency F can be determined as: F=fs / k (fs is the current sampling rate of the audio signal, which can be obtained directly).
[0149] For example, taking one of the base frequencies in the above set as an example, Figure 9 This is a schematic diagram of another optional audio processing method according to an embodiment of this application. The step of determining the fundamental frequency may include, but is not limited to, the following: Figure 9 As shown:
[0150] S902, acquire the audio signal of one of the singing audios in the above set of singing audios;
[0151] S904 performs preprocessing on audio signals, including but not limited to noise removal and filtering.
[0152] S906 calculates the autocorrelation output value R(k) of the audio signal using a preset autocorrelation function (R(k)=Σ(x(n)*x(nk))).
[0153] S908, when R(k) takes the maximum value, k is determined as the fundamental period;
[0154] S910 defines the fundamental frequency as the ratio of the sampling rate to the pitch period.
[0155] In an exemplary embodiment, assuming that the preset autocorrelation function model of the first group of singing audio calculates that the fundamental frequency of one of the singing audios is 100Hz, while the standard fundamental frequency is 110Hz, then the fundamental frequency is converted, for example, 100Hz is converted to C major (261.63Hz), and 110Hz is converted to D major (293.66Hz) to obtain the pitch corresponding to the fundamental frequency. The pitch corresponding to the singing audio is C major, while the standard pitch is D major.
[0156] As an optional approach, the above-mentioned method of assigning a first set of weight coefficients to the first set of singing audio based on the aforementioned set of matching parameters, and performing a mixing operation on the first set of singing audio using the aforementioned first set of weight coefficients to determine the target singing audio includes: determining whether the aforementioned set of matching parameters meets a preset matching condition; in response to the second matching parameter in the aforementioned set of matching parameters meeting the aforementioned preset matching condition, assigning a first set of weight coefficients to the second matching parameter according to the value of the second matching parameter, performing a mixing operation on the first set of singing audio using the aforementioned first set of weight coefficients to determine the target singing audio; in response to the third matching parameter in the aforementioned set of matching parameters not meeting the aforementioned preset matching condition, deleting the third matching parameter from the aforementioned set of matching parameters, and assigning the aforementioned first set of weight coefficients to the remaining fourth matching parameter in the aforementioned set of matching parameters according to the value of the fourth matching parameter, performing a mixing operation on the first set of singing audio using the aforementioned first set of weight coefficients to determine the target singing audio.
[0157] For example, after determining the above set of matching parameters, the matching parameters that can participate in the weight parameter allocation can be determined using the above preset matching conditions, as in the embodiments of this application. Then, the above mixing operation is performed on the singing audio corresponding to the matching parameters that meet the preset matching conditions. That is, in the process of processing the above set of matching parameters, matching parameters that do not meet the conditions are deleted, and weight coefficients are allocated to the remaining parameters to achieve the purpose of finally determining the target singing audio and ensuring the chorus quality of the target singing audio.
[0158] Optionally, in the embodiments of this application, the above-mentioned preset matching conditions can be flexibly set, and this application does not impose any limitations.
[0159] In an exemplary embodiment, the 20% of matching parameters with the largest values in the above set of matching parameters are considered to meet the above preset matching conditions, and the above first set of weight coefficients can be assigned to these 20% of matching parameters.
[0160] In another exemplary embodiment, the 10% of matching parameters with smaller values in the above set of matching parameters are considered not to meet the above preset matching conditions. The third matching parameter refers to these 10% of matching parameters. After deleting the third matching parameter in the set of matching parameters, the first set of weight coefficients is assigned to the remaining 90% of the fourth matching parameters in the above set of matching parameters.
[0161] As an optional approach, the above-mentioned method of assigning a first set of weight coefficients to the first set of singing audio based on the aforementioned set of matching parameters, and performing a mixing operation on the first set of singing audio using the aforementioned first set of weight coefficients to determine the target singing audio includes: determining the value of each matching parameter in the aforementioned set of matching parameters and multiple pre-set value intervals, wherein different value intervals in the multiple value intervals correspond to different weight coefficients; assigning a first set of weight coefficients to the first set of singing audio according to the distribution of each matching parameter in the aforementioned set of matching parameters in the multiple value intervals, and performing a mixing operation on the first set of singing audio using the aforementioned first set of weight coefficients to determine the target singing audio.
[0162] For example, in the process of allocating the first set of weight coefficients, as in the embodiments of this application, a corresponding weight coefficient can be assigned to each matching parameter by pre-setting multiple value ranges and the values of matching parameters. Then, the singing audio is mixed using these weight coefficients to finally determine the target singing audio. The first set of weight coefficients can be flexibly adjusted to obtain the best mixing effect.
[0163] In an exemplary embodiment, the aforementioned value range includes value range A, value range B, and value range C. Then, if the value of the matching parameter falls into value range A, a weight coefficient A is set for the matching parameter; if the value of the matching parameter falls into value range B, a weight coefficient B is set for the matching parameter; and if the value of the matching parameter falls into value range C, a weight coefficient C is set for the matching parameter. The weight coefficients A, B, and C can be the same or different values.
[0164] For example, the setting of the above value range may include, but is not limited to, things related to the target music, the number of singers involved, etc.
[0165] In an exemplary embodiment, the setting of the above-mentioned value intervals is related to the target music. For example, if the target music is difficult, resulting in low values for each matching parameter in the above set of matching parameters, a smaller number of value intervals can be set. For example, only two value intervals can be set. Value interval 1 consists of the maximum and median values in this set of matching parameters, represented as (median value, maximum value). Value interval 2 consists of the minimum and median values in this set of matching parameters, represented as [minimum value, median value]. The 50% of matching parameters with larger values in value interval 1 are assigned a weight of 0.6, while the 50% of matching parameters with larger values in value interval 2 are assigned a weight of 0.4. This allows for further optimization and adjustment of the value interval settings and weight allocation based on the actual application scenario, ensuring the robustness and adaptability of the weight coefficients.
[0166] As an optional solution, the above-mentioned acquisition of the first set of singing audio includes: when playing the first segment of the target music, determining the above-mentioned set of matching parameters corresponding to each frame in the first segment, wherein the first segment represents the segment pre-marked for the target music; assigning a first set of weight coefficients to the first set of singing audio using the above-mentioned set of matching parameters; performing a mixing operation on the first set of singing audio using the above-mentioned set of weight coefficients; and determining the first target singing audio.
[0167] When the second segment of the target music is played on the target application, the first set of weighting coefficients is used to directly perform a mixing operation on the first set of vocal audio to determine the second target vocal audio, wherein the first segment is different from the second segment.
[0168] For example, the above-mentioned audio processing method can also be as described in the embodiments of this application, firstly determining the matching parameters corresponding to each singing audio in the first paragraph of the target music, and using these parameters to assign weight coefficients to the singing audio. Then, when singing the second paragraph, the mixing operation is directly performed using the weight coefficients determined in the first paragraph, without needing to re-obtain the weight coefficients, simplifying the operation process, thereby improving the above-mentioned mixing processing efficiency.
[0169] Optionally, in the embodiments of this application, the first paragraph may include, but is not limited to, the beginning paragraph of the target music, the beginning paragraph of the climax, the ending paragraph of the climax, the leading-to-ending paragraph, etc., and the second paragraph is different from the first paragraph.
[0170] In an exemplary embodiment, assuming that the first segment is the starting segment of the target music and the first segment lasts for 10 frames, the matching parameters and weight parameters can be determined in each of these 10 frames. The weight coefficients of each vocal audio in the vocal audio determined in these 10 frames are determined as the first set of weight coefficients. Then, in at least one segment of the subsequent music segment (the second segment mentioned above), the first set of weight coefficients is used to directly mix the vocal audio without having to re-determine the matching parameters and weight parameters, further saving mixing time.
[0171] In another exemplary embodiment, the target music may include, but is not limited to, multiple first segments. Assuming the first segment is the starting segment and the climax starting segment of the target music, the matching parameters and weight parameters can be determined in each frame of the starting segment of the target music. The weight coefficients of each vocal audio in each frame of the starting segment are determined as the first set of weight coefficients. Then, the music segment between the starting segment and the climax starting segment of the target music will be determined as the second segment. The first set of weight coefficients is used directly to mix the vocal audio. Next, the matching parameters and weight parameters are determined again in each frame of the climax starting segment of the target music. The weight coefficients of each vocal audio in each frame of the climax starting segment are re-determined as the first set of weight coefficients. The first set of weight coefficients is used to mix the vocal audio in at least one segment (the second segment) of the subsequent music segment.
[0172] As an optional solution, obtaining the first set of vocal audio includes: obtaining the first set of vocal audio during the playback of the target music on the target application, wherein the first set of vocal audio corresponds to a first set of accounts, and the first set of vocal audio includes audio input by each account in the first set of accounts for the target music; or obtaining the first set of vocal audio during the playback of the target music on the target application, wherein the first set of vocal audio corresponds to a target account, and the first set of vocal audio includes audio input by the target account for the target music at different times.
[0173] Optionally, in this embodiment of the application, the first group of accounts refers to a group of user accounts authorized to play and sing target music on the target application. Each account in the first group of accounts starts to sing, and the input audio of each account is the audio in the first group of singing audio.
[0174] It is understandable that the audio input from the first group of accounts mentioned above may include, but is not limited to, the actual singing audio of the song performed by the singer using the account, and may also include interference audio from the singer's environment, such as background noise and echoes. Background noise and echoes can be pre-denoised, and the audio after noise removal can be used as the first group of singing audio mentioned above.
[0175] For example, the first group of accounts may include, but is not limited to, meeting certain preset conditions, such as having singing permissions enabled, or users who have successfully registered in the aforementioned target application have singing permissions enabled by default. The aforementioned target application can provide users with a duet function, allowing users to sing with other users synchronously or asynchronously, including but not limited to real-time duets and offline duets.
[0176] It should be noted that there is no limit to the number of accounts in the first group mentioned above. These accounts can be not only accounts that have been successfully registered in the target application mentioned above, but also accounts that have been successfully registered in a third-party application. The third-party application and the target application mentioned above have a cooperative or integrated relationship. This cooperative or integrated relationship allows users to log in to or use the target application using an account registered in the third-party application.
[0177] As an optional approach, after assigning a first set of weighting coefficients to the first set of singing audio based on the aforementioned set of matching parameters, performing a mixing operation on the first set of singing audio using the aforementioned first set of weighting coefficients, and determining the target singing audio, the above method further includes at least one of the following:
[0178] When playing similar music on the target application, obtain the second set of singing audio corresponding to the first set of accounts. The second set of singing audio includes audio input by each account in the first set of accounts for the similar music. The similar music refers to music that is marked as being of the same type as the target music. Assign the first set of weight coefficients to the second set of singing audio based on the first set of matching parameters. Perform a mixing operation on the second set of singing audio using the first set of weight coefficients to determine the third target singing audio.
[0179] When the target music is played on the target application, the third set of singing audio corresponding to the second set of accounts is obtained. The third set of singing audio includes the audio input by each account in the second set of accounts for the target music. When the first account in the first set of accounts and the second account in the second set of accounts have the same tag, the matching parameters determined for the first account are used as the matching parameters for the second account to determine the second set of weight coefficients corresponding to the second set of accounts. The third set of singing audio is mixed using the second set of weight coefficients to determine the fourth target singing audio.
[0180] For example, after obtaining the target singing audio, as in the embodiments of this application, when playing similar music on the target application, the first set of weight coefficients can be directly used to mix the second set of singing audio of the same music sung by the first group of accounts. For the second account with the same tag as the first account, the weight coefficients corresponding to the first account can be directly used for mixing, without having to re-determine the weight parameters using the singing audio of the second account. This achieves a more efficient audio mixing and playback process, saves audio processing time and computing resources, and improves the user experience for the singer.
[0181] Optionally, in this application embodiment, the aforementioned similar music may include, but is not limited to, music with the same rhythm, style, melody, creation period, cultural background, author, etc. as the aforementioned target music, and the audio input by the aforementioned first group of accounts may include, but is not limited to, real-time captured audio, pre-recorded audio, etc.
[0182] In an exemplary embodiment, suppose an account participates in a chorus of a rock song A, and the account selects another rock song B from the same music list to participate in a chorus. In this case, during the chorus of the aforementioned rock song B, the weight parameter corresponding to the account in rock song A can be directly used as the weight parameter corresponding to the account in rock song B to achieve the above-mentioned mixing operation.
[0183] Optionally, in the embodiments of this application, the aforementioned identical label refers to two accounts having the same attributes or classifications in certain characteristics. These characteristics may include, but are not limited to, the account's interests, preferences, timbre, volume, performance of the terminal device used, etc.
[0184] In an exemplary embodiment, assuming that both the first account and the second account are skilled at singing upbeat music, both the first account and the second account have the tag "Skilled music style: upbeat". When the second account participates in a duet of an upbeat song, it can be determined whether the first account has participated in the duet of the song before. If it has, the weight coefficient corresponding to the first account is directly used as the weight coefficient corresponding to the second account for mixing operations.
[0185] For example, the audio processing method proposed in this application can be applied to the application scenario of online song chorus. Specifically, online chorus refers to the function of multi-user real-time chorus by using the Internet through the recording and broadcasting equipment and singing processing software of each user's local terminal. Through the embodiments of this application, the problem of poor synthesized sound effect caused by the difference in pitch and rhythm of the singers in online chorus is solved.
[0186] Specifically, in this embodiment, the pitch of each singer is first detected in real time and compared with the standard pitch. Then, short-term relative matching degree and long-term relative matching degree are calculated. All singers are ranked according to the short-term and long-term relative matching degrees. Singers with high matching degrees are assigned higher mixing weights (as mentioned above weight parameters), while singers with low matching degrees are assigned lower mixing weights. That is, while each singer is recording in real time using local devices, the local devices or servers will detect the relative pitch of the singer's voice in real time and provide its matching degree with the standard pitch. The chorus mixing is adjusted based on the relative pitch matching degree to achieve the optimal chorus effect, thereby achieving the technical effect of adaptively controlling the overall chorus effect and avoiding the impact of individual singers' pitch and rhythm problems on the overall chorus effect.
[0187] In one exemplary embodiment, Figure 10 This is a schematic diagram of another optional audio processing method according to an embodiment of this application, such as... Figure 10 As shown, each singer in the chorus records their voice using local audio acquisition devices (such as mobile phones, microphones, etc.). The recording signal (the audio signal of the singing) is processed by frame segmentation (the continuous recording signal is divided into frames every 20ms). Each frame of the audio signal is processed by a fundamental frequency detection algorithm to detect the current fundamental frequency value. The absolute pitch is then calculated using the fundamental frequency to pitch conversion formula. Alternatively, a standard pitch file for the current song can be retrieved from a song server. This file provides the standard pitch values at different points in the song. By comparing the currently detected pitch value with the standard pitch value at the current moment, the relative pitch can be calculated, including the note difference Tdif and the interval distance Tdist. Based on Tdif and Tdist... ist calculates the current relative pitch matching value Mcur, and then calculates the short-term relative pitch matching value Mst and the long-term relative pitch matching value Mlt for each singer, finally obtaining the comprehensive pitch matching value Mf. The singers' Mf values are sorted from largest to smallest, with higher rankings indicating relatively higher singing levels and lower rankings indicating relatively lower singing levels. Therefore, singers with higher rankings will be assigned a larger mixing weight value during choral mixing, while singers with lower rankings will be assigned a smaller mixing weight value. The goal is to give singers with higher singing levels ample opportunity to showcase their skills, while singers with lower singing levels will be significantly weakened to avoid interfering with the choral effect, ultimately achieving an ideal choral effect.
[0188] Furthermore, the mixing weighting coefficients (or mixing weight coefficients) for each singer in the final mixing process are strongly correlated with their respective matching parameters. The higher the matching parameter value, the greater the mixing weight. Here, we first sort the singers by their matching parameter values from largest to smallest. We can also set a matching threshold, Thrd. If a singer's matching parameter value is less than Thrd, it indicates that the singer is far from the standard, and their voice will negatively impact the final synthesized effect. Therefore, the mixing weight for that singer is set to the minimum value, Wmin, for example, Wmin = 0.01. Singers whose matching parameter values exceed Thrd are assigned corresponding mixing weights according to their ranking. For example, the top 25% of singers have a mixing weight of 0.4, singers ranked between 25% and 60% have a mixing weight of 0.2, and the remaining singers have a mixing weight of 0.08. Then, after calculating the mixing weights for each singer, the mixing process is performed... To obtain the final choral mixed audio result (the target singing audio mentioned above).
[0189] It should be noted that the audio detection algorithm implemented using the audio processing method proposed in this application can be deployed on a local client or a mixing server. When deployed on a local client, the input audio signal required by the detection algorithm can be directly obtained from the local device. However, when the detection algorithm is deployed on a mixing server, the input audio signal required by the detection algorithm needs to be obtained by the server first decoding and encoding the singing audio from the client.
[0190] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0191] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0192] According to another aspect of the embodiments of this application, an audio processing apparatus for implementing the above-described audio processing method is also provided. For example... Figure 11 As shown, the device includes:
[0193] The acquisition module 1102 is used to acquire a first set of singing audio, wherein the first set of singing audio includes multiple audios input for the target music;
[0194] The determination module 1104 is used to assign a first set of weight coefficients to the first set of singing audio based on a set of matching parameters, perform a mixing operation on the first set of singing audio using the first set of weight coefficients, and determine the target singing audio.
[0195] The execution module 1106 is used to assign a first set of weight coefficients to the first set of singing audio based on a set of matching parameters, perform a mixing operation on the first set of singing audio using the first set of weight coefficients, and determine the target singing audio.
[0196] As an optional solution, the above-mentioned device is used to determine a set of matching parameters based on the first set of singing videos and the corresponding standard singing videos in the following manner: determining a set of pitches corresponding to the first set of singing audios and a standard pitch corresponding to the standard singing audios, wherein the set of pitches and the standard pitches correspond to the same moment of the target music, and the first matching parameters are determined by the set of pitches, the first pitch corresponding to the first singing audios, and the standard pitches; comparing each pitch in the set of pitches with the standard pitches to determine the set of matching parameters.
[0197] As an optional solution, the above-mentioned device is used to determine the set of matching parameters by comparing each pitch in the set of pitches with the standard pitch in the following manner: performing pitch conversion operations on the set of pitches and the standard pitch respectively to determine a set of note difference parameters and a set of interval distance parameters, wherein the note difference parameters are used to indicate the pitch difference between a pitch in the set of pitches and the standard pitch, and the interval distance parameters are used to indicate the interval relationship between a pitch in the set of pitches and the standard pitch; determining the set of matching parameters based on a predetermined pitch weighting coefficient, the set of note difference parameters and the set of interval distance parameters, wherein the pitch weighting coefficient is used to indicate the degree of influence of the note difference parameters and the interval distance parameters on the matching parameters respectively.
[0198] As an optional solution, the above-mentioned device is used to determine the above-mentioned set of matching parameters by comparing each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch in the following manner: by comparing each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine a first set of matching parameters, wherein the above-mentioned first set of matching parameters represents the degree of matching between the first singing audio corresponding to the i-th frame and the above-mentioned standard singing audio when the target music is played to the i-th frame, i is a positive integer; using the above-mentioned first set of matching parameters and a second set of matching parameters to determine the above-mentioned set of matching parameters, wherein the above-mentioned second set of matching parameters represents the degree of matching between the first singing audio in the above-mentioned set of singing audio corresponding to the in-th frame and the above-mentioned standard singing audio when the target music is played to the in-th frame, n≤i, and n is a positive integer.
[0199] As an optional solution, the above-mentioned device is used to determine the above-mentioned set of matching parameters by comparing each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch in the following manner: when the target music is played to the j-th frame, a comparison operation is performed on each pitch in the above-mentioned set of pitches with the above-mentioned standard pitch to determine a set of initial matching parameters corresponding to the j-th frame, wherein the above-mentioned set of initial matching parameters represents the initial matching degree between the first singing audio and the standard singing audio corresponding to the j-th frame when the target music is played to the j-th frame, j>1, and j is a positive integer; the above-mentioned set of initial matching parameters and a set of historical matching parameters are used to determine the above-mentioned set of matching parameters, wherein the above-mentioned historical matching parameters correspond to the jm-th frame of the target music, and the above-mentioned set of matching parameters corresponds to the j-th frame of the target music, m<j, and m is a positive integer.
[0200] As an optional solution, the above-mentioned device is used to determine the set of matching parameters using the set of initial matching parameters and the set of historical matching parameters in the following manner: obtaining a first smoothing coefficient, and determining a set of second short-term matching parameters using the set of initial matching parameters, the first smoothing coefficient, and a set of first short-term matching parameters, wherein the set of historical matching parameters includes the set of first short-term matching parameters, the first short-term matching parameters correspond to the jm-th frame of the target music, the second short-term matching parameters correspond to the j-th frame of the target music, and the first smoothing coefficient is used to indicate the degree of influence of the first short-term matching parameters and the initial matching parameters on the second short-term matching parameters; obtaining a second smoothing coefficient, and determining a set of second short-term matching parameters using the set of initial matching parameters, the second smoothing coefficient, and a set of first long-term matching parameters. A set of second long-term matching parameters is determined by periodic matching parameters, wherein the set of historical matching parameters includes the set of first long-term matching parameters, the first long-term matching parameters correspond to the jm-th frame of the target music, the second long-term matching parameters correspond to the j-th frame of the target music, and the second smoothing coefficient is used to indicate the degree of influence of the first long-term matching parameters and the initial matching parameters on the second long-term matching parameters, wherein the first smoothing coefficient is less than the second smoothing coefficient; a third smoothing coefficient is obtained, and the set of matching parameters is determined by the set of second short-term matching parameters, the set of second long-term matching parameters, and the third smoothing coefficient, wherein the third smoothing coefficient is used to indicate the degree of influence of the second short-term matching parameters and the second long-term matching parameters on the matching parameters.
[0201] As an optional solution, the above-mentioned device is used to determine the set of pitches corresponding to the first set of singing audio and the standard pitch corresponding to the standard singing audio in the following manner: inputting the first set of singing audio and the standard singing audio into a preset autocorrelation function model respectively to determine a set of fundamental frequencies and a standard fundamental frequency, wherein the preset autocorrelation function is used to calculate the similarity between the signal and its own delayed version to determine a set of fundamental frequency periods and a standard fundamental frequency period, and determining the set of fundamental frequencies and the standard fundamental frequency based on the set of fundamental frequency periods, the standard fundamental frequency period and a preset sampling rate; performing a conversion operation on the set of fundamental frequencies and the standard fundamental frequency to determine the set of pitches and the standard pitch.
[0202] As an optional solution, the above-mentioned device is used to assign a first set of weight coefficients to the first set of singing audio based on the above-mentioned set of matching parameters, and to perform a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio: determining whether the above-mentioned set of matching parameters meets a preset matching condition; in response to the second matching parameter in the above-mentioned set of matching parameters meeting the preset matching condition, assigning the first set of weight coefficients to the second matching parameter according to the value of the second matching parameter, and performing a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio; in response to the third matching parameter in the above-mentioned set of matching parameters not meeting the preset matching condition, deleting the third matching parameter from the above-mentioned set of matching parameters, and assigning the first set of weight coefficients to the remaining fourth matching parameter in the above-mentioned set of matching parameters according to the value of the fourth matching parameter, and performing a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio.
[0203] As an optional solution, the above-mentioned device is used to assign a first set of weight coefficients to the first set of singing audio based on the above-mentioned set of matching parameters in the following manner, and to perform a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio: determining the value of each matching parameter in the above-mentioned set of matching parameters and a plurality of pre-set value intervals, wherein different value intervals in the plurality of value intervals correspond to different weight coefficients; assigning the first set of weight coefficients to the first set of singing audio according to the distribution of each matching parameter in the above-mentioned set of matching parameters in the plurality of value intervals, and performing a mixing operation on the first set of singing audio using the first set of weight coefficients to determine the target singing audio.
[0204] As an optional solution, the above-mentioned device is further configured to: when playing a first segment of the target music, determine the set of matching parameters corresponding to each frame in the first segment, wherein the first segment represents a segment pre-marked for the target music; assign a first set of weight coefficients to the first group of vocal audio using the set of matching parameters; perform a mixing operation on the first group of vocal audio using the first set of weight coefficients to determine a first target vocal audio; when playing a second segment of the target music on the target application, directly perform a mixing operation on the first group of vocal audio using the first set of weight coefficients to determine a second target vocal audio, wherein the first segment is different from the second segment.
[0205] As an optional solution, the above-mentioned device is used to acquire the first set of singing audio in the following manner: acquiring the first set of singing audio during the playback of the target music on the target application, wherein the first set of singing audio corresponds to a first set of accounts, and the first set of singing audio includes audio input by each account in the first set of accounts for the target music; or acquiring the first set of singing audio during the playback of the target music on the target application, wherein the first set of singing audio corresponds to a target account, and the first set of singing audio includes audio input by the target account for the target music at different times.
[0206] As an optional solution, the above-mentioned device is used to assign a first set of weight coefficients to the first set of singing audio based on the aforementioned set of matching parameters in at least one of the following ways: performing a mixing operation on the first set of singing audio using the first set of parameters; determining the target singing audio; and, when playing similar music on the target application, obtaining a second set of singing audio corresponding to the first set of accounts, wherein the second set of singing audio includes audio input by each account in the first set of accounts for the aforementioned similar music, and the aforementioned similar music refers to music marked as being of the same type as the target music; assigning the first set of weight coefficients to the second set of singing audio based on the aforementioned set of matching parameters; and performing a mixing operation on the first set of parameters using the first set of parameters. The weighted coefficients are used to perform a mixing operation on the second group of singing audio to determine the third target singing audio. When the target music is played on the target application, the third group of singing audio corresponding to the second group of accounts is obtained. The third group of singing audio includes the audio input by each account in the second group of accounts for the target music. When the first account in the first group of accounts and the second account in the second group of accounts have the same tag, the matching parameters determined for the first account are used as the matching parameters for the second account to determine the second group of weighted coefficients corresponding to the second group of accounts. The second group of weighted coefficients are used to perform a mixing operation on the third group of singing audio to determine the fourth target singing audio.
[0207] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0208] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0209] According to one aspect of this application, a computer program product is provided, the computer program product comprising a computer program.
[0210] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0211] Figure 12 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.
[0212] It should be noted that, Figure 12 The computer system 1200 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0213] like Figure 12 As shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1202 or programs loaded from storage section 1208 into random access memory (RAM). The RAM 1203 also stores various programs and data required for system operation. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output interface 1205 (I / O interface) is also connected to the bus 1204.
[0214] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.
[0215] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit 1201, it performs various functions defined in the system of this application.
[0216] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, it performs various functions provided in the embodiments of this application.
[0217] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described audio processing method is also provided. This electronic device may be... Figure 1 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps of any of the above method embodiments through the computer program.
[0218] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0219] Optionally, in this embodiment, the processor may be configured to execute the methods in the embodiments of this application via a computer program.
[0220] Alternatively, as those skilled in the art will understand, Figure 13 The structure shown is for illustrative purposes only. Figure 13 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 13 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 13 The different configurations shown.
[0221] The memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio processing method and apparatus in this embodiment. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, thereby implementing the aforementioned audio processing method. The memory 1302 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include memory remotely located relative to the processor 1304, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1302 may be used, but is not limited to, to store information such as singing audio. As an example, such as... Figure 13 As shown, the memory 1302 may include, but is not limited to, the playback module 1102, acquisition module 1104, determination module 1106, and execution module 1108 of the audio processing device. Furthermore, it may include, but is not limited to, other module units of the audio processing device, which will not be described in detail in this example.
[0222] Optionally, the transmission device 1306 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1306 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1306 is a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0223] In addition, the aforementioned electronic device also includes: a display 1308 for displaying the chorus interface of the aforementioned target application to play the target singing audio; and a connection bus 1310 for connecting the various module components in the aforementioned electronic device.
[0224] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0225] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of an electronic device reads computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the electronic device to perform the audio processing method provided in various alternative implementations of the above-described audio processing aspect.
[0226] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store methods for performing the embodiments of this application.
[0227] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0228] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0229] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of this application.
[0230] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0231] In the several embodiments provided in this application, it should be understood that the disclosed application can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0232] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0233] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0234] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method of processing audio, characterized by, The method comprises: obtaining a first set of singing audios, wherein the first set of singing audios comprises a plurality of audios for a target music input; determining a set of matching parameters based on the first set of singing audios and a corresponding standard singing audio, wherein the standard singing audio is a singing audio preset for the target music, and a first matching parameter in the set of matching parameters is used to indicate a matching degree between a first singing audio in the first set of singing audios and the standard singing audio; assigning a first set of weight coefficients to the first set of singing audios based on the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a target singing audio.
2. The method of claim 1, wherein, The method comprises: determining a set of standard pitches corresponding to the standard singing audio and a set of pitches corresponding to the first set of singing audios, wherein the set of pitches and the standard pitch correspond to a same time point of the target music, and the first matching parameter is determined by the set of pitches, a first pitch corresponding to the first singing audio, and the standard pitch; respectively comparing each pitch in the set of pitches with the standard pitch to determine the set of matching parameters.
3. The method of claim 2, wherein, The method comprises: respectively performing a pitch conversion operation on the set of pitches and the standard pitch to determine a set of pitch difference parameters and a set of interval distance parameters, wherein the pitch difference parameter is used to indicate a pitch difference between one pitch in the set of pitches and the standard pitch, and the interval distance parameter is used to indicate an interval relationship between one pitch in the set of pitches and the standard pitch; determining the set of matching parameters based on a preset pitch weight coefficient, the set of pitch difference parameters, and the set of interval distance parameters, wherein the pitch weight coefficient is used to indicate an influence degree of the pitch difference parameter and the interval distance parameter on the matching parameter.
4. The method of claim 2, wherein, The method comprises: respectively comparing each pitch in the set of pitches with the standard pitch to determine a set of first matching parameters, wherein the set of first matching parameters represents a matching degree between the first singing audio corresponding to the i-th frame and the standard singing audio when the target music plays to the i-th frame, i is a positive integer; determining the set of matching parameters using the set of first matching parameters and a set of second matching parameters, wherein the set of second matching parameters represents a matching degree between the first singing audio in the first set of singing audios corresponding to the i-n-th frame and the standard singing audio when the target music plays to the i-n-th frame, n≤i, and n is a positive integer.
5. The method of claim 2, wherein, The method comprises: In a case where the target music plays to the jth frame, each pitch in the set of pitches is compared with the standard pitch to determine a set of initial matching parameters corresponding to the jth frame, wherein the set of initial matching parameters represents an initial matching degree between the first singing audio corresponding to the jth frame and the standard singing audio when the target music plays to the jth frame, j>1, j is a positive integer; The set of initial matching parameters and a set of historical matching parameters are used to determine the set of matching parameters, wherein the historical matching parameters correspond to the j-mth frame of the target music, the set of matching parameters correspond to the jth frame of the target music, m 6. The method of claim 5, wherein, The set of initial matching parameters and a set of historical matching parameters are used to determine the set of matching parameters, including: A first smoothing coefficient is obtained, and a set of second short-term matching parameters is determined using the set of initial matching parameters, the first smoothing coefficient and a set of first short-term matching parameters, wherein the set of historical matching parameters includes the set of first short-term matching parameters, the first short-term matching parameters correspond to the j-mth frame of the target music, the second short-term matching parameters correspond to the jth frame of the target music, and the first smoothing coefficient is used to indicate the influence degree of the first short-term matching parameters and the initial matching parameters on the second short-term matching parameters respectively; A second smoothing coefficient is obtained, and a set of second long-term matching parameters is determined using the set of initial matching parameters, the second smoothing coefficient and a set of first long-term matching parameters, wherein the set of historical matching parameters includes the set of first long-term matching parameters, the first long-term matching parameters correspond to the j-mth frame of the target music, the second long-term matching parameters correspond to the jth frame of the target music, the second smoothing coefficient is used to indicate the influence degree of the first long-term matching parameters and the initial matching parameters on the second long-term matching parameters respectively, and the first smoothing coefficient is less than the second smoothing coefficient; A third smoothing coefficient is obtained, and the set of matching parameters is determined using the set of second short-term matching parameters, the set of second long-term matching parameters and the third smoothing coefficient, wherein the third smoothing coefficient is used to indicate the influence degree of the second short-term matching parameters and the second long-term matching parameters on the matching parameters respectively.
7. The method of claim 2, wherein, The set of initial matching parameters and a set of historical matching parameters are used to determine the set of matching parameters, including: The first set of singing audios and the standard singing audio are respectively input into a preset autocorrelation function model to determine a set of fundamental frequencies and a standard fundamental frequency, wherein the preset autocorrelation function is used to calculate the similarity between a signal and its delayed version to determine a set of pitch periods and a standard pitch period, and the set of fundamental frequencies and the standard fundamental frequency are determined based on the set of pitch periods, the standard pitch period and a preset sampling rate; The set of fundamental frequencies and the standard fundamental frequency are transformed to determine the set of pitches and the standard pitch.
8. The method of claim 1, wherein, The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph.
9. The method of claim 1, wherein, The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; 10. The method of claim 1, wherein, In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph. The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; 11. The method of claim 1, wherein, In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph. The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph. The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph. The method further comprises: In a case where a first paragraph of the target music is played, determining the set of matching parameters corresponding to each frame in the first paragraph, wherein the first paragraph represents a paragraph pre-labeled for the target music; assigning a first set of weight coefficients to the first set of singing audios using the set of matching parameters, performing a mixing operation on the first set of singing audios using the first set of weight coefficients, and determining a first target singing audio; In a case where a second paragraph of the target music is played on the target application, directly performing a mixing operation on the first set of singing audios using the first set of weight coefficients to determine a second target singing audio, wherein the first paragraph is different from the second paragraph.
12. The method of claim 11, wherein, After the first set of singing audios is assigned the first set of weight coefficients based on the set of matching parameters, the mixing operation is performed on the first set of singing audios using the first set of weight coefficients, and the target singing audio is determined, the method further includes at least one of the following: In a case where the target application plays the same type of music, a second set of singing audios corresponding to the first set of accounts is obtained, where the second set of singing audios includes audios input by each account in the first set of accounts respectively for the same type of music, and the same type of music represents music that is marked as having the same type as the target music; a third target singing audio is determined by assigning the first set of weight coefficients to the second set of singing audios based on the set of matching parameters and performing the mixing operation on the second set of singing audios using the first set of weight coefficients; In a case where the target application plays the target music, a third set of singing audios corresponding to a second set of accounts is obtained, where the third set of singing audios includes audios input by each account in the second set of accounts respectively for the target music; in a case where a first account in the first set of accounts and a second account in the second set of accounts have the same label, a second set of weight coefficients corresponding to the second set of accounts is determined using the matching parameters determined for the first account as the matching parameters of the second account, and a fourth target singing audio is determined by performing the mixing operation on the third set of singing audios using the second set of weight coefficients.
13. An apparatus for processing audio, characterized by comprise: an obtaining module configured to obtain a first set of singing audios, where the first set of singing audios includes a plurality of audios input for target music; a determining module configured to determine a set of matching parameters based on the first set of singing audios and a corresponding standard singing audio, where the standard singing audio is a singing audio set in advance for the target music, and a first matching parameter in the set of matching parameters is used to indicate a matching degree between a first singing audio in the first set of singing audios and the standard singing audio; an executing module configured to assign a first set of weight coefficients to the first set of singing audios based on the set of matching parameters, perform a mixing operation on the first set of singing audios using the first set of weight coefficients, and determine a target singing audio.
14. A computer readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, where the computer program can be run by an electronic device to execute the method described in any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the method described in any one of claims 1 to 12.
16. An electronic device comprising a memory and a processor, characterized in that The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 12 by using the computer program. The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 12 by using the computer program.