Audio processing method, device, equipment and storage medium
By using gain processing and loudness value analysis in the song element extraction model to determine the target gain coefficient, the problem of poor audio quality in the existing technology is solved, and an audio processing effect with a loudness value closer to the actual one is achieved.
Patent Information
- Application Number
- CN202111196144.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-10-14
AI Technical Summary
In the prior art, the audio quality of vocal audio or accompaniment audio extracted by machine learning models is poor, the loudness value is different from the actual loudness value, and the loudness value of the vocal audio frame changes unevenly.
By inputting multiple audio frames of the target song into the trained song element extraction model, using different gain coefficients to perform gain processing on the initial audio frame, determining the loudness value of the difference audio frame, and selecting the target gain coefficient closest to the actual loudness value for processing, the target audio frame is obtained.
The audio quality of vocal audio and accompaniment audio has been improved, making their loudness values closer to actual values and enhancing the effect of audio processing.
Smart Images

Figure CN113963707B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of technology, people can perform certain data processing on song audio containing vocals and accompaniment to separate the vocals and accompaniment, and obtain the vocal audio and accompaniment audio corresponding to the song audio. Some music applications provide users with richer and more diverse music entertainment methods by setting the audio of the two elements, vocal audio and accompaniment audio, separated from the song audio. For example, for karaoke applications, users can choose between original singing mode and accompaniment mode when singing karaoke. The original singing mode plays the song audio containing vocal audio and accompaniment audio for the user, while the accompaniment mode plays only the accompaniment audio for the user. For another example, for listening to music applications, users can choose to play only vocal audio or only accompaniment audio through operations, and so on.
[0003] The traditional method of separating vocal audio or accompaniment audio from a song is to use two different machine learning models to extract vocal audio and accompaniment audio respectively.
[0004] However, the loudness value of the vocal audio or accompaniment audio extracted by the machine learning model is somewhat different from the actual loudness value of the vocals or accompaniment in the song audio, and the loudness value changes of each vocal audio frame in the vocal audio are different, and the same is true for each accompaniment audio frame in the accompaniment audio, resulting in poor audio quality of the extracted vocal audio or accompaniment audio. Summary of the Invention
[0005] An embodiment of the present application provides an audio processing method that can solve the technical problem in the prior art of poor audio quality of vocal audio or accompaniment audio extracted by a song element extraction model.
[0006] In a first aspect, an audio processing method is provided, the method comprising:
[0007] Inputting multiple song audio frames of a target song into a trained song element extraction model to obtain initial audio frames of first-category elements corresponding to the song audio frames output by the song element extraction model, wherein the first-category elements are vocals or accompaniment;
[0008] Performing gain processing on the initial audio frames using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients;
[0009] Determine the difference audio frames between the song audio frame and each gain-processed initial audio frame, and determine the loudness value of the difference audio frame corresponding to each gain coefficient;
[0010] Based on the loudness value of the difference audio frame corresponding to each gain coefficient, determining a target gain coefficient corresponding to the actual loudness value of the first-category element in the song audio frame from among the different gain coefficients, and determining the gain-processed initial audio frame corresponding to the target gain coefficient as the target audio frame of the first-category element corresponding to the song audio frame;
[0011] The target audio frames of the first category elements corresponding to the song audio frames of each frame are combined into an audio segment of the first category elements corresponding to the target song.
[0012] In a possible implementation, the different gain coefficients are a plurality of gain coefficients distributed with equal intervals within a preset value range.
[0013] In a possible implementation, determining the loudness value of the difference audio frame corresponding to each gain coefficient includes:
[0014] For the difference audio frame corresponding to each gain coefficient, a root mean square of the loudness value of each sampling point in the difference audio frame is determined as the loudness value of the difference audio frame.
[0015] In a possible implementation, determining, among the different gain coefficients, a target gain coefficient corresponding to an actual loudness value of the first-category element in the song audio frame based on the loudness value of the difference audio frame corresponding to each gain coefficient includes:
[0016] The gain coefficient corresponding to the smallest loudness value among the loudness values of the difference audio frames corresponding to the gain coefficients is determined as the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame.
[0017] In a possible implementation, after determining the initial audio frame after gain processing corresponding to the target gain coefficient as the target audio frame of the first-category element corresponding to the song audio frame, the method further includes:
[0018] Determining the difference audio frame corresponding to the target gain coefficient as the target audio frame of the second type of element corresponding to the song audio frame, wherein the second type of element is vocals or accompaniment, and the second type of element is different from the first type of element;
[0019] The target audio frames of the second category elements corresponding to the song audio frames of each frame are combined into an audio segment of the second category elements corresponding to the target song.
[0020] In a possible implementation, the method further includes:
[0021] For each song audio frame, determining a target adjustment coefficient corresponding to the song audio frame based on a time interval between the song audio frame and a start time point of the target song, wherein the target adjustment coefficient of the song audio frame is positively correlated or negatively correlated with the time interval;
[0022] Using the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the song audio frame, gain processing is performed on the initial audio frame of the first-category element corresponding to the song audio frame to obtain an adjusted audio frame of the first-category element corresponding to the song audio frame;
[0023] The difference audio frames between the multiple song audio frames and the corresponding adjusted audio frames of the first category elements are respectively determined to form the adjusted audio clip corresponding to the target song.
[0024] In a second aspect, an audio processing method is provided, the method comprising:
[0025] Displaying a loudness adjustment interface corresponding to the target song, wherein a vocal loudness adjustment control and an accompaniment loudness adjustment control are provided in the loudness adjustment interface;
[0026] Obtaining a target vocal adjustment coefficient input through the vocal loudness adjustment control and a target accompaniment adjustment coefficient input through the accompaniment loudness adjustment control;
[0027] Sending an adjustment request to a server, wherein the adjustment request carries identification information of the target song, the target vocal adjustment coefficient, and the target accompaniment adjustment coefficient;
[0028] Receive the adjusted audio corresponding to the target song sent by the server.
[0029] According to a third aspect, an audio processing method is provided, the method comprising:
[0030] Receive an adjustment request sent by a target terminal, wherein the adjustment request carries identification information of a target song, a target vocal adjustment coefficient, and a target accompaniment adjustment coefficient;
[0031] Based on the identification information of the target song, obtaining multiple song audio frames of the target song;
[0032] Determine the vocal audio frames and the corresponding accompaniment audio frames corresponding to the multiple song audio frames;
[0033] Performing gain processing on the vocal audio frame corresponding to each song audio frame using the target vocal adjustment coefficient to obtain a gain-processed vocal audio frame corresponding to each song audio frame;
[0034] Performing gain processing on the accompaniment audio frame corresponding to each song audio frame using the target accompaniment adjustment coefficient to obtain a gain-processed accompaniment audio frame corresponding to each song audio frame;
[0035] The gain-processed vocal audio frame and the gain-processed accompaniment audio frame corresponding to each song audio frame are combined into the adjusted audio corresponding to the target song;
[0036] The adjusted audio corresponding to the target song is sent to the target terminal.
[0037] According to a fourth aspect, an audio processing device is provided, the device comprising:
[0038] A first determination module is configured to input multiple song audio frames of a target song into a trained song element extraction model to obtain initial audio frames of first-category elements corresponding to the song audio frames output by the song element extraction model, wherein the first-category elements are vocals or accompaniment;
[0039] a gain module, configured to perform gain processing on the initial audio frames using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients;
[0040] A second determining module is used to respectively determine a difference audio frame between the song audio frame and each gain-processed initial audio frame, and determine a loudness value of the difference audio frame corresponding to each gain coefficient;
[0041] a third determining module, configured to determine, based on the loudness value of the difference audio frame corresponding to each gain coefficient, a target gain coefficient corresponding to the actual loudness value of the first-category element in the song audio frame from among the different gain coefficients, and determine the gain-processed initial audio frame corresponding to the target gain coefficient as the target audio frame of the first-category element corresponding to the song audio frame;
[0042] The composition module is used to compose the target audio frames of the first type of elements corresponding to the song audio frames of each frame into an audio clip of the first type of elements corresponding to the target song.
[0043] In a possible implementation, the different gain coefficients are a plurality of gain coefficients distributed with equal intervals within a preset value range.
[0044] In a possible implementation, the second determining module is configured to:
[0045] For the difference audio frame corresponding to each gain coefficient, a root mean square of the loudness value of each sampling point in the difference audio frame is determined as the loudness value of the difference audio frame.
[0046] In a possible implementation, the third determining module is configured to:
[0047] The gain coefficient corresponding to the smallest loudness value among the loudness values of the difference audio frames corresponding to the gain coefficients is determined as the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame.
[0048] In a possible implementation, the apparatus further includes a fourth determining module, configured to:
[0049] Determining the difference audio frame corresponding to the target gain coefficient as the target audio frame of the second type of element corresponding to the song audio frame, wherein the second type of element is vocals or accompaniment, and the second type of element is different from the first type of element;
[0050] The target audio frames of the second category elements corresponding to the song audio frames of each frame are combined into an audio segment of the second category elements corresponding to the target song.
[0051] In a possible implementation, the apparatus further includes a fifth determining module, configured to:
[0052] For each song audio frame, determining a target adjustment coefficient corresponding to the song audio frame based on a time interval between the song audio frame and a start time point of the target song, wherein the target adjustment coefficient of the song audio frame is positively correlated or negatively correlated with the time interval;
[0053] Using the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the song audio frame, gain processing is performed on the initial audio frame of the first-category element corresponding to the song audio frame to obtain an adjusted audio frame of the first-category element corresponding to the song audio frame;
[0054] The difference audio frames between the multiple song audio frames and the corresponding adjusted audio frames of the first category elements are respectively determined to form the adjusted audio clip corresponding to the target song.
[0055] The beneficial effect of the technical solution provided by the embodiment of the present application is as follows: the embodiment of the present application can first extract the initial audio frame of the first-category element corresponding to the song audio frame based on the song element extraction model, and then determine the target gain coefficient corresponding to the actual loudness value of the first-category element in the song audio frame based on the loudness value of the difference audio frame after gain processing using different gain coefficients. Based on the target gain coefficient, a target audio frame of the first-category element with a loudness value closer to the actual loudness value is obtained, thereby obtaining an audio clip of the first-category element with better audio quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0058] Figure 2 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0059] Figure 3 Schematic diagram of a changing relationship between a gain coefficient and a loudness value of a difference audio frame provided in an embodiment of the present application;
[0060] Figure 4 This is a flow chart of a method for determining and adjusting an audio segment provided by an embodiment of the present application;
[0061] Figure 5 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0062] Figure 6 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0063] Figure 7 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0064] Figure 8 is a structural diagram of an audio processing device provided in an embodiment of the present application;
[0065] Figure 9 This is a structural block diagram of a terminal provided in an embodiment of the present application;
[0066] Figure 10 This is a structural block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0068] An embodiment of the present application provides an audio processing method that can be implemented by a computer device. The computer device can be a device for extracting vocal audio or accompaniment audio from the target song audio, or a device that needs to extract vocal audio or accompaniment audio from the target song audio. For example, it can be a background server of a music application, or it can be a terminal for a user who can listen to music. The computer device can be a terminal or server, etc. The terminal can be a desktop computer, laptop computer, tablet computer, mobile phone, etc. The computer device may include a processor, memory, communication components, etc.
[0069] The processor can be a central processing unit (CPU), which can be used to determine the initial audio frame of the first category element corresponding to the song audio frame based on the song element extraction model, determine the loudness value of the difference audio frame corresponding to each gain coefficient, determine the target gain coefficient, determine the target audio frame of the first category element corresponding to the song audio frame, and so on.
[0070] The memory may be various volatile memories or non-volatile memories, such as a solid state disk (SSD), a dynamic random access memory (DRAM), etc. The memory may be used for data storage, for example, data storage of song audio frames, data storage of song element extraction models, data storage of initial audio frames of first-category elements corresponding to determined song audio frames, data storage of difference audio frames corresponding to different determined gain coefficients, data storage of loudness values of difference audio frames corresponding to each gain coefficient, data storage of target audio frames of first-category elements corresponding to determined song audio frames, and the like.
[0071] The communication component may be a wired network connector, a wireless fidelity (WiFi) module, a Bluetooth module, a cellular network communication module, etc. The communication component may be used to transmit data with other devices. For example, the communication component may be used to send the determined audio clip of the first category element or the audio clip of the second category element to a specified device, etc.
[0072] Figure 1 and Figure 2 This is a flow chart of an audio processing method provided by an embodiment of the present application. Figure 1 and Figure 2 , the embodiment includes:
[0073] 101. Input multiple song audio frames of the target song into the trained song element extraction model to obtain initial audio frames of the first category of elements corresponding to the song audio frames output by the song element extraction model.
[0074] Among them, the first type of element is vocals or accompaniment.
[0075] In implementation, when a song audio needs to be processed, for the sake of ease of description, the song audio to be processed can be referred to as a target song, and the audio frames contained in the target song can be referred to as song audio frames. The target song contains vocal audio and accompaniment audio, and the song element extraction model can be used to extract the vocal audio or accompaniment audio in the target song to obtain the vocal audio or accompaniment audio corresponding to the target song. The first-category element in the embodiment of the present application can be a vocal or an accompaniment, and the audio of the first-category element can be a vocal audio or an accompaniment audio, and the corresponding song element extraction model can include a vocal extraction model and an accompaniment extraction model. When the first-category element is a vocal, the vocal extraction model can be used to extract the vocal audio in the target song to obtain the vocal audio corresponding to the target song. When the first-category element is an accompaniment, the accompaniment extraction model can be used to extract the accompaniment audio in the target song to obtain the accompaniment audio corresponding to the target song.
[0076] Optionally, there are multiple methods for extracting the audio of the first-category element in the target song using the song element extraction model, and the following is one of them:
[0077] At least one song audio frame of the target song is input into the trained song element extraction model to obtain an initial audio frame of the first category element corresponding to each song audio frame in the at least one song audio frame.
[0078] During implementation, the song audio of the target song can be input into the trained song element extraction model. The song element extraction model processes the song audio frame and outputs the processed audio frame. In order to distinguish it from other audio frames, the audio frame output by the song element extraction model can be called the initial audio frame. The audio composed of the initial audio frame is the initial audio of the first type of element corresponding to the target song audio.
[0079] Alternatively, a preset number of song audio frames can be input into the song element extraction model each time, that is, the song audio can be divided into multiple input data according to the preset number of audio frames. If the number of song audio frames in the last input data is less than the preset number of audio frames, silent audio frames can be used to complete it. Then, the input data can be input into the trained song element extraction model respectively, and corresponding multiple output data can be obtained. Each output data is the initial audio frame of the first category element corresponding to the song audio frame in the input data. The initial audio frame in the output data corresponding to the silent audio frame in the last input data is deleted, and the remaining output data is the initial audio frame of the first category element corresponding to each song audio frame in the song audio.
[0080] 102. Perform gain processing on the initial audio frames using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients.
[0081] In implementation, for each initial audio frame, different gain coefficients are used to perform gain processing on the amplitude of each time domain sampling point in the initial audio frame, and the amplitude is multiplied by the gain coefficient to obtain the gain-processed initial audio frame corresponding to the different gain coefficients.
[0082] 103. Determine the difference audio frames between the song audio frame and each gain-processed initial audio frame, and determine the loudness value of the difference audio frame corresponding to each gain coefficient.
[0083] In practice, the amplitude of each time-domain sampling point in the song audio frame is subtracted from the amplitude of the time-domain sampling point corresponding to the initial audio frame after gain processing to obtain the difference audio frame between the song audio frame and the initial audio frame after gain processing. For each different gain coefficient, the difference audio frame corresponding to each initial audio frame can be obtained in the above manner. If the song audio frame can be represented by Y, the initial audio frame can be represented by X, the difference audio frame can be represented by R, and the gain coefficient can be represented by a, then the formula for the difference audio frame can be expressed as: R = Y - aX.
[0084] Because multiple different gain coefficients are used to gain process the initial audio frame, multiple initial audio frames gain-processed according to the gain coefficients can be obtained, thereby obtaining multiple difference audio frames corresponding to the different gain coefficients. Optionally, these multiple different gain coefficients can be set to multiple increasing or decreasing values, so that the effects of different gain coefficients on the loudness values of the difference audio frames can be compared. The different gain coefficients can be set to multiple gain coefficients with equal differences distributed within a preset value range.
[0085] Since the loudness value of the initial audio frame may deviate from the actual loudness value of the first type of element in the song audio frame, a preset numerical range can be pre-set for the value of the gain coefficient, and the value of the gain coefficient can be a plurality of gain coefficients with equal difference distribution within the preset numerical range. In an embodiment of the present application, the gain coefficient can be a plurality of values uniformly distributed with a difference of 0.01 within the preset numerical range of [0,2], that is, the value of the gain coefficient can be 0, 0.01, 0.02, 0.03...1.98, 1.99, 2. Of course, the preset numerical range of the gain coefficient can also be other ranges, and the embodiment of the present application does not limit this.
[0086] Optionally, after obtaining a plurality of difference audio frames corresponding to different gain coefficients, the loudness value of each difference audio frame can be calculated. There are multiple methods for calculating the loudness value of the difference audio frame, and the following is one of them:
[0087] For the difference audio frame corresponding to each gain coefficient, the root mean square of the loudness value of each sampling point in the difference audio frame is determined as the loudness value of the difference audio frame.
[0088] In implementation, gain processing is performed on the initial audio frame using multiple different gain coefficients to obtain multiple gain-processed initial audio frames corresponding to the different gain coefficients, and thus multiple difference audio frames corresponding to the different gain coefficients can be obtained. For each difference audio frame corresponding to the gain coefficient, the loudness value of each sampling point in the difference audio frame can be obtained, and then the root mean square of the loudness values of the multiple sampling points of the difference audio frame can be calculated to serve as the loudness value of the difference audio frame. The corresponding formula can be as follows:
[0089] RMS(R)=10lg(sum(R 2 / n))
[0090] sum(R 2 / n)=(R1 2 +R2 2 +……+R n 2 ) / n
[0091] Wherein, R is the difference audio frame, RMS(R) is the loudness value of the difference audio frame, and n is the sampling point number.
[0092] Optionally, other calculation methods may be selected to represent the loudness value of the difference audio frame, which is not limited in this embodiment of the present application.
[0093] 104. Based on the loudness value of the difference audio frame corresponding to each gain coefficient, determine the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame among different gain coefficients, and determine the initial audio frame after gain processing corresponding to the target gain coefficient as the target audio frame of the first category element corresponding to the song audio frame.
[0094] In implementation, after obtaining the loudness values of the difference audio frames corresponding to multiple different gain coefficients, a loudness value that is equal to or closest to the actual loudness value of the first-category element in the song audio frame can be determined from these multiple loudness values. The gain coefficient corresponding to the loudness value that is equal to or closest to the actual loudness value is determined as the target gain coefficient.
[0095] Optionally, the process of determining the target gain coefficient in the embodiment of the present application may be as follows:
[0096] The gain coefficient corresponding to the smallest loudness value among the loudness values of the difference audio frames corresponding to the gain coefficients is determined as the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame.
[0097] In implementations, different gain coefficients correspond to different loudness values of difference audio frames. When the loudness value of the difference audio frames corresponding to the multiple determined different gain coefficients reaches a minimum value, it means that after the initial audio frame is gain-processed using the gain coefficient corresponding to the minimum loudness value, the loudness value of the initial audio frame obtained after gain processing is closest to the actual loudness value of the first-category element in the song audio frame.
[0098] When the selected gain coefficient is smaller than the target gain coefficient, the loudness value of the difference audio frame corresponding to the gain coefficient has not yet reached the minimum value, indicating that there are still more sounds of the first category elements in the difference audio frame.
[0099] When the selected gain coefficient is greater than the target gain coefficient, the loudness value of the initial audio frame after gain processing using the gain coefficient will be greater than the actual loudness value of the first-category element in the song audio frame. Then, the difference audio frame obtained by subtracting the initial audio frame after gain processing from the song audio frame is the song audio frame minus the audio frame of the first-category element contained in it, plus the audio frame of the first-category element with reverse amplitude. At this time, the sound of the first-category element will still exist in the difference audio frame, and as the gain coefficient becomes larger and larger, the loudness value of the first-category element in the difference audio frame will gradually increase.
[0100] Therefore, when the selected gain coefficient is not equal to the target gain coefficient, the difference between the loudness value of the initial audio frame after gain processing using the gain coefficient and the actual loudness value of the first-category element in the song audio frame will be more obvious, indicating that the difference audio frame still contains the sound of the first-category element. Therefore, the gain coefficient corresponding to the minimum loudness value is determined as the target gain coefficient. Then, the loudness value of the initial audio frame after gain processing using the target gain coefficient is equal to or closest to the actual loudness value of the first-category element in the song audio frame.
[0101] For example, Figure 3 As shown, the horizontal axis is a plurality of gain coefficients with equal difference distribution in the preset value range [0,2], and the vertical axis is the loudness value of the difference audio frame corresponding to the gain coefficients with different values. Figure 3 It can be seen that when the gain coefficient is 1.52, the loudness value of the difference audio frame is the smallest, which is -18.63 dB. Therefore, the target gain coefficient can be determined to be 1.52.
[0102] From the above, it can be seen that the loudness value of the initial audio frame after gain processing using the target gain coefficient is closest to the actual loudness value of the first-category element in the song audio frame. Therefore, the initial audio frame after gain processing corresponding to the target gain coefficient can be determined as the target audio frame of the first-category element corresponding to the song audio frame.
[0103] When the first type of element is human voice, the above steps can be used to obtain the human voice audio frame corresponding to the song audio frame; when the first type of element is accompaniment, the above steps can be used to obtain the accompaniment audio frame corresponding to the song audio frame.
[0104] Optionally, after determining the target audio frame of the first-category element corresponding to the song audio frame, the difference audio frame corresponding to the target audio frame of the first-category element may also be determined. The corresponding processing may be as follows:
[0105] The difference audio frame corresponding to the target gain coefficient is determined as the target audio frame of the second type of element corresponding to the song audio frame, wherein the second type of element is vocals or accompaniment, and the second type of element is different from the first type of element.
[0106] In implementation, after determining the target gain coefficient, the difference audio frame corresponding to the target gain coefficient can also be determined as the audio frame of the second type of element corresponding to the song audio frame, that is, the difference audio frame between the song audio frame and the target audio frame of the first type of element is determined as the audio frame of the second type of element. When the first type of element is a human voice, the target audio frame of the first type of element is the human voice audio frame in the song audio frame, and the audio frame of the second type of element is the accompaniment audio frame in the song audio frame; when the first type of element is an accompaniment, the target audio frame of the first type of element is the accompaniment audio frame in the song audio frame, and the audio frame of the second type of element is the human voice audio frame in the song audio frame.
[0107] 105. Target audio frames of the first category elements corresponding to the song audio frames are combined into an audio clip of the first category elements corresponding to the target song.
[0108] In implementation, after obtaining the target audio frame of the first category element corresponding to each song audio frame in multiple song audio frames, the target audio frames of the first category element corresponding to these multiple song audio frames can be arranged and combined according to the order of the song audio frames in the target song to form an audio clip of the first category element corresponding to the target song.
[0109] Similarly, the target audio frames of the second type of elements corresponding to the obtained multi-frame song audio frames can also be combined according to the arrangement order of the multi-frame song audio frames in the target song to form an audio clip of the second type of elements corresponding to the target song.
[0110] The vocal audio and accompaniment audio obtained in the above manner can be used separately. Alternatively, the target gain coefficient can be adjusted to obtain a dynamic song effect in which the vocals gradually appear, or a dynamic song effect in which the vocals gradually disappear, or even a dynamic song effect in which the vocals appear and disappear intermittently. Similarly, a dynamic song effect in which the accompaniment gradually appears, or a dynamic song effect in which the accompaniment gradually disappears, or even a dynamic song effect in which the accompaniment appears and disappears intermittently.
[0111] like Figure 4 As shown, the corresponding processing flow can be as follows:
[0112] 401. For each song audio frame, determine a target adjustment coefficient corresponding to the song audio frame based on a time interval between the song audio frame and a start time point of a target song.
[0113] The target adjustment coefficient of the song audio frame is positively or negatively correlated with the time interval.
[0114] In implementation, after determining the target gain coefficient corresponding to each initial audio frame, the target adjustment coefficient corresponding to each song audio frame can also be determined. The target adjustment coefficient can be a value in the range of [0, 1] and is positively correlated or negatively correlated with the time interval. That is, the target adjustment coefficient corresponding to multiple consecutive song audio frames can have a larger value the farther away from the start time point of the target song, or, the smaller the value the farther away from the start time point of the target song.
[0115] 402. Use the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the song audio frame to perform gain processing on the initial audio frame of the first category element corresponding to the song audio frame to obtain the adjusted audio frame of the first category element corresponding to the song audio frame.
[0116] In implementation, for a song audio frame, the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the initial audio frame of the first-category element corresponding to the song audio frame are used to perform gain processing on the initial audio frame of the first-category element corresponding to the song audio frame, thereby obtaining an adjusted audio frame of the first-category element corresponding to the song audio frame. By performing gain processing on each song audio frame in the above manner, an adjusted audio frame of the first-category element corresponding to each song audio frame in a plurality of song audio frames can be obtained.
[0117] 403. Determine difference audio frames between the plurality of song audio frames and the corresponding adjusted audio frames of the first category elements to form an adjusted audio clip corresponding to the target song.
[0118] In practice, the amplitude of each time-domain sampling point in the song audio frame is subtracted from the amplitude of the time-domain sampling point of the adjusted audio frame corresponding to the song audio frame to obtain a difference audio frame between the song audio frame and the corresponding adjusted audio frame. By processing each song audio frame in the above manner, a difference audio frame for each adjusted audio frame in multiple song audio frames can be obtained.
[0119] By arranging the difference audio frames corresponding to the multiple adjusted audio frames according to the arrangement order of the song audio frames in the target song, an audio frequency band can be obtained, which is the adjusted audio segment corresponding to the target song.
[0120] When the target adjustment coefficient corresponding to the song audio frame is positively correlated with the time interval, that is, when the target adjustment coefficients corresponding to multiple song audio frames become larger the farther away from the starting time point, the loudness of the adjusted audio frames of the first-category elements corresponding to the multiple song audio frames will become larger and larger. However, since the value range of the target adjustment coefficient is [0, 1], the maximum value of the loudness of the adjusted audio frame will not be greater than the loudness of the target audio frame of the first-category element corresponding to the adjusted audio frame.
[0121] In the adjusted audio frequency band determined based on the song audio frame and the adjusted audio frame, the loudness of the first-category element in the resulting adjusted audio frequency band will decrease as the loudness of the adjusted audio frame of the subtracted first-category element increases. For example, if the first-category element is a human voice, and the target adjustment coefficient corresponding to the song audio frame is positively correlated with the time interval, then in the resulting adjusted audio clip, the human voice will become increasingly quieter over time, creating a dynamic song effect where the human voice gradually fades away.
[0122] Similarly, when the target adjustment coefficient for a song audio frame is negatively correlated with the time interval, the loudness of the resulting adjusted audio frame for the first-category element will decrease, and the loudness of the first-category element in the resulting adjusted audio clip will increase. For example, if the first-category element is a human voice, and the target adjustment coefficient for a song audio frame is negatively correlated with the time interval, then in the resulting adjusted audio clip, the human voice will become louder over time, creating a dynamic song effect where the human voice gradually emerges.
[0123] You can also first divide the target song into multiple audio frequency bands, process the odd-numbered audio segments so that the vocals gradually disappear, and process the even-numbered audio segments so that the vocals gradually appear, or process the even-numbered audio segments so that the vocals gradually disappear, and process the odd-numbered audio segments so that the vocals gradually appear, and you will get a dynamic song effect where the vocals appear and disappear.
[0124] When the first type of element is accompaniment, the processing method is the same as above, and a dynamic song effect in which the accompaniment gradually appears, or a dynamic song effect in which the accompaniment gradually disappears, or a dynamic song effect in which the accompaniment appears and disappears can also be obtained. The processing method will not be repeated here.
[0125] The present application also provides an audio processing method. Figure 5 , the corresponding processing flow is as follows:
[0126] 501. Display a loudness adjustment interface corresponding to the target song, wherein a vocal loudness adjustment control and an accompaniment loudness adjustment control are provided in the loudness adjustment interface.
[0127] In practice, a music application is installed on the user's target terminal. The user can open the music application and enter the loudness adjustment interface (also known as the volume adjustment interface) of the target song. The loudness adjustment interface includes a vocal loudness adjustment control and an accompaniment loudness adjustment control. The loudness adjustment control can be a sliding control with a minimum selectable value of 0 and a maximum selectable value of 1. The user can control the vocal loudness of the target song by sliding the corresponding vocal loudness adjustment control button, and control the accompaniment loudness of the target song by sliding the corresponding accompaniment loudness adjustment control button.
[0128] 502. Obtain a target vocal adjustment coefficient input through a vocal loudness adjustment control and a target accompaniment adjustment coefficient input through an accompaniment loudness adjustment control.
[0129] In implementation, after the user adjusts the vocal loudness adjustment control or the accompaniment loudness adjustment control, the target terminal may obtain the target vocal adjustment coefficient input through the vocal loudness adjustment control and the target accompaniment adjustment coefficient input through the accompaniment loudness adjustment control.
[0130] 503. Send a request for adjustment to the server.
[0131] The adjustment request carries the identification information of the target song, the target vocal adjustment coefficient, and the target accompaniment adjustment coefficient.
[0132] 504. Receive the adjusted audio corresponding to the target song sent by the server.
[0133] During implementation, after sending an adjustment request to the server, the server will process multiple song audio frames in the target song based on the target vocal adjustment coefficient and the target accompaniment adjustment coefficient, and send the adjusted audio corresponding to the target song obtained after processing back to the target terminal. The target terminal receives the adjusted audio and plays it.
[0134] The present application also provides an audio processing method. Figure 6 , the corresponding processing flow is as follows:
[0135] 601. Receive an adjustment request sent by a target terminal.
[0136] The adjustment request carries the identification information of the target song, the target vocal adjustment coefficient, and the target accompaniment adjustment coefficient.
[0137] 602. Based on the identification information of the target song, obtain multiple song audio frames of the target song.
[0138] 603. Determine the vocal audio frames and the corresponding accompaniment audio frames corresponding to the multiple song audio frames.
[0139] In implementation, steps 101-105 can be used to obtain the target audio frames of vocals and the target audio frames of accompaniment corresponding to all the song audio frames of the target song, that is, the vocal audio frames and the accompaniment audio frames corresponding to the song audio frames.
[0140] 604. Perform gain processing on the vocal audio frame corresponding to each song audio frame using the target vocal adjustment coefficient to obtain a gain-processed vocal audio frame corresponding to each song audio frame.
[0141] 605. Perform gain processing on the accompaniment audio frame corresponding to each song audio frame using the target accompaniment adjustment coefficient to obtain the accompaniment audio frame corresponding to each song audio frame after gain processing. In implementation, there is no order between step 604 and step 605.
[0142] 606. The gain-processed vocal audio frames and the gain-processed accompaniment audio frames corresponding to each song audio frame are combined to form an adjusted audio corresponding to the target song.
[0143] 607. Send the adjusted audio corresponding to the target song to the target terminal.
[0144] The present application also provides an audio processing method. Figure 7 , the corresponding processing flow is as follows:
[0145] 701. The target terminal displays a loudness adjustment interface corresponding to the target song, where a vocal loudness adjustment control and an accompaniment loudness adjustment control are provided.
[0146] 702. The target terminal obtains a target vocal adjustment coefficient input through a vocal loudness adjustment control and a target accompaniment adjustment coefficient input through an accompaniment loudness adjustment control.
[0147] 703. The target terminal sends an adjustment request to the server.
[0148] The adjustment request carries the identification information of the target song, the target vocal adjustment coefficient, and the target accompaniment adjustment coefficient.
[0149] 704. The server receives the adjustment request sent by the target terminal.
[0150] 705. The server obtains multiple song audio frames of the target song based on the identification information of the target song.
[0151] 706. The server determines the vocal audio frames and the corresponding accompaniment audio frames corresponding to the multiple song audio frames.
[0152] 707. The server performs gain processing on the vocal audio frame corresponding to each song audio frame using the target vocal adjustment coefficient to obtain the gain-processed vocal audio frame corresponding to each song audio frame.
[0153] 708. The server performs gain processing on the accompaniment audio frame corresponding to each song audio frame using the target accompaniment adjustment coefficient to obtain the gain-processed accompaniment audio frame corresponding to each song audio frame.
[0154] 709. The server combines the gain-processed vocal audio frames and the gain-processed accompaniment audio frames corresponding to each song audio frame into an adjusted audio corresponding to the target song.
[0155] 710. The server sends the adjusted audio corresponding to the target song to the target terminal.
[0156] 711. The target terminal receives the adjusted audio corresponding to the target song sent by the server.
[0157] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0158] The solution mentioned in the embodiment of the present application can first extract the initial audio frame of the first category element corresponding to the song audio frame based on the song element extraction model, and then determine the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame according to the loudness value of the difference audio frame after gain processing using different gain coefficients. Based on the target gain coefficient, the target audio frame of the first category element with a loudness value closer to the actual loudness value is obtained, thereby obtaining an audio clip of the first category element with better audio quality.
[0159] The present application embodiment provides an audio processing device, which may be the computer device in the above embodiment, such as Figure 8 As shown, the device includes:
[0160] A first determination module 810 is configured to input multiple song audio frames of a target song into a trained song element extraction model to obtain initial audio frames of first-category elements corresponding to the song audio frames output by the song element extraction model, wherein the first-category elements are vocals or accompaniment;
[0161] A gain module 820 is configured to perform gain processing on the initial audio frame using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients;
[0162] The second determining module 830 is used to respectively determine the difference audio frames between the song audio frame and each gain-processed initial audio frame, and determine the loudness value of the difference audio frame corresponding to each gain coefficient;
[0163] A third determining module 840 is configured to determine, based on the loudness value of the difference audio frame corresponding to each gain coefficient, a target gain coefficient corresponding to the actual loudness value of the first-category element in the song audio frame from among the different gain coefficients, and determine the gain-processed initial audio frame corresponding to the target gain coefficient as the target audio frame of the first-category element corresponding to the song audio frame;
[0164] The composition module 850 is used to combine the target audio frames of the first category elements corresponding to the song audio frames of each frame into an audio segment of the first category elements corresponding to the target song.
[0165] In a possible implementation, the different gain coefficients are a plurality of gain coefficients distributed with equal intervals within a preset value range.
[0166] In a possible implementation, the second determining module 830 is configured to:
[0167] For the difference audio frame corresponding to each gain coefficient, a root mean square of the loudness value of each sampling point in the difference audio frame is determined as the loudness value of the difference audio frame.
[0168] In a possible implementation, the third determining module 840 is configured to:
[0169] The gain coefficient corresponding to the smallest loudness value among the loudness values of the difference audio frames corresponding to the gain coefficients is determined as the target gain coefficient corresponding to the actual loudness value of the first category element in the song audio frame.
[0170] In a possible implementation, the apparatus further includes a fourth determining module, configured to:
[0171] Determining the difference audio frame corresponding to the target gain coefficient as the target audio frame of the second type of element corresponding to the song audio frame, wherein the second type of element is vocals or accompaniment, and the second type of element is different from the first type of element;
[0172] The target audio frames of the second category elements corresponding to the song audio frames of each frame are combined into an audio segment of the second category elements corresponding to the target song.
[0173] In a possible implementation, the apparatus further includes a fifth determining module, configured to:
[0174] For each song audio frame, determining a target adjustment coefficient corresponding to the song audio frame based on a time interval between the song audio frame and a start time point of the target song, wherein the target adjustment coefficient of the song audio frame is positively correlated or negatively correlated with the time interval;
[0175] Using the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the song audio frame, gain processing is performed on the initial audio frame of the first-category element corresponding to the song audio frame to obtain an adjusted audio frame of the first-category element corresponding to the song audio frame;
[0176] The difference audio frames between the multiple song audio frames and the corresponding adjusted audio frames of the first category elements are respectively determined to form the adjusted audio clip corresponding to the target song.
[0177] It should be noted that the audio processing device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate audio processing. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device provided in the above embodiment and the audio processing method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0178] Figure 9 The following is a block diagram of a terminal 900 according to an exemplary embodiment of the present application. The terminal may be the computer device described in the aforementioned embodiments. The terminal 900 may be a smartphone, a tablet computer, an MP3 player (moving picture experts group audio layer III), an MP4 player (moving picture experts group audio layer IV), a laptop computer, or a desktop computer. The terminal 900 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.
[0179] Typically, the terminal 900 includes a processor 901 and a memory 902 .
[0180] The processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (digital signal processing), FPGA (field-programmable gate array), or PLA (programmable logic array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (graphics processing unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI (artificial intelligence) processor, which is used to process computing operations related to machine learning.
[0181] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one instruction, which is executed by the processor 901 to implement the audio processing method provided in the method embodiment of the present application.
[0182] In some embodiments, terminal 900 may optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, memory 902, and peripheral device interface 903 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 903 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 904, a display screen 905, a camera 906, an audio circuit 907, a positioning component 908, and a power supply 909.
[0183] The peripheral device interface 903 can be used to connect at least one I / O (input / output)-related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0184] The radio frequency circuit 904 is used to receive and transmit RF (radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 904 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi networks. In some embodiments, the radio frequency circuit 904 may also include circuits related to NFC (near field communication), which is not limited in this application.
[0185] Display screen 905 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. If display screen 905 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 905. These touch signals can be input as control signals to processor 901 for processing. Display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 905, located on the front panel of terminal 900. In other embodiments, there can be at least two display screens 905, located on different surfaces of terminal 900 or in a foldable design. In still other embodiments, display screen 905 can be a flexible display, located on a curved or foldable surface of terminal 900. Furthermore, display screen 905 can be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 905 can be made of materials such as LCD (liquid crystal display) and OLED (organic light-emitting diode).
[0186] The camera assembly 906 is used to capture images or videos. Optionally, the camera assembly 906 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (virtual reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0187] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 901 for processing, or input into the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.
[0188] Positioning component 908 is used to locate the current geographic location of terminal 900 to implement navigation or LBS (location-based service). Positioning component 908 can be based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Greninja system, or the European Union's Galileo system.
[0189] Power supply 909 is used to power various components in terminal 900. Power supply 909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 909 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0190] In some embodiments, the terminal 900 further includes one or more sensors 910 , including but not limited to: an acceleration sensor 911 , a gyroscope sensor 912 , a pressure sensor 913 , a fingerprint sensor 914 , an optical sensor 915 , and a proximity sensor 916 .
[0191] The accelerometer 911 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal 900. For example, the accelerometer 911 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 901 can control the display screen 905 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 911. The accelerometer 911 can also be used to collect game or user motion data.
[0192] The gyroscope sensor 912 can detect the orientation and rotation angle of the terminal 900. It can work with the accelerometer 911 to collect the user's 3D movements on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0193] The pressure sensor 913 can be set on the side frame of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is set on the side frame of the terminal 900, it can detect the user's grip signal of the terminal 900, and the processor 901 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 913. When the pressure sensor 913 is set on the lower layer of the display screen 905, the processor 901 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0194] The fingerprint sensor 914 is used to collect the user's fingerprint. The processor 901 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 901 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 914 can be set on the front, back, or side of the terminal 900. When a physical button or manufacturer logo is set on the terminal 900, the fingerprint sensor 914 can be integrated with the physical button or manufacturer logo.
[0195] The optical sensor 915 is used to detect ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 based on the ambient light intensity detected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera assembly 906 based on the ambient light intensity detected by the optical sensor 915.
[0196] Proximity sensor 916, also known as a distance sensor, is typically located on the front panel of terminal 900. Proximity sensor 916 is used to detect the distance between the user and the front of terminal 900. In one embodiment, when proximity sensor 916 detects that the distance between the user and the front of terminal 900 is gradually decreasing, processor 901 controls display screen 905 to switch from the screen-on state to the screen-off state. When proximity sensor 916 detects that the distance between the user and the front of terminal 900 is gradually increasing, processor 901 controls display screen 905 to switch from the screen-off state to the screen-on state.
[0197] Those skilled in the art will understand that Figure 9 The structure shown in the figure does not constitute a limitation on the terminal 900, and the terminal 900 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0198] Figure 10 1 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1000 may vary significantly due to different configurations or performance, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one instruction, which is loaded and executed by the processor 1001 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which are not described in detail here.
[0199] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The instructions can be executed by a processor in a terminal to perform the audio processing method in the above embodiment. The computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be a ROM (read-only memory), RAM (random access memory), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0200] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0201] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An audio processing method, characterized in that: The method comprises: Receive an adjustment request sent by a target terminal, wherein the adjustment request carries identification information of a target song, a target vocal adjustment coefficient, and a target accompaniment adjustment coefficient; Based on the identification information of the target song, obtaining multiple song audio frames of the target song; Inputting multiple song audio frames of the target song into a trained vocal extraction model to obtain initial audio frames corresponding to the vocals in the song audio frames output by the vocal extraction model; Performing gain processing on the initial audio frames using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients; Subtracting the amplitude of the time domain sampling point corresponding to the initial audio frame after the gain processing from the amplitude of each time domain sampling point in the song audio frame to obtain a difference audio frame, and determining the loudness value of the difference audio frame corresponding to each gain coefficient; Determining the gain coefficient corresponding to the difference audio frame with the smallest loudness value as the target gain coefficient corresponding to the actual loudness value of the human voice in the song audio frame, determining the initial audio frame after gain processing corresponding to the target gain coefficient as the human voice audio frame, and determining the difference audio frame corresponding to the target gain coefficient as the corresponding accompaniment audio frame; Performing gain processing on the vocal audio frame using the target vocal adjustment coefficient to obtain a gain-processed vocal audio frame, and performing gain processing on the accompaniment audio frame using the target accompaniment adjustment coefficient to obtain a gain-processed accompaniment audio frame; The gain-processed vocal audio frame and the gain-processed accompaniment audio frame form an adjusted audio corresponding to the target song; The adjusted audio corresponding to the target song is sent to the target terminal.
2. The method according to claim 1, characterized in that The different gain coefficients are multiple gain coefficients distributed with equal intervals within a preset value range.
3. The method according to claim 1, characterized in that Determining the loudness value of the difference audio frame corresponding to each gain coefficient includes: For the difference audio frame corresponding to each gain coefficient, a root mean square of the loudness value of each sampling point in the difference audio frame is determined as the loudness value of the difference audio frame.
4. The method according to claim 1, wherein The method further comprises: For each song audio frame, determining a target adjustment coefficient corresponding to the song audio frame based on a time interval between the song audio frame and a start time point of the target song, wherein the target adjustment coefficient of the song audio frame is positively correlated or negatively correlated with the time interval; Using the target adjustment coefficient corresponding to the song audio frame and the target gain coefficient corresponding to the song audio frame, gain processing is performed on the initial audio frame of the human voice corresponding to the song audio frame to obtain an adjusted audio frame of the human voice corresponding to the song audio frame; The difference audio frames between the multiple song audio frames and the adjusted audio frames corresponding to the human voice are respectively determined to form the adjusted audio clip corresponding to the target song.
5. An audio processing device, characterized in that: The device is used to: Receive an adjustment request sent by a target terminal, wherein the adjustment request carries identification information of a target song, a target vocal adjustment coefficient, and a target accompaniment adjustment coefficient; Based on the identification information of the target song, obtaining multiple song audio frames of the target song; The device includes a first determination module for inputting multiple song audio frames of the target song into a trained vocal extraction model to obtain initial audio frames corresponding to the vocals in the song audio frames output by the vocal extraction model; The device includes a gain module, configured to perform gain processing on the initial audio frames using different gain coefficients to obtain gain-processed initial audio frames corresponding to the different gain coefficients; The device includes a second determining module for subtracting the amplitude of the time domain sampling point corresponding to the initial audio frame after the gain processing from the amplitude of each time domain sampling point in the song audio frame to obtain a difference audio frame, and determining the loudness value of the difference audio frame corresponding to each gain coefficient; The device includes a third determining module, configured to determine a gain coefficient corresponding to a difference audio frame having a minimum loudness value as a target gain coefficient corresponding to an actual loudness value of a human voice in the song audio frame, determine an initial audio frame after gain processing corresponding to the target gain coefficient as a human voice audio frame, and determine a difference audio frame corresponding to the target gain coefficient as a corresponding accompaniment audio frame; The device includes a component module for performing gain processing on the vocal audio frame using the target vocal adjustment coefficient to obtain a gain-processed vocal audio frame, performing gain processing on the accompaniment audio frame using the target accompaniment adjustment coefficient to obtain a gain-processed accompaniment audio frame; and combining the gain-processed vocal audio frame and the gain-processed accompaniment audio frame to form an adjusted audio corresponding to the target song; The device is further configured to send the adjusted audio corresponding to the target song to the target terminal.
6. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the operation performed by the audio processing method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the operation performed by the audio processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for dynamically adjusting volume
CN102098606A
Lyric alignment method and related product
CN111210850A
Method, apparatus and computer program for calculating and adjusting the perceived loudness of an audio signal
EP1835487A2