Audio processing method and device, electronic equipment and storage medium
By extracting audio features and correlation features from multi-channel audio and using a prediction model to determine the target bitrate, the problem of low bitrate accuracy in multi-channel audio coding is solved, and the coding effect is improved.
Patent Information
- Application Number
- CN202310522004.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-05-10
AI Technical Summary
In existing technologies, the accuracy of the encoding bitrate is low during multi-channel audio encoding, resulting in poor encoding processing performance.
The audio features of a specified channel in a multi-channel audio dataset are extracted, and the correlation features between the audio features and single-channel audio features are obtained. A prediction model is used to obtain multiple coding bitrates and their corresponding sound quality, and the target bitrate of the multi-channel audio is determined.
The accuracy of the multi-channel audio encoding bitrate was improved, thereby enhancing the encoding effect.
Smart Images

Figure CN116564319B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to an audio processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] At present, as an audio signal processing technology, audio coding has been widely applied. Through audio coding, audio signals can be compressed to reduce the transmission bandwidth required for audio transmission and the storage space required for audio as much as possible. Among them, the audio quality is closely related to the audio coding mode adopted.
[0003] In the prior art, the coding rate of the audio is often directly determined based on the overall content of the audio, and subsequent encoding processing of the audio is performed based on the coding rate. In this way, when processing multi-channel audio, the accuracy of the determined coding rate is low, which further leads to poor subsequent encoding processing effect. SUMMARY
[0004] The present disclosure provides an audio processing method and device, electronic equipment and storage medium to at least solve the problem of low accuracy of the coding rate in the related art, which further leads to poor subsequent encoding processing effect. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, an audio processing method is provided, comprising:
[0006] extracting an audio feature of a specified channel audio in a multi-channel audio to be encoded; the specified channel audio is obtained based on a single-channel audio included in the multi-channel audio;
[0007] extracting a correlation feature between at least part of the single-channel audio included in the multi-channel audio;
[0008] inputting the audio feature and the correlation feature into a preset prediction model to obtain a plurality of coding rates output by the prediction model for the multi-channel audio and an audio quality corresponding to each of the plurality of coding rates;
[0009] determining a target coding rate of the multi-channel audio based on the plurality of coding rates and the audio quality corresponding to each of the plurality of coding rates.
[0010] Optionally, before the extracting the audio feature of the specified channel audio in the multi-channel audio to be encoded, the method further comprises:
[0011] selecting N single-channel audios from the single-channel audios included in the multi-channel audio as the specified channel audio;
[0012] and / or,
[0013] generating N audio groups based on the mono-channel audios included in the multi-channel audio; the N is a positive integer, and there is at least one audio group including at least two mono-channel audios in the N audio groups;
[0014] For any of the audio groups, generating one of the specified channel audios based on the mono-channel audios included in the audio group.
[0015] Optionally, the extracting of the correlation features between the at least part of the mono-channel audios included in the multi-channel audio comprises:
[0016] determining a channel group based on the multiple audio channels corresponding to the multi-channel audio; at least two mono-channels are included in one of the channel groups;
[0017] For any of the channel groups, extracting correlation features between the mono-channel audios corresponding to the at least two mono-channels included in the channel group.
[0018] Optionally, the extracting of the correlation features between the mono-channel audios corresponding to the at least two mono-channels included in the channel group comprises:
[0019] obtaining inter-channel correlation parameters between specified audio frames in the mono-channel audios corresponding to the at least two mono-channels; the inter-channel correlation parameters comprise a parameter for representing a degree of correlation between the mono-channel audios corresponding to the at least two mono-channels;
[0020] For any of the inter-channel correlation parameters, determining a feature corresponding to the inter-channel correlation parameter according to the inter-channel correlation parameters between the specified audio frames;
[0021] generating the correlation features based on the features corresponding to all the inter-channel correlation parameters.
[0022] Optionally, the determining of the channel group based on the multiple audio channels corresponding to the multi-channel audio comprises:
[0023] in a case where a total number of the audio channels corresponding to the multi-channel audio is equal to 2, determining two audio channels corresponding to the multi-channel audio as one channel group;
[0024] in a case where the total number of the audio channels corresponding to the multi-channel audio is greater than 2, dividing at least two audio channels corresponding to audio contents with a similarity meeting a preset requirement into a same channel group.
[0025] Optionally, the prediction model is obtained by training in the following manner:
[0026] obtaining multiple sample encoding bit rates of a sample multi-channel audio and audio qualities corresponding to the multiple sample encoding bit rates respectively;
[0027] input the audio features and the correlation features of the sample multi-channel audio into a preset prediction model, and acquire a plurality of encoding bit rates output by the prediction model for the multi-channel audio and audio qualities corresponding to the plurality of encoding bit rates respectively;
[0028] adjust model parameters of the prediction model based on the plurality of sample encoding bit rates, the audio qualities corresponding to the plurality of sample encoding bit rates respectively, the plurality of encoding bit rates output by the prediction model and the audio qualities corresponding to the plurality of encoding bit rates respectively;
[0029] determine the prediction model as the trained prediction model in a case where the prediction model converges.
[0030] Optionally, the inputting the audio features and the correlation features into the preset prediction model comprises:
[0031] splicing the audio features and the correlation features to obtain spliced features;
[0032] inputting the spliced features into the prediction model.
[0033] According to a second aspect of the embodiments of the present disclosure, an audio processing apparatus applied to a terminal is provided, comprising:
[0034] a first extraction module configured to perform extraction of audio features of specified channel audio in multi-channel audio to be encoded, the specified channel audio being obtained based on single-channel audio included in the multi-channel audio;
[0035] a second extraction module configured to perform extraction of correlation features between at least part of single-channel audio included in the multi-channel audio;
[0036] a first acquisition module configured to perform inputting of the audio features and the correlation features into a preset prediction model, and acquiring a plurality of encoding bit rates output by the prediction model for the multi-channel audio and audio qualities corresponding to the plurality of encoding bit rates respectively;
[0037] a first determination module configured to perform determination of a target bit rate of the multi-channel audio based on the plurality of encoding bit rates and the audio qualities corresponding to the plurality of encoding bit rates respectively.
[0038] Optionally, the apparatus further comprises:
[0039] a selection module configured to perform selection of N single-channel audios from single-channel audio included in the multi-channel audio as the specified channel audio;
[0040] and / or,
[0041] a first generating module configured to generate N audio groups based on the mono-channel audios included in the multi-channel audio, wherein N is a positive integer, and at least two mono-channel audios are included in at least one of the N audio groups;
[0042] a second generating module configured to generate the specified channel audio based on the mono-channel audio included in the audio group for any of the audio groups.
[0043] Optionally, the second extracting module is specifically configured to:
[0044] determine a channel group based on the multiple audio channels corresponding to the multi-channel audio, wherein at least two mono-channels are included in the channel group;
[0045] extract a correlation feature between the mono-channel audios corresponding to the at least two mono-channels included in the channel group for any of the channel groups.
[0046] Optionally, the second extracting module is specifically further configured to:
[0047] obtain an inter-channel correlation parameter between the specified audio frames in the mono-channel audios corresponding to the at least two mono-channels, wherein the inter-channel correlation parameter includes a parameter for representing a degree of correlation between the mono-channel audios corresponding to the at least two mono-channels;
[0048] determine a feature corresponding to the inter-channel correlation parameter according to the inter-channel correlation parameter between the specified audio frames for any of the inter-channel correlation parameters;
[0049] generate the correlation feature based on the features corresponding to all the inter-channel correlation parameters.
[0050] Optionally, the second extracting module is specifically further configured to:
[0051] in a case where a total number of audio channels corresponding to the multi-channel audio is equal to 2, determine two audio channels corresponding to the multi-channel audio as a channel group;
[0052] in a case where the total number of audio channels corresponding to the multi-channel audio is greater than 2, divide at least two audio channels corresponding to audio content with a similarity meeting a preset requirement into a same channel group.
[0053] Optionally, the prediction model is obtained by training the following modules:
[0054] a second obtaining module configured to obtain multiple sample encoding bit rates of a sample multi-channel audio and audio qualities corresponding to the multiple sample encoding bit rates respectively;
[0055] a third obtaining module configured to obtain, as input of the to-be-trained prediction model, audio features and correlation features of the sample multi-channel audio, and obtain a plurality of encoding bit rates output by the to-be-trained prediction model and audio qualities corresponding to the plurality of encoding bit rates respectively;
[0056] an adjusting module configured to adjust model parameters of the to-be-trained prediction model based on the plurality of sample encoding bit rates, the audio qualities corresponding to the plurality of sample encoding bit rates respectively, the plurality of encoding bit rates output by the to-be-trained prediction model, and the audio qualities corresponding to the plurality of encoding bit rates respectively;
[0057] a second determining module configured to determine the to-be-trained prediction model as the prediction model in a case where the to-be-trained prediction model converges.
[0058] Optionally, the first obtaining module is specifically configured to:
[0059] concatenate the audio features and the correlation features to obtain concatenated features;
[0060] input the concatenated features into the prediction model.
[0061] According to a third aspect of embodiments of the present disclosure, an electronic device is provided, including:
[0062] a processor;
[0063] a memory for storing instructions executable by the processor;
[0064] wherein the processor is configured to execute the instructions to implement the method according to any one of the first aspect.
[0065] According to a fourth aspect of embodiments of the present disclosure, a storage medium is provided, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device performs the method according to any one of the first aspect.
[0066] According to a fifth aspect of embodiments of the present disclosure, a computer program product is provided, the computer program product includes readable program instructions, when the readable program instructions are executed by a processor of an electronic device, the electronic device performs the method according to any one of the first aspect.
[0067] The embodiments of the present disclosure provide at least the following beneficial effects: In the embodiments of the present disclosure, an audio feature of a specified channel audio in the multi-channel audio to be encoded is extracted, and the specified channel audio is obtained based on a single-channel audio included in the multi-channel audio. A correlation feature between at least part of the single-channel audios included in the multi-channel audio is extracted. The audio feature and the correlation feature are input into a preset prediction model, and a plurality of encoding code rates and a plurality of audio qualities corresponding to the plurality of encoding code rates respectively, which are output by the prediction model for the multi-channel audio, are obtained. Based on the plurality of encoding code rates and the plurality of audio qualities corresponding to the plurality of encoding code rates respectively, a target code rate of the multi-channel audio is determined. In this way, compared with a manner of directly determining an encoding code rate based on audio content, in the embodiments of the present disclosure, when processing the multi-channel audio, the audio feature of the specified channel audio obtained based on the single-channel audio in the multi-channel audio and the correlation feature between at least part of the single-channel audios included in the multi-channel audio are extracted, and the target code rate of the multi-channel audio is determined based on the audio feature and the correlation feature. Since the channel correlation of the multi-channel audio will affect the required encoding code rate, the audio feature and the correlation feature can more comprehensively represent the multi-channel audio, so that the determined target code rate can be more suitable for the multi-channel audio to a certain extent, and thus the accuracy of the target code rate determined for the multi-channel audio is improved, thereby improving the subsequent encoding effect.
[0068] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0069] The accompanying drawings, which are incorporated into and form part of the specification, illustrate an embodiment consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not limit the present disclosure.
[0070] Figure 1 is a flowchart of an audio processing method according to an exemplary embodiment;
[0071] Figure 2 is a schematic diagram of an audio processing process according to an exemplary embodiment;
[0072] Figure 3 is a block diagram of an audio processing device according to an exemplary embodiment;
[0073] Figure 4 is a block diagram of a device for audio processing according to an exemplary embodiment;
[0074] Figure 5 is a block diagram of another device for audio processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0075] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings.
[0076] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0077] Figure 1 is a flowchart of an audio processing method according to an exemplary embodiment, as shown in Figure 1 may include the following steps:
[0078] Step 101, extracting an audio feature of a specified channel audio in the multi-channel audio to be encoded; the specified channel audio is obtained based on a single-channel audio included in the multi-channel audio.
[0079] In the embodiments of the present disclosure, the multi-channel audio to be encoded can be any multi-channel audio that needs to be encoded. The multi-channel audio can include at least two audio channels, and accordingly, based on the multi-channel audio, a single-channel audio corresponding to each audio channel included in the multi-channel audio can be extracted, i.e., the multi-channel audio can be decomposed into at least two single-channel audios. The multi-channel audio can also be referred to as multi-channel audio, and one audio channel is one channel, and a single-channel audio corresponding to one audio channel is a channel signal corresponding to one channel.
[0080] Further, the specified channel audio can be a single-channel audio itself, or can be mixed based on a single-channel audio, and the embodiments of the present disclosure do not limit this. Accordingly, for any specified channel audio, an audio feature of the specified channel audio can be extracted. In this way, the obtained audio feature can be characterized from the dimension of a single channel audio included in the multi-channel audio. The type of audio feature can be set according to actual needs, and the embodiments of the present disclosure do not limit this.
[0081] Step 102, extracting a correlation feature between at least part of the single-channel audios included in the multi-channel audio.
[0082] The inter-channel correlation of the single-channel audio included in the multi-channel audio is different, and the required encoding code rate often differs. For example, taking a two-channel as an example, when the inter-channel correlation of the two-channel audio is strong, the encoder often only needs to increase a small amount of code rate on the basis of single-channel encoding to retain the sound details of the two-channel. Conversely, when the inter-channel correlation of the two-channel audio is weak, the encoder needs to spend more code rate on the basis of single-channel encoding to retain the sound details. That is, the inter-channel correlation of the single-channel audio can be reflected in the code rate required for the multi-channel audio to a certain extent while ensuring the encoding quality. Therefore, in the embodiments of the present disclosure, the correlation features between at least part of the single-channel audio included in the multi-channel audio can be further extracted. In this way, the obtained correlation features can be characterized from the inter-channel dimension of the multiple single-channel audio included in the multi-channel audio. The correlation features can include features for characterizing the inter-channel correlation, and the specific types of the correlation features can be set according to actual needs, which are not limited in the embodiments of the present disclosure.
[0083] Step 103, input the audio features and the correlation features into a preset prediction model to obtain multiple encoding code rates output by the prediction model for the multi-channel audio and audio qualities corresponding to the multiple encoding code rates respectively.
[0084] The preset prediction model can be a pre-trained model, which can be a neural network model. The prediction model can output multiple encoding code rates for the multi-channel audio based on the input, and the audio qualities corresponding to each of the multiple encoding code rates. The number of the output encoding code rates can be pre-set, for example, the number of the output encoding code rates can be 10.
[0085] Step 104, determine a target code rate of the multi-channel audio based on the multiple encoding code rates and the audio qualities corresponding to the multiple encoding code rates respectively.
[0086] In the embodiments of the present disclosure, it can be determined whether the target sound quality exists in the sound qualities corresponding to the plurality of encoding code rates. If the target sound quality exists, the encoding code rate corresponding to the target sound quality can be directly selected from the plurality of encoding code rates as the target code rate. If the target sound quality does not exist, a sound quality-code rate curve can be constructed based on the plurality of encoding code rates and the sound qualities corresponding to the plurality of encoding code rates. Alternatively, the sound quality-code rate curve can be directly constructed. One encoding code rate and the sound quality corresponding to the encoding code rate can be taken as one data pair. The sound quality of the obtained audio after the multi-channel audio is encoded at the encoding code rate in the data pair can be embodied by the data pair. The sound quality-code rate curve can be obtained by curve fitting based on the plurality of data pairs. The sound quality-code rate curve can represent the sound quality of the obtained audio after the multi-channel audio is encoded at different encoding code rates. The target sound quality can be the sound quality required by the current encoding operation, and the target sound quality can be pre-set. The encoding code rate corresponding to the target sound quality can be searched from the sound quality-code rate curve, and the target code rate is obtained. The encoding code rate corresponding to the target sound quality is adaptively determined based on the plurality of encoding code rates and the sound qualities corresponding to the plurality of encoding code rates, and the encoding code rate is taken as the target code rate. In this way, it can be ensured that the finally determined encoding code rate can meet the requirement of the encoding operation on the sound quality.
[0087] Further, the multi-channel audio can be encoded based on the target code rate. For example, the target code rate can be set as the encoding code rate used by an encoder, and the multi-channel audio can be taken as the input of the encoder, and the multi-channel audio can be encoded by the encoder according to the target code rate. Correspondingly, the encoder can output the encoded multi-channel audio.
[0088] To sum up, the audio processing method provided in the embodiments of the present disclosure extracts the audio features of the specified channel audio in the multi-channel audio to be encoded, and the specified channel audio is obtained based on the mono-channel audio included in the multi-channel audio. The correlation features between at least part of the mono-channel audios included in the multi-channel audio are extracted. The audio features and the correlation features are input into a preset prediction model to obtain the multiple encoding code rates output by the prediction model for the multi-channel audio and the audio qualities corresponding to the multiple encoding code rates respectively. Based on the multiple encoding code rates and the audio qualities corresponding to the multiple encoding code rates respectively, the target code rate of the multi-channel audio is determined. In this way, compared with the way of directly determining the encoding code rate based on the audio content, in the embodiments of the present disclosure, when the multi-channel audio is processed, the audio features of the specified channel audio obtained based on the mono-channel audio in the multi-channel audio and the correlation features between at least part of the mono-channel audios included in the multi-channel audio are extracted, and the target code rate of the multi-channel audio is determined based on the audio features and the correlation features. Since the channel correlation of the multi-channel audio will affect the required encoding code rate, the audio features and the correlation features can more comprehensively represent the multi-channel audio, so that the determined target code rate can be more suitable for the multi-channel audio to some extent, and thus the accuracy of the target code rate determined for the multi-channel audio is improved, thereby improving the subsequent encoding effect.
[0089] Optionally, before the step of extracting the audio features of the specified channel audio in the multi-channel audio to be encoded, the specified channel audio can be determined through the following steps:
[0090] Step 201: selecting N mono-channel audios from the mono-channel audios included in the multi-channel audio as the specified channel audios.
[0091] And / or, step 202: generating N audio groups based on the mono-channel audios included in the multi-channel audio; the N is a positive integer, and there is an audio group including at least two mono-channel audios in the N audio groups.
[0092] Step 203: for any audio group, generating one specified channel audio based on the mono-channel audios included in the audio group.
[0093] In the embodiments of the present disclosure, the specific value of N can be set according to actual conditions, and N is not greater than the total number of single-channel audios included in the multi-channel audio. For example, the specific value of N can be set according to the computing power, and the size of N can be positively correlated with the computing power. For example, when the computing power is high, a larger N is set. When the computing power is low, a smaller N is set to avoid too many specified channel audios required to be processed, resulting in too low processing efficiency. At the same time, the specific value of N is adaptively set based on the computing power, which can ensure that the set N matches the computing power, avoid too many specified channel audios required to be processed, resulting in too low processing efficiency, while providing as many specified channel audios as possible for subsequent processing, thereby providing more dimensional audio features for subsequent determination of the target code rate, thereby ensuring the subsequent determination effect.
[0094] Further, in one implementation, one specified channel audio can correspond to one single-channel audio. Specifically, N single-channel audios can be directly selected from the single-channel audios included in the multi-channel audio as the specified channel audios. In this way, the determination efficiency of the specified channel audio can be ensured to some extent, thereby improving the overall processing efficiency. When selecting the N single-channel audios, the N single-channel audios can be randomly selected from the multiple single-channel audios included in the multi-channel audio. Alternatively, the N single-channel audios can be selected based on a preset rule, for example, the single-channel audios corresponding to the first N channels with the maximum energy.
[0095] In another implementation, there can be a specified channel audio corresponding to at least two single-channel audios. A plurality of single-channel audios included in the multi-channel audio can be selected and mixed to obtain a specified channel audio. Specifically, the single-channel audios included in the multi-channel audio are grouped to obtain N audio groups. For example, N can be set to be less than the total number of single-channel audios included in the multi-channel audio, and the multiple single-channel audios included in the multi-channel audio are divided into N audio groups to ensure that there is an audio group including at least two single-channel audios. Alternatively, only part of the single-channel audios included in the multi-channel audio can be grouped, which is not limited in the embodiments of the present disclosure. For example, the multi-channel audio includes single-channel audios: audio A, audio B, audio C, audio D, and audio E, audio A and audio B can be divided into one audio group, and audio D and audio E can be divided into one audio group.
[0096] In grouping, the single-channel audios can be randomly selected or selected based on preset rules, for example, selecting single-channel audios with similar audio content meeting preset requirements. The preset requirements can include that the similarity of the audio content is not less than a preset similarity threshold. The audio channels can be grouped in advance according to the characteristics of different audio channels. For example, for 5.1 multi-channel audio, the 5.1 multi-channel audio includes six audio channels: left channel, middle channel, right channel, left surround channel, right surround channel, and bass channel. Since the audio collected by the left channel, middle channel, and right channel often belongs to the same type, the left channel, middle channel, and right channel can be grouped as one audio channel. The left surround channel and the right surround channel are grouped as one audio channel, and the bass channel is grouped as one audio channel.
[0097] Correspondingly, the single-channel audios corresponding to the audio channels grouped in the same audio channel can be determined as single-channel audios with similar audio content meeting preset requirements. The left channel audio, middle channel audio, and right channel audio can be grouped as one audio group, the left surround channel audio and the right surround channel audio can be grouped as one audio group, and the bass channel audio can be grouped as one audio group. For any audio group, the single-channel audios included in the audio group can be mixed to obtain a specified channel audio. For example, the single-channel audios included in the audio group can be added and averaged to achieve mixing. After addition, it can also be detected whether there is reverse cancellation of phase, and if so, phase adjustment can be performed to avoid the problem of reverse cancellation.
[0098] It should be noted that in the embodiments of the present disclosure, the multi-channel audio can be first input into the channel extraction / mixing module, and the multi-channel audio is first separated based on the channel extraction / mixing module to obtain a plurality of single-channel audios included therein, and then the single-channel audios are extracted as specified channel audios. Alternatively, the plurality of single-channel audios are mixed to obtain a specified channel audio. Further, N can be less than the total number of single-channel audios included in the multi-channel audio, so that the data dimension required for subsequent processing can be reduced, thereby improving processing efficiency.
[0099] In another implementation manner, N single-channel audios can be selected from the single-channel audios included in the multi-channel audio as specified channel audios, and N audio groups can be generated based on the single-channel audios included in the multi-channel audio, and a specified channel audio can be generated based on the single-channel audios included in each audio group. That is, the final specified channel audio includes the directly selected single-channel audio and the channel audio generated based on the audio group.
[0100] Further, for any specified channel audio, a single-channel feature extraction operation can be performed on the specified channel audio, specifically, a specified type of feature of the specified channel audio can be extracted. The specified type of feature can include statistical mean and variance of Mel Frequency Cepstral Coefficients (MFCC) of multiple audio frames (e.g., all audio frames) in the specified channel audio, statistical mean and variance of sub-band energy ratio of multiple audio frames. Further, it can also include audio effective bandwidth and audio richness of an audio segment in the specified channel audio. The length of the audio segment can be set according to actual needs, for example, the audio segment can be an audio segment with a time length of 3 seconds. When extracting the sub-band energy ratio of the audio frame, a short-time Fourier transform can be performed on the specified channel audio first, and then the frequency spectrum data corresponding to the multiple audio frames in the specified channel audio is obtained. Then, the frequency spectrum can be divided into sub-bands, and the ratio of the energy of the sub-band to the total energy of the spectrum is calculated to obtain the sub-band energy ratio of the audio frame.
[0101] For any specified channel audio, multiple specified types of features of the specified channel audio can be spliced to ultimately obtain a D1-dimensional feature vector. Accordingly, for N specified channel audios, a feature with a dimension of D1*N can be ultimately obtained as the first audio feature. The specific value of D1 can be set according to actual needs, for example, D1 can be 40.
[0102] In the embodiments of the present disclosure, when extracting the audio features of the specified channel audio, N single-channel audios are directly selected as the specified channel audio. In this way, the processing amount of selecting the specified channel audio can be saved to some extent, and the processing efficiency is further improved. By dividing the single-channel audios included in the multi-channel audio into N audio groups and generating the specified channel audio based on the single-channel audios included in the audio groups, it can be ensured that the N specified channel audios ultimately obtained can represent all single-channel audios included in the multi-channel audio, and the comprehensiveness of the information provided for subsequent operations is further ensured.
[0103] Optionally, the step of extracting the correlation feature between the single-channel audios included in the multi-channel audio can specifically include:
[0104] In step 1021, a channel group is determined based on the multiple audio channels corresponding to the multi-channel audio, and at least two single channels are included in one channel group.
[0105] In this step, the number of single channels included in one channel group can be set according to actual needs, for example, two single channels can be included in one channel group. Assuming that the multi-channel audio includes M audio channels, the M audio channels can be combined two by two to obtain a plurality of channel groups. The plurality of audio channels corresponding to the multi-channel audio are a plurality of single channels corresponding to the multi-channel audio, wherein the plurality of single channels corresponding to the multi-channel audio represent the single channels included in the multi-channel audio. Then K channel groups are selected therefrom. For example, the combination can be randomly two by two, and K channel groups can be randomly selected. Wherein the K channel groups are not greater than the total number of combinations obtained by combining the plurality of audio channels included in the multi-channel audio two by two, and K is an integer not less than 1. The audio channels included in different channel groups can be different, that is, there are no two channel groups that are exactly the same. For example, if the channel group 1 includes a left channel and a middle channel, and the channel group 2 includes a left channel and a right channel, the channel group 1 and the channel group 2 are different, that is, the channel group 1 and the channel group 2 are partially the same. If the channel group 1 includes a left channel and a middle channel, and the channel group 2 includes a left surround channel and a right surround channel, the channel group 1 and the channel group 2 are two completely different channel groups. If the channel group 1 includes a left channel and a middle channel, and the channel group 2 includes a left channel and a middle channel, the channel group 1 and the channel group 2 are two completely same channel groups.
[0106] It should be noted that for one channel group, the single channels included in the channel group can be different or completely the same. For example, assuming that the multi-channel audio is pseudo dual-channel audio, that is, the left channel and the right channel included in the multi-channel audio are completely the same. Accordingly, for a channel group including a left channel and a right channel, the single channels included in the channel group are completely the same.
[0107] Step 1022, for any of the channel groups, extracting a correlation feature between the single channel audios corresponding to the at least two single channels included in the channel group.
[0108] For any channel group, a multi-channel feature extraction operation can be performed on the single channel audios corresponding to the at least two single channels included in the channel group. Specifically, the inter-channel correlation parameter between the single channel audios corresponding to the at least two single channels included in the channel group can be extracted, and the correlation feature is determined based on the inter-channel correlation parameter. Wherein the inter-channel correlation parameter can be set according to actual needs, and the embodiments of the present disclosure do not make any limitation.
[0109] In the embodiments of the present disclosure, a channel group is determined based on the plurality of audio channels corresponding to the multi-channel audio, and at least two single channels are included in one channel group. Then, for any channel group, a correlation feature between the single-channel audio corresponding to the at least two single channels included in the channel group is extracted. In this way, by grouping first and then extracting, the correlation feature between the single-channel audio included in the multi-channel audio can be conveniently extracted to some extent.
[0110] Optionally, the step of determining the channel group based on the plurality of audio channels corresponding to the multi-channel audio can specifically include:
[0111] In the step 1021a, in a case where the total number of audio channels corresponding to the multi-channel audio is equal to 2, the two audio channels corresponding to the multi-channel audio are determined as one channel group.
[0112] In the step 1021b, in a case where the total number of audio channels corresponding to the multi-channel audio is greater than 2, at least two audio channels corresponding to audio content with a similarity meeting a preset requirement are divided into the same channel group.
[0113] In the embodiments of the present disclosure, if the total number of audio channels corresponding to the multi-channel audio is equal to 2, that is, only two audio channels are included in the channel audio, the two audio channels can be directly selected as one channel group, and then the operation of selecting at least one channel group from the plurality of audio channels included in the multi-channel audio can be conveniently implemented. For example, for double-channel audio, the left channel and the right channel can be directly selected as one channel group. Further, if the total number of audio channels corresponding to the multi-channel audio is greater than 2, that is, the audio channels included in the channel audio can form a plurality of channel groups. Therefore, at least two audio channels corresponding to audio content with a similarity meeting a preset requirement can be divided into the same channel group. In this way, since the similarity of the audio content corresponding to the single channels in the same channel group meets the preset requirement, the reference of the correlation feature generated for the channel group can be improved to some extent.
[0114] The preset requirement can include that the similarity of the audio content is not less than a preset similarity threshold. The audio channels can be grouped in advance according to the characteristics of different audio channels. For example, for 5.1 multi-channel audio, the 5.1 multi-channel audio includes six audio channels: a left channel, a middle channel, a right channel, a left surround channel, a right surround channel, and a bass channel. Since the audio collected by the left channel, the middle channel, and the right channel often belongs to the same type, the left channel, the middle channel, and the right channel can be grouped as one audio channel. The left surround channel and the right surround channel are grouped as one audio channel. Accordingly, the audio channels belonging to the same audio channel group can be determined as the audio content similarity of the audio channel that meets the preset requirement. In this step, two audio channels belonging to the same pre-divided audio channel group can be selected as one channel group. For example, [left and right], [left and middle], [right and middle], and [left surround and right surround] are selected as channel groups.
[0115] Optionally, the step of extracting the correlation feature between the single-channel audios corresponding to the at least two single channels included in the channel group can specifically include:
[0116] In step 1022a, an inter-channel correlation parameter between specified audio frames in the single-channel audios corresponding to the at least two single channels is obtained. The inter-channel correlation parameter includes a parameter for representing the degree of correlation between the single-channel audios corresponding to the at least two single channels.
[0117] The number of specified audio frames is not less than 2, and the inter-channel correlation parameter can include one or more of an inter-channel phase difference, an inter-channel energy difference, and an inter-channel cross-correlation coefficient. In the embodiment of the present disclosure, the inter-channel correlation parameter can include one or more of an inter-channel phase difference (IPD), an inter-channel level difference (ILD), and an inter-channel cross-correlation coefficient (ICC). Of course, the inter-channel correlation parameter can also include other parameters, such as an inter-channel time difference (ICTD).
[0118] The specified audio frames can be all the audio frames included in the single-channel audio, can be part of the audio frames, for example, can be the 10th audio frame to the 30th audio frame. For any channel group, the at least two single-channel audios corresponding to the at least two single channels in the channel group are the at least two single-channel audios corresponding to the channel group. Accordingly, the inter-channel correlation parameters between the specified audio frames at the same position in the at least two single-channel audios can be calculated. Assuming that the at least two single-channel audios corresponding to the channel group are audio 1 and audio 2, the inter-channel correlation parameters between the 10th audio frame in audio 1 and the 10th audio frame in audio 2 can be calculated, the inter-channel correlation parameters between the 11th audio frame in audio 1 and the 11th audio frame in audio 2 can be calculated, and so on, and the inter-channel correlation parameters between the 30th audio frame in audio 1 and the 30th audio frame in audio 2 can be calculated. The calculation method of each type of inter-channel correlation parameter can refer to the existing calculation method, for example, a preset parameter calculation formula is used to calculate the inter-channel correlation parameters between the specified audio frames.
[0119] Step 1022b, for any of the inter-channel correlation parameters, determining the feature corresponding to the inter-channel correlation parameter according to the inter-channel correlation parameters between the specified audio frames.
[0120] Taking the inter-channel correlation parameters including the inter-channel phase difference, the inter-channel energy difference, and the inter-channel cross-correlation coefficient as an example. The feature corresponding to the inter-channel phase difference can be determined according to the inter-channel phase difference between the specified audio frames, specifically, the statistical mean and variance of the inter-channel phase difference between all the specified audio frames can be calculated as the feature corresponding to the inter-channel phase difference. The feature corresponding to the inter-channel energy difference can be determined according to the inter-channel energy difference between the specified audio frames, specifically, the statistical mean and variance of the inter-channel energy difference between all the specified audio frames can be calculated as the feature corresponding to the inter-channel energy difference. The feature corresponding to the inter-channel cross-correlation coefficient can be determined according to the inter-channel cross-correlation coefficient between the specified audio frames, specifically, the statistical mean and variance of the inter-channel cross-correlation coefficient between all the specified audio frames can be calculated as the feature corresponding to the inter-channel cross-correlation coefficient.
[0121] Step 1022c, generating the correlation feature based on the features corresponding to all the inter-channel correlation parameters.
[0122] In this step, the features corresponding to the various inter-channel correlation parameters extracted for the channel group can be spliced. For example, the features corresponding to the inter-channel phase difference, the features corresponding to the inter-channel energy difference, and the features corresponding to the inter-channel cross-correlation coefficient are spliced to obtain the correlation features corresponding to the channel group. The correlation features corresponding to the channel group can be D2 dimensions. Accordingly, for K channel groups, a feature with a dimension of D2*K can be finally obtained as the above correlation features. The specific value of D1 can be set according to actual needs, for example, D2 dimensions can be 40 dimensions.
[0123] It should be noted that when the number of single channels included in the channel group is greater than 2, the single channels included in the channel group can be divided into a group two by two. For any group, the inter-channel correlation parameters between the specified audio frames in the two single channel audios in the group are calculated. Then, the features corresponding to any inter-channel correlation parameter corresponding to the group are determined. For any inter-channel correlation parameter corresponding to the feature, the features corresponding to the inter-channel correlation parameters of the multiple groups can be averaged. Finally, the features corresponding to all the inter-channel correlation parameters obtained after averaging are spliced to obtain the final correlation features of the channel group.
[0124] Exemplarily, assuming that the channel group corresponds to three single channel audios: audio 1, audio 2, and audio 3, then audio 1 and audio 2 can be divided into group A, audio 1 and audio 3 can be divided into group B, and audio 2 and audio 3 can be divided into group C. The inter-channel phase difference, the inter-channel energy difference, and the inter-channel cross-correlation coefficient corresponding to group A are calculated based on audio 1 and audio 2, the inter-channel phase difference, the inter-channel energy difference, and the inter-channel cross-correlation coefficient corresponding to group B are calculated based on audio 1 and audio 3, and the inter-channel phase difference, the inter-channel energy difference, and the inter-channel cross-correlation coefficient corresponding to group C are calculated based on audio 2 and audio 3.
[0125] Then, the features corresponding to the inter-channel phase difference, the features corresponding to the inter-channel energy difference, and the features corresponding to the inter-channel cross-correlation coefficient corresponding to group A are determined, the features corresponding to the inter-channel phase difference, the features corresponding to the inter-channel energy difference, and the features corresponding to the inter-channel cross-correlation coefficient corresponding to group B are determined, and the features corresponding to the inter-channel phase difference, the features corresponding to the inter-channel energy difference, and the features corresponding to the inter-channel cross-correlation coefficient corresponding to group C are determined.
[0126] Then, the features corresponding to the inter-channel phase differences of the A group, the B group and the C group are averaged respectively to obtain the final features corresponding to the inter-channel phase differences. The features corresponding to the inter-channel energy differences of the A group, the B group and the C group are averaged respectively to obtain the final features corresponding to the inter-channel energy differences. The features corresponding to the inter-channel cross-correlation coefficients of the A group, the B group and the C group are averaged respectively to obtain the final features corresponding to the inter-channel cross-correlation coefficients. Finally, the final features corresponding to the inter-channel phase differences, the final features corresponding to the inter-channel energy differences and the final features corresponding to the inter-channel cross-correlation coefficients are spliced to obtain the correlation features between the three single channels included in the channel group.
[0127] In the embodiments of the present disclosure, by obtaining the inter-channel correlation parameters between the specified audio frames in the single-channel audios corresponding to the at least two single channels, the features corresponding to the inter-channel correlation parameters are determined based on the inter-channel correlation parameters between the specified audio frames, and finally, the correlation features are generated based on the features corresponding to all the inter-channel correlation parameters. Since the inter-channel correlation parameters can accurately represent the correlation between channels, the accuracy of the generated correlation features can be ensured to some extent.
[0128] Optionally, the step of inputting the audio features and the correlation features into the preset prediction model can specifically include:
[0129] In step 1031, the audio features and the correlation features are spliced to obtain spliced features.
[0130] In step 1032, the spliced features are input into the prediction model.
[0131] In the embodiments of the present disclosure, the features obtained by splicing the audio features and the correlation features can be used as the input of the prediction model. For example, the spliced features can be represented as D1*N+D2*K. Figure 2 is a schematic diagram of an audio processing process according to an example embodiment. As shown in Figure 2 The original audio is the multi-channel audio to be encoded, and the processing process can be implemented based on an encoding audio quality prediction module. Specifically, a plurality of specified channel audios can be obtained through a channel extraction / mixing module, and the audio features FEAintra of each specified channel audio can be extracted based on a single-channel feature extraction module. At the same time, the correlation features FEAinter can be extracted based on an inter-channel feature extraction module. Then, FEAintra and FEAinter are spliced and input into a neural network to finally obtain an audio quality rate curve. FEAintra can be understood as a single-channel feature, FEAinter can be understood as an inter-channel correlation feature, and the neural network can be a preset prediction model.
[0132] Further, the above-mentioned coding timbre prediction module can belong to an audio coding system framework. In the audio coding system framework, the original audio input can be further encoded by an encoder, and a code rate is calculated based on a timbre-code rate curve and a target timbre, i.e., a code rate corresponding to the target timbre is searched to determine a target code rate. Finally, the original audio is encoded by the encoder at the target code rate to obtain encoded audio.
[0133] In the embodiments of the present disclosure, for a multi-channel audio to be encoded, a target code rate is determined for the multi-channel audio based on single-channel audio features of the multi-channel audio and correlation features between audios, and the multi-channel audio to be encoded is adaptively encoded based on the target code rate, so that appropriate and more accurate target code rates can be allocated for different multi-channel audios to be encoded, and thus the processing effect of the encoding process is ensured.
[0134] In the embodiments of the present disclosure, the audio features and the correlation features are spliced, and the spliced features are taken as inputs of a preset prediction model. In this way, the prediction model can process the audio features and the correlation features, and thus the processing efficiency is improved to a certain extent.
[0135] Optionally, the prediction model is obtained by the following steps:
[0136] Step A, obtaining a plurality of sample encoding code rates of a sample multi-channel audio and timbres corresponding to the plurality of sample encoding code rates respectively.
[0137] The sample multi-channel audio can be randomly selected. The sample multi-channel audio can be multiple, and can be speech, music, environmental sound, or a mixture of several contents. For any sample multi-channel audio, an audio encoding algorithm such as a high-efficiency advanced audio coding (HE-AAC) algorithm can be used for L code rate encoding. The L code rates are sample encoding code rates, and the L code rates can be represented as R = [r1, r2, …, rL]. The encoded output audio can be represented as Y = [y1, y2, …, yL]. The specific value of L can be determined by the network structure of the prediction model to be trained, for example, the number of output neurons included in the last layer of the prediction model to be trained. In an implementation manner, L can be equal to 7, and R = [16, 24, 32, 40, 48, 56, 64] kbps (kilobits per second).
[0138] Further, the L kinds of audio Y = [y1, y2, …, yL] obtained after encoding the sample multi-channel audio according to the sample multi-channel audio and the L kinds of code rates can be used to determine the objective audio quality S = [s1, s2, …, sL] corresponding to each of the L kinds of code rates by using an objective audio quality evaluation algorithm. The objective audio quality can be a perceptual evaluation of audio quality (PEAQ) parameter or a parameter obtained by linearly fusing multiple objective audio quality indicators. Alternatively, the audio quality corresponding to each of the L kinds of audio obtained after encoding using the L kinds of code rates can also be determined by manual labeling, so that the audio quality corresponding to the sample encoding code rate is closer to the actual subjective perception of humans.
[0139] Correspondingly, the L kinds of code rates and the audio quality corresponding to each of the L kinds of code rates, i.e., L kinds of data pairs, can be obtained, wherein one data pair can include one kind of code rate and the audio quality corresponding to the code rate.
[0140] Step B, using the audio features and correlation features of the sample multi-channel audio as inputs of the to-be-trained prediction model, obtaining the multiple encoding code rates output by the to-be-trained prediction model and the audio quality corresponding to each of the multiple encoding code rates.
[0141] The implementation of extracting the audio features and correlation features of the sample multi-channel audio can refer to the implementation of extracting the audio features and correlation features of the to-be-encoded multi-channel audio. Further, the audio features and correlation features of the sample multi-channel audio can be spliced to obtain sample spliced features.
[0142] The to-be-trained prediction model can be a neural network, for example, specifically a multi-layer fully connected network including P layers, each layer having Q nodes, wherein P can be equal to 2 and Q can be equal to 100, to adapt to a low-computing-lightweight scenario. Of course, in the case of sufficient computing power, a multi-layer convolutional neural network (CNN), a long short-term memory (LSTM) network, and a deep neural network (DNN) can also be used as the to-be-trained prediction model. Correspondingly, a multi-channel spectrum input prediction model of the multi-channel audio can also be further obtained to provide more dimensions of data for the prediction model, thereby improving the accuracy of the prediction result.
[0143] Further, the to-be-trained prediction model can output L kinds of encoding code rates and the audio quality corresponding to each of the L kinds of encoding code rates based on the input.
[0144] Step D, adjusting model parameters of the to-be-trained prediction model based on the plurality of sample encoding code rates, the sound quality corresponding to each of the plurality of sample encoding code rates, the plurality of encoding code rates output by the to-be-trained prediction model, and the sound quality corresponding to each of the plurality of encoding code rates.
[0145] The sound quality corresponding to the sample encoding code rate is denoted as S, and the sound quality corresponding to the encoding code rate output by the to-be-trained prediction model is denoted as S'. In the embodiments of the present disclosure, a mean square error (MSE) function can be used as a loss function, and an error value can be calculated based on S and S' corresponding to the same code rate in the plurality of sample encoding code rates and the plurality of encoding code rates output by the to-be-trained prediction model. In order to minimize the loss function, the model parameters of the to-be-trained prediction model are adjusted in a gradient descent manner based on the error value.
[0146] Step E, in the case where the to-be-trained prediction model converges, the to-be-trained prediction model is determined as the prediction model.
[0147] In this step, the to-be-trained prediction model can be determined to converge in the case where the loss function reaches a minimum, or the number of times of adjusting the model parameters reaches a preset number threshold, or the calculated error value is less than a preset numerical threshold. Accordingly, the to-be-trained prediction model that converges is the prediction model.
[0148] It should be noted that the execution subject of the above model training process can be the same device as the execution subject of the above encoding process, or can be a different device.
[0149] In the embodiments of the present disclosure, the audio features and correlation features of the sample multi-channel audio are used as the input of the to-be-trained prediction model, so that the to-be-trained prediction model can learn more comprehensive features capable of representing the multi-channel audio in the training process, and thus the prediction model obtained by training can more accurately determine the predicted encoding code rate for the multi-channel audio.
[0150] Figure 3 is a block diagram of an audio processing apparatus according to an example embodiment, as shown in Figure 3 The apparatus 30 can include:
[0151] The first extraction module 301 is configured to perform extraction of audio features of a specified channel audio in the multi-channel audio to be encoded; the specified channel audio is obtained based on a single-channel audio included in the multi-channel audio;
[0152] The second extraction module 302 is configured to perform extraction of correlation features between at least part of the single-channel audios included in the multi-channel audio;
[0153] The first obtaining module 303 is configured to input the audio features and the correlation features into a preset prediction model, and obtain a plurality of encoding code rates of the multi-channel audio output and audio qualities corresponding to the plurality of encoding code rates respectively, which are output by the prediction model.
[0154] The first determining module 304 is configured to determine a target code rate of the multi-channel audio based on the plurality of encoding code rates and the audio qualities corresponding to the plurality of encoding code rates respectively.
[0155] In an optional embodiment, the apparatus 30 further comprises:
[0156] The selecting module is configured to select N single-channel audios from the single-channel audios included in the multi-channel audio as the specified channel audios.
[0157] And / or,
[0158] The first generating module is configured to generate N audio groups based on the single-channel audios included in the multi-channel audio; the N is a positive integer, and there is at least one audio group including at least two single-channel audios in the N audio groups.
[0159] The second generating module is configured to generate one specified channel audio based on the single-channel audios included in any audio group for the audio group.
[0160] In an optional embodiment, the second extracting module 302 is specifically configured to:
[0161] determine a channel group based on the plurality of audio channels corresponding to the multi-channel audio; at least two single channels are included in one channel group;
[0162] extract a correlation feature between the single-channel audios corresponding to the at least two single channels included in any channel group for the channel group.
[0163] In an optional embodiment, the second extracting module 302 is specifically further configured to:
[0164] obtain an inter-channel correlation parameter between specified audio frames in the single-channel audios corresponding to the at least two single channels; the inter-channel correlation parameter includes a correlation degree between the single-channel audios corresponding to the at least two single channels;
[0165] determine a feature corresponding to any inter-channel correlation parameter according to the inter-channel correlation parameter between the specified audio frames for the inter-channel correlation parameter;
[0166] generate the correlation features based on the features corresponding to all the inter-channel correlation parameters.
[0167] In an optional embodiment, the second extraction module 302 is further configured to perform:
[0168] In a case where the total number of audio channels corresponding to the multi-channel audio is equal to 2, two audio channels corresponding to the multi-channel audio are determined as one channel group.
[0169] In a case where the total number of audio channels corresponding to the multi-channel audio is greater than 2, at least two audio channels corresponding to audio content with a similarity meeting a preset requirement are divided into the same channel group.
[0170] In an optional embodiment, the prediction model is obtained by training the following modules:
[0171] The second acquisition module is configured to perform acquisition of a plurality of sample encoding bit rates of a sample multi-channel audio and audio qualities corresponding to the plurality of sample encoding bit rates.
[0172] The third acquisition module is configured to perform acquisition of a plurality of encoding bit rates output by the to-be-trained prediction model and audio qualities corresponding to the plurality of encoding bit rates, by taking the audio features and the correlation features of the sample multi-channel audio as inputs of the to-be-trained prediction model.
[0173] The adjustment module is configured to perform adjustment of model parameters of the to-be-trained prediction model based on the plurality of sample encoding bit rates, the audio qualities corresponding to the plurality of sample encoding bit rates, the plurality of encoding bit rates output by the to-be-trained prediction model, and the audio qualities corresponding to the plurality of encoding bit rates.
[0174] The second determination module is configured to perform determination of the to-be-trained prediction model as the prediction model in a case where the to-be-trained prediction model converges.
[0175] In an optional embodiment, the first acquisition module 303 is configured to perform:
[0176] The audio features and the correlation features are spliced to obtain spliced features.
[0177] The spliced features are input into the prediction model.
[0178] To sum up, the audio processing apparatus provided by the embodiment of the present disclosure extracts the audio features of the specified channel audio in the multi-channel audio to be encoded, and the specified channel audio is obtained based on the mono-channel audio included in the multi-channel audio. The correlation features between at least part of the mono-channel audio included in the multi-channel audio are extracted. The audio features and the correlation features are input into a preset prediction model to obtain the multiple encoding code rates output by the prediction model for the multi-channel audio and the audio quality corresponding to each of the multiple encoding code rates. Based on the multiple encoding code rates and the audio quality corresponding to each of the multiple encoding code rates, the target code rate of the multi-channel audio is determined. In this way, compared with the way of directly determining the encoding code rate based on the audio content, in the embodiment of the present disclosure, when the multi-channel audio is processed, the audio features of the specified channel audio obtained based on the mono-channel audio in the multi-channel audio and the correlation features between at least part of the mono-channel audio included in the multi-channel audio are extracted, and the target code rate of the multi-channel audio is determined based on the audio features and the correlation features. Since the channel correlation of the multi-channel audio will affect the required encoding code rate, the audio features and the correlation features can more comprehensively represent the multi-channel audio, so that the determined target code rate can be more suitable for the multi-channel audio to some extent, and thus the accuracy of the target code rate determined for the multi-channel audio is improved, thereby improving the subsequent encoding effect.
[0179] According to an embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory for storing processor-executable instructions, wherein the processor is configured to execute the steps of the audio processing method in any one of the above embodiments.
[0180] According to an embodiment of the present disclosure, a storage medium is also provided, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can perform the steps of the audio processing method in any one of the above embodiments.
[0181] According to an embodiment of the present disclosure, a computer program product is also provided, the computer program product comprises readable program instructions, when the readable program instructions are executed by the processor of the electronic device, the electronic device can perform the steps of the audio processing method in any one of the above embodiments.
[0182] Figure 4is a block diagram of an apparatus for audio processing according to an example embodiment. Wherein, the apparatus 900 can include a processing component 902, a memory 904, a power supply component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, a communication component 916, and a processor 920. The processing component 902 can include one or more processors 920 to execute instructions to complete all or part of steps of the above method for audio processing. In an example embodiment, a storage medium can also provide a storage medium including instructions, e.g., the memory 904 including instructions, which can be executed by the processor 920 of the apparatus 900 to complete the above method. Optionally, the storage medium can be a non-transitory computer-readable storage medium, e.g., the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0183] Figure 5 is a block diagram of another apparatus for audio processing according to an example embodiment.
[0184] Wherein, the apparatus 1000 can include a processing component 1022, a memory 1032, an input / output interface 1058, a network interface 1050, and a power supply component 1026. The apparatus 1000 can be provided as a server. The applications stored in the memory 1032 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1022 is configured to execute the instructions to perform the above method for audio processing.
[0185] The user information (including but not limited to the user's device information, the user's personal information, etc.) and related data involved in the present disclosure are information authorized by the user or authorized by each party.
[0186] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The present disclosure is intended to cover any and all variations of the present disclosure that come within the scope of the following claims, along with all equivalents and alternatives to the features of the present disclosure that are apparent to those skilled in the art in light of the specification. The specification and examples given are intended as illustrative only and are not intended to limit the true scope and spirit of the present disclosure.
[0187] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes in the embodiments can be effected by those skilled in the art without departing from the scope of the application. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An audio processing method, characterized by, The method comprises: extracting audio features of specified channel audio in the multi-channel audio to be encoded; the specified channel audio is obtained based on single-channel audio included in the multi-channel audio; extracting correlation features between single-channel audios corresponding to at least two single channels included in the multi-channel audio based on multiple single-channel audios corresponding to the multi-channel audio; inputting the audio features and the correlation features into a preset prediction model to obtain multiple encoding code rates output by the prediction model for the multi-channel audio and audio qualities corresponding to the multiple encoding code rates respectively; determining a target code rate of the multi-channel audio based on the multiple encoding code rates and the audio qualities corresponding to the multiple encoding code rates respectively.
2. The method of claim 1, wherein, Before the step of extracting the audio features of the specified channel audio in the multi-channel audio to be encoded, the method further comprises: selecting N single-channel audios from single-channel audios included in the multi-channel audio as the specified channel audio; N is a positive integer; and / or, generating N audio groups based on single-channel audios included in the multi-channel audio; there is an audio group including at least two single-channel audios in the N audio groups; for any audio group, generating one specified channel audio based on single-channel audios included in the audio group.
3. The method of claim 1, wherein, The step of extracting the correlation features between single-channel audios corresponding to at least two single channels included in the multi-channel audio based on multiple single-channel audios corresponding to the multi-channel audio comprises: determining a channel group based on multiple audio channels corresponding to the multi-channel audio; at least two single channels are included in one channel group; for any channel group, extracting correlation features between single-channel audios corresponding to the at least two single channels included in the channel group.
4. The method of claim 3, wherein, The step of extracting the correlation features between single-channel audios corresponding to the at least two single channels included in the channel group comprises: obtaining inter-channel correlation parameters between specified audio frames in the single-channel audios corresponding to the at least two single channels; the inter-channel correlation parameters include a parameter for representing a degree of correlation between the single-channel audios corresponding to the at least two single channels; for any inter-channel correlation parameter, determining a feature corresponding to the inter-channel correlation parameter according to inter-channel correlation parameters between the specified audio frames; generating the correlation features based on features corresponding to all the inter-channel correlation parameters.
5. The method of claim 3, wherein, The step of determining a channel group based on multiple audio channels corresponding to the multi-channel audio comprises: in a case where a total number of audio channels corresponding to the multi-channel audio is equal to 2, determining two audio channels corresponding to the multi-channel audio as one channel group; in a case where the total number of audio channels corresponding to the multi-channel audio is greater than 2, dividing at least two audio channels corresponding to audio content with a similarity meeting a preset requirement into a same channel group.
6. The method according to any one of claims 1 to 5, characterized in that, The prediction model is trained by the following method: obtaining multiple sample encoding code rates of sample multi-channel audios and audio qualities corresponding to the multiple sample encoding code rates respectively; The audio features and the correlation features of the sample multi-channel audio are input into a preset prediction model, and a plurality of encoding code rates and audio qualities corresponding to the plurality of encoding code rates output by the prediction model are obtained. Based on the plurality of sample encoding code rates, the audio qualities corresponding to the plurality of sample encoding code rates, the plurality of encoding code rates output by the prediction model, and the audio qualities corresponding to the plurality of encoding code rates, the model parameters of the prediction model are adjusted. In a case where the prediction model converges, the prediction model is determined as the trained prediction model.
7. The method according to any one of claims 1 to 5, characterized in that, The inputting of the audio features and the correlation features into the preset prediction model comprises: The audio features and the correlation features are spliced to obtain spliced features; The spliced features are input into the prediction model.
8. An audio processing apparatus, characterized by comprising: The device comprises: A first extraction module configured to extract audio features of specified channel audio in multi-channel audio to be encoded; the specified channel audio is obtained based on single-channel audio included in the multi-channel audio; A second extraction module configured to extract correlation features between single-channel audios included in the multi-channel audio based on a plurality of single-channel audios corresponding to the multi-channel audio; A first acquisition module configured to input the audio features and the correlation features into a preset prediction model, and acquire a plurality of encoding code rates and audio qualities corresponding to the plurality of encoding code rates output by the prediction model for the multi-channel audio; A first determination module configured to determine a target code rate of the multi-channel audio based on the plurality of encoding code rates and the audio qualities corresponding to the plurality of encoding code rates.
9. An electronic device, comprising: Comprise: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method of any one of claims 1 to 7.
10. A storage medium, characterized by When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device performs the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Joint stereo audio coding and decoding method and device
CN115346540A
Audio encoding device, decoding device, method, and program
CN1969318A