Coding and decoding processing method and device for audios with different sound channel numbers, equipment and medium

By decomposing the two-channel audio into mono audio in audio processing and using a mono network for encoding and decoding, the problem of different audio encoding and decoding models in the prior art is solved, and efficient and unified encoding and decoding processing is achieved.

CN120148528APending Publication Date: 2025-06-13PANORAMIC SOUND (BEIJING) INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311714570.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, different audio codec models are needed to process audio with different channel numbers, resulting in complexity and inconsistency.

Method used

By obtaining mono audio and/or dual-channel audio in the to-processed audio, the mono audio is encoded and decoded using a mono network, and the two-channel audio is decomposed into the main component channel and the residual component channel, and the mono network is encoded and coded using a mono network to obtain the codec stream corresponding to the original channel.

Benefits of technology

It realizes that the deep learning network can only use mono audio encoding to complete the encoding and decoding of audio numbers of different channels, improving the encoding and decoding efficiency and quality, and avoiding the complexity and inconsistency caused by using different networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148528A_ABST
    Figure CN120148528A_ABST
Patent Text Reader

Abstract

The invention provides a coding and decoding processing method and device for audios with different sound channel numbers, equipment and a medium. The method comprises the following steps: acquiring a monaural audio and / or a dual-channel audio in to-be-processed audios; for the monaural audio, adopting a monaural audio coding and decoding deep learning network to carry out coding and decoding processing on the monaural audio to obtain a corresponding coding and decoding code stream; the method comprises the following steps: for a dual-channel audio, carrying out dual-channel conversion processing based on a left channel audio and a right channel audio in the dual-channel audio to obtain a principal component channel audio and a residual component channel audio, and respectively carrying out coding and decoding processing on the principal component channel audio and the residual component channel audio by adopting a single-channel network, and finally, coding and decoding code streams respectively corresponding to the left channel audio and the right channel audio are obtained. According to the method and the device, the technical effects that the multi-channel audio is decomposed into a plurality of monaural audios, and then the monaural audios obtained through decomposition are coded and decoded by adopting the monaural network can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to audio processing technologies, and in particular, to a method, apparatus, device, and medium for encoding and decoding different-channel audio. Background Art

[0002] Currently, for audio encoding and decoding algorithms based on deep learning, different models are used for audio with different numbers of channels. For example, for the most common mono-channel audio and stereo audio, two types of deep learning networks need to be designed and trained using different training sets; moreover, in applications, according to whether the input sound source is mono-channel or stereo, the corresponding deep learning network is called for encoding and decoding.

[0003] It can be seen that in the related art, there is a problem that different audio encoding and decoding models need to be used to process audio with different numbers of channels. Summary of the Invention

[0004] The present application provides a method, apparatus, device, and medium for encoding and decoding different-channel audio, which is used to solve the problem in the related art that different audio encoding and decoding models need to be used to process audio with different numbers of channels, and achieve the technical effect of adaptively decomposing stereo audio into mono-channel audio and then processing the decomposed mono-channel audio using the processing method of mono-channel audio.

[0005] On the one hand, the present application provides a method for encoding and decoding different-channel audio, and the method includes:

[0006] Obtain mono-channel audio and / or stereo audio in the audio to be processed;

[0007] For the mono-channel audio, perform encoding and decoding processing on the mono-channel audio using a mono-channel network to obtain a corresponding encoded and decoded bitstream;

[0008] For the stereo audio, perform dual-channel transformation processing on the left-channel audio and the right-channel audio in the stereo audio to obtain a principal component channel audio and a residual component channel audio, and respectively perform encoding and decoding processing on the principal component channel audio and the residual component channel audio using a mono-channel network to finally obtain encoded and decoded bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

[0009] An optional implementation manner, obtaining mono-channel audio and / or stereo audio in the audio to be processed includes:

[0010] If the audio to be processed is multi-channel audio, obtain the natural attributes of the channels of the audio to be processed;

[0011] Split the audio to be processed into the stereo audio and / or the mono audio according to the natural attributes of the channels of the audio to be processed.

[0012] An optional implementation manner is to separately perform encoding and decoding processing on the principal component channel audio and the residual component channel audio by using a mono network, so as to finally obtain the encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively, including:

[0013] Separate the principal component channel audio and the residual component channel audio by using a mono network to perform encoding and decoding processing, and obtain the encoding and decoding bitstreams corresponding to the principal component channel audio and the residual component channel audio respectively;

[0014] Perform inverse stereo transformation processing on the encoding and decoding bitstreams corresponding to the principal component channel audio and the residual component channel audio respectively to obtain the encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

[0015] An optional implementation manner is to separately perform encoding and decoding processing on the principal component channel audio and the residual component channel audio by using a mono network, including:

[0016] Obtain the correlation between the left-channel audio and the right-channel audio in the stereo audio;

[0017] Determine the bitrates of the principal component channel audio and the residual component channel audio respectively according to the correlation between the left-channel audio and the right-channel audio in the stereo audio;

[0018] Perform encoding and decoding processing on the principal component channel audio and the residual component channel audio according to the bitrates of the principal component channel audio and the residual component channel audio respectively.

[0019] An optional implementation manner is to perform stereo transformation processing on the left-channel audio and the right-channel audio in the stereo audio to obtain the principal component channel audio and the residual component channel audio, including:

[0020] Perform sample-by-sample addition on the left-channel audio and the right-channel audio in the stereo audio, and then divide by 2 to obtain the principal component channel audio;

[0021] Perform sample-by-sample subtraction on the left-channel audio and the right-channel audio in the stereo audio, and then divide by 2 to obtain the residual component channel audio.

[0022] An optional implementation manner is to perform stereo transformation processing on the left-channel audio and the right-channel audio in the stereo audio to obtain the principal component channel audio and the residual component channel audio, including:

[0023] Perform principal component analysis on the left-channel audio and the right-channel audio in the stereo audio to obtain the angular parameter of the principal component transformation;

[0024] Based on the principal component transformation angular parameter, perform transformation processing on the left-channel audio and the right-channel audio in the stereo audio respectively to obtain the principal component channel audio and the residual component channel audio.

[0025] An optional implementation manner is to perform encoding and decoding processing on the mono audio by using a mono network to obtain the corresponding encoding and decoding bitstreams, including:

[0026] Perform signal processing on the mono audio to obtain multiple audio segments;

[0027] Use multiple encoder networks in the mono network to continuously perform dimensionality reduction processing on the multiple audio segments to obtain multi-dimensional vectors of the multiple audio segments;

[0028] Use the quantization network in the mono network to retrieve the codeword corresponding to the multi-dimensional vector in the codebook, so that the index value of the codeword is used to replace the multi-dimensional vector as the encoding bitstream of the mono audio in subsequent processing;

[0029] Use the inverse quantization network in the mono network to retrieve the corresponding multi-dimensional vector from the codebook according to the index value of the codeword;

[0030] Use multiple decoder networks in the mono network to continuously perform dimensionality increase processing on the multi-dimensional vector to obtain the decoding bitstream of the mono audio.

[0031] An optional implementation manner is that before using the quantization network in the mono network to retrieve the codeword corresponding to the multi-dimensional vector, the method further includes:

[0032] If the multi-dimensional vector does not match the codebook tensor of the quantization network, use the mapping network in the mono network to process the multi-dimensional vector to match the codebook tensor of the quantization network.

[0033] On the other hand, the present application provides an encoding and decoding processing device for audio with different numbers of channels, and the device includes:

[0034] An acquisition module, configured to acquire mono audio and / or stereo audio in the audio to be processed;

[0035] A mono audio processing module, configured to perform encoding and decoding processing on the mono audio by using a mono network to obtain the corresponding encoding and decoding bitstreams for the mono audio;

[0036] A two-channel audio processing module is configured to perform two-channel transformation processing on the two-channel audio based on the left-channel audio and the right-channel audio in the two-channel audio to obtain a principal component channel audio and a residual component channel audio, and respectively perform encoding and decoding processing on the principal component channel audio and the residual component channel audio by using a mono network, so as to finally obtain encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

[0037] On the other hand, the present application provides an electronic device, including: a processor, and a memory connected to the above-mentioned processor; the above-mentioned memory stores computer-executable instructions; the above-mentioned processor executes the computer-executable instructions stored in the above-mentioned memory to implement the method as described in any one of the above.

[0038] On the other hand, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the above.

[0039] On the other hand, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method as described in any one of the above.

[0040] The encoding and decoding processing method, device, equipment and medium for audio with different numbers of channels provided by the present application obtain mono audio and / or two-channel audio in the audio to be processed; for the mono audio, perform encoding and decoding processing on the mono audio by using a mono network to obtain a corresponding encoding and decoding bitstream; for the two-channel audio, perform two-channel transformation processing on the left-channel audio and the right-channel audio in the two-channel audio to obtain a principal component channel audio and a residual component channel audio, and respectively perform encoding and decoding processing on the principal component channel audio and the residual component channel audio by using a mono network, so as to finally obtain encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

[0041] In the embodiment of the present application, for mono audio, encoding and decoding processing is performed on the mono audio by using a mono network to obtain a corresponding encoding and decoding bitstream; for two-channel audio, first perform two-channel transformation processing on the left-channel audio and the right-channel audio in the two-channel audio to obtain a principal component channel audio and a residual component channel audio, and then respectively perform encoding and decoding processing on the principal component channel audio and the residual component channel audio by using a mono network, so as to finally obtain encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively. That is, only a mono network needs to be used in the embodiment of the present application, which can solve the problem in the related art that different audio encoding and decoding models need to be used to process audio with different numbers of channels, and achieve the technical effect of decomposing two-channel audio into mono audio and then processing the decomposed mono audio by using the processing method of mono audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0043] Figure 1 It is a schematic flowchart of a method for encoding and decoding audio with different numbers of channels provided by an embodiment of the present application;

[0044] Figure 2 It is a schematic diagram of encoding and decoding processing for an optional stereo video provided by an embodiment of the present application;

[0045] Figure 3 It is a schematic diagram of encoding and decoding processing for another optional stereo video provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram of encoding and decoding processing for yet another optional stereo video provided by an embodiment of the present application;

[0047] Figure 5 It is a schematic diagram of obtaining a mono audio by processing an optional stereo audio provided by an embodiment of the present application;

[0048] Figure 6 It is a schematic diagram of obtaining a mono audio by processing an optional multi-channel audio provided by an embodiment of the present application;

[0049] Figure 7 It is a schematic diagram of an optional training network structure provided by an embodiment of the present application;

[0050] Figure 8 It is a schematic diagram of encoding processing for an optional mono audio (obtained by processing a multi-channel audio as shown by Figure 5 ) provided by an embodiment of the present application;

[0051] Figure 9a It is a schematic diagram of decoding processing for an optional stereo audio provided by an embodiment of the present application;

[0052] Figure 9b It is a schematic diagram of an optional inverse transform processing for a stereo provided by an embodiment of the present application;

[0053] Figure 10 It is a block diagram of a device for encoding and decoding audio with different numbers of channels provided by an embodiment of the present application;

[0054] Figure 11 It is a schematic diagram of an electronic device provided by an embodiment of the present application.

[0055] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments

[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0057] Currently, for audio codecs based on deep learning, different models are adopted for audio with different numbers of channels. For example, for the most common mono audio and stereo audio, two types of deep learning networks need to be designed and trained using different training sets; and in applications, according to whether the input sound source is mono or stereo, the corresponding deep learning network is called for encoding and decoding.

[0058] The encoding and decoding processing method for audio with different numbers of channels provided by the present application aims to solve the above technical problems in the prior art.

[0059] The following uses specific embodiments to detail the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0060] Figure 1 is a schematic flowchart of an encoding and decoding processing method for audio with different numbers of channels provided by an embodiment of the present application, as Figure 1 shown, the method includes:

[0061] S101, obtaining mono audio and / or stereo audio in the audio to be processed;

[0062] S102, for the above mono audio, using a mono network to perform encoding and decoding processing on the mono audio to obtain the corresponding encoding and decoding bitstream;

[0063] S103. For the above-mentioned stereo audio, perform stereo transformation processing on the left-channel audio and the right-channel audio in the above-mentioned stereo audio to obtain a principal component channel audio and a residual component channel audio, and respectively use a mono network to perform encoding and decoding processing on the above-mentioned principal component channel audio and the above-mentioned residual component channel audio, so as to finally obtain encoding and decoding bitstreams corresponding to the above-mentioned left-channel audio and the above-mentioned right-channel audio respectively.

[0064] The encoding and decoding processing method for audio with different numbers of channels provided by the embodiments of the present application can be understood as a deep learning-based adaptive encoding and decoding processing method for audio with different numbers of channels. It should be noted that the solution of the present application can implement the use of only a mono audio encoding deep learning network (hereinafter referred to as a mono network), that is, it can complete the encoding and decoding processing of audio with different numbers of channels.

[0065] First, obtain the mono audio and / or stereo audio in the audio to be processed. For example, the audio to be processed can directly be one or more mono audio or stereo audio, or the audio to be processed can also be multi-channel audio. In this case, it is necessary to first split the audio to be processed into multiple stereo audio and multiple mono audio according to the natural attributes between channels.

[0066] For the above-mentioned mono audio, use a mono network to perform encoding and decoding processing on the above-mentioned mono audio to obtain the corresponding encoding and decoding bitstream.

[0067] In one example, as Figure 2 shown in a schematic diagram of an optional encoding and decoding processing of stereo video, as Figure 2 shown, for stereo audio (for example, left-channel audio L and right-channel audio R), the following processing steps need to be performed, and then use a mono network for encoding and decoding processing: Extract the principal component channel audio L' and the residual component channel audio R' from L and R.

[0068] After that, during encoding, use a mono network to encode the principal component channel audio L' to generate an encoded bitstream a; use a mono network to encode the residual component channel audio R' to generate an encoded bitstream b; during decoding, use a mono network to decode the encoded bitstream a and the encoded bitstream b respectively to generate the corresponding decoded bitstreams a' and b'; perform an inverse transformation on the decoded bitstreams a' and b' to obtain the final decoded bitstreams L" and R", where L" is the decoded bitstream corresponding to the left-channel audio L, and R" is the decoded bitstream corresponding to the right-channel audio R.

[0069] In the embodiment of the present application, for monophonic audio, a monophonic network is used to perform encoding and decoding processing on the monophonic audio to obtain the corresponding encoded and decoded bitstreams. For stereophonic audio, first, based on the left-channel audio and the right-channel audio in the stereophonic audio, stereophonic transformation processing is performed to obtain the principal component channel audio and the residual component channel audio, and then the monophonic network is respectively used to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio, so as to finally obtain the encoded and decoded bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

[0070] By adopting the embodiment of the present application, encoding and decoding processing of audio with different channel numbers can be realized, and the encoding and decoding efficiency and quality are improved. This method can automatically identify monophonic audio and stereophonic audio in multichannel audio according to the natural attributes of the channels of the audio to be processed, and perform encoding and decoding processing using the same monophonic network, avoiding the complexity and inconsistency caused by using different networks. This method can also utilize the correlation between the left and right channels in stereophonic audio to convert it into the principal component channel and the residual component channel, thereby reducing the redundancy and data volume of the encoded and decoded bitstreams.

[0071] That is, the embodiment of the present application only needs to use a monophonic network, which can solve the problem in the related art that different audio encoding and decoding models need to be used to process audio with different channel numbers, and realize the technical effect of decomposing stereophonic audio into monophonic audio and then processing the decomposed monophonic audio in the processing manner of monophonic audio.

[0072] An optional implementation manner, obtaining monophonic audio and / or stereophonic audio in the audio to be processed, includes:

[0073] S201, if the audio to be processed is multichannel audio, obtain the natural attributes of the channels of the audio to be processed;

[0074] S202, according to the natural attributes of the channels of the audio to be processed, split the audio to be processed to obtain the stereophonic audio and / or the monophonic audio.

[0075] This optional embodiment specifies the manner of obtaining monophonic audio and stereophonic audio in multichannel audio, which can enhance the adaptability and flexibility of this method to different types of audio to be processed, and can perform corresponding splitting operations according to whether the audio to be processed contains monophonic or stereophonic.

[0076] In the embodiments of the present application, for the case where the audio to be processed is multi-channel audio, the number of channels can be first split into multiple stereo audio and multiple mono audio according to the natural attributes between the channels, and then the stereo audio obtained by splitting is encoded and decoded according to the above stereo audio processing method, and the mono audio obtained by splitting is encoded and decoded according to the above mono audio processing method. For example, the left-channel audio L and the right-channel audio R can be used as a stereo audio, and also as Figure 3 shown, the left surround audio Ls and the right surround audio Rs are used as a stereo channel, as Figure 4 shown, the center c and the low frequency effects Lfe are used as two mono audio. In this way, the stereo audio and mono audio obtained by the above splitting can be encoded and decoded respectively according to the above stereo audio processing method and mono audio processing method.

[0077] An alternative embodiment is to encode and decode the above main component channel audio and the above residual component channel audio respectively using a mono network to finally obtain the encoded and decoded bitstreams corresponding to the above left-channel audio and the above right-channel audio respectively, including:

[0078] S301, encoding and decoding the above main component channel audio and the above residual component channel audio respectively using a mono network to obtain the encoded and decoded bitstreams corresponding to the above main component channel audio and the above residual component channel audio respectively;

[0079] S302, performing a stereo inverse transform process on the encoded and decoded bitstreams corresponding to the above main component channel audio and the above residual component channel audio respectively to obtain the encoded and decoded bitstreams corresponding to the above left-channel audio and the above right-channel audio respectively.

[0080] Using the above embodiments, it can be ensured that the original stereo information is not lost during the encoding and decoding process, and the encoded and decoded bitstreams of the original stereo can be restored to the encoded and decoded bitstreams that are the same as or approximate to the original stereo audio.

[0081] In an alternative example, specifically during encoding, in the embodiments of the present application, the main component channel audio L' can be encoded using a mono network to generate an encoded bitstream a; the residual component channel audio R' can be encoded using a mono network to generate an encoded bitstream b.

[0082] Specifically during decoding, in the embodiments of the present application, the encoded bitstream a and the encoded bitstream b can be decoded respectively using a mono network to generate the corresponding decoded bitstream a' and decoded bitstream b'; an inverse transform is performed on the decoded bitstream a' and the decoded bitstream b' to obtain the final decoded bitstreams L” and R”. Among them, L” is the decoded bitstream corresponding to the left-channel audio L, and R” is the decoded bitstream corresponding to the right-channel audio R.

[0083] An optional implementation method is to separately perform encoding and decoding processing on the above-mentioned principal component channel audio and the above-mentioned residual component channel audio by using a mono network, including:

[0084] S401. Obtain the correlation between the left-channel audio and the right-channel audio in the above-mentioned stereo audio;

[0085] S402. Determine the respective bitrates of the above-mentioned principal component channel audio and the above-mentioned residual component channel audio according to the correlation between the left-channel audio and the right-channel audio in the above-mentioned stereo audio;

[0086] S403. Perform encoding and decoding processing on the above-mentioned principal component channel audio and the above-mentioned residual component channel audio according to the respective bitrates of the above-mentioned principal component channel audio and the above-mentioned residual component channel audio.

[0087] In the above embodiment, by separately performing encoding and decoding processing on the principal component channel and the residual component channel by using a mono network, it is possible to dynamically adjust the respective bitrates of the principal component channel and the residual component channel according to the correlation between the left and right channels in the stereo audio, thereby optimizing the data volume and quality of the encoded and decoded bitstreams. It is also possible to improve the adaptability and flexibility of this method for different types of stereo audio, and corresponding encoding and decoding processing can be performed according to the high or low correlation between the left and right channels in the stereo audio.

[0088] As an optional embodiment, considering the strong correlation between the left and right channels in the stereo audio, in actual encoding and decoding, a relatively high bitrate can be used for channel L', and a relatively low bitrate can be used for channel R'. Then, encoding and decoding processing can be performed on the above-mentioned principal component channel audio and the above-mentioned residual component channel audio according to the respective bitrates of the above-mentioned principal component channel audio and the above-mentioned residual component channel audio.

[0089] As an optional embodiment, according to actual needs, a mono network different from L' can be used to perform encoding and decoding on R'.

[0090] As another optional embodiment, according to actual needs, a mono network different from L' can be used to perform encoding and decoding on Lfe.

[0091] An optional implementation method is to perform stereo transformation processing on the basis of the left-channel audio and the right-channel audio in the above-mentioned stereo audio to obtain the principal component channel audio and the residual component channel audio, including:

[0092] S501. Perform a per-sample addition on the left-channel audio and the right-channel audio in the above-mentioned stereo audio, and then divide by 2 to obtain the above-mentioned principal component channel audio;

[0093] S502, perform sample - by - sample subtraction on the left - channel audio and the right - channel audio in the above - mentioned stereo audio, and then divide by 2 to obtain the above - mentioned residual component channel audio.

[0094] In an alternative embodiment, for the mono audio obtained by processing stereo audio, as Figure 5 shown, the principal component channel audio L’ and the residual component channel audio R’ obtained by converting the stereo audio (left channel L and right channel R). Among them, the principal component channel audio L’=(L + R) / 2, and the residual component channel audio R’=(L - R) / 2; where ‘+’ is sample - by - sample addition and ‘-’ is sample - by - sample subtraction.

[0095] For the mono audio obtained by processing multi - channel audio, adopt the transformation method as Figure 5 or as Figure 6 shown, transform the multi - channel audio channels (L, R, C, Lfe, Ls, Rs) to obtain each mono audio L’, R’, C, Lfe, Ls’, Rs’. Among them, L’=(L + R) / 2, R’=(L - R) / 2, Ls’=(Ls + Rs) / 2, Rs’=(Ls - Rs) / 2, the two channels of C and Lfe remain unchanged, and the audio with other numbers of channels can be processed to obtain each mono audio by referring to the alternative methods described in the above embodiments.

[0096] An alternative implementation manner, based on the left - channel audio and the right - channel audio in the stereo audio, perform stereo transformation processing to obtain the principal component channel audio and the residual component channel audio, including:

[0097] S601, perform principal component analysis on the left - channel audio and the right - channel audio in the above - mentioned stereo audio to obtain the angle parameter beta of the principal component transformation;

[0098] S602, based on the above - mentioned principal component transformation angle parameter beta, perform transformation processing on the left - channel audio and the right - channel audio in the above - mentioned stereo audio respectively to obtain the above - mentioned principal component channel audio and the residual component channel audio.

[0099] In an alternative embodiment, perform principal component analysis (Principal Component Analysis, PCA) on the stereo audio (left channel L and right channel R) to obtain the angle parameter beta of the principal component transformation of the stereo audio,

[0100] The principal component transformation matrix of the stereo audio can be expressed as:

[0101] Perform principal component transformation on the left channel L and the right channel R based on the principal component transformation angle parameter beta of the principal component transformation to obtain the principal component channel audio L' and the residual component channel audio R': The principal component channel audio L' = (cos(beta)L + sin(beta)R), and the residual component channel audio R' = (-sin(beta)L + cos(beta)R); where, "+" is addition by sample point, "-" is subtraction by sample point, cos() is the cosine function, and sin() is the sine function.

[0102] For the mono audio obtained by processing multi-channel audio, transform the multi-channel audio channels (L, R, C, Lfe, Ls, Rs) to obtain each mono audio L', R', C, Lfe, Ls', Rs', where L' = (cos(beta1)L + sin(beta1)R);

[0103] R' = (-sin(beta1)L + cos(beta1)R);

[0104] Ls' = (cos(beta2)Ls + sin(beta2)Rs);

[0105] Rs' = (-sin(beta2)Ls + cos(beta2)Rs).

[0106] Among them, beta1 is the principal component transformation angle parameter obtained by performing principal component analysis on the L and R audio, and beta2 is the principal component transformation angle parameter obtained by performing principal component analysis on the Ls and Rs audio; the C and Lfe channels remain unchanged, and the audio of other channel numbers can be processed to obtain each mono audio according to the optional methods described in the above embodiments.

[0107] In the above embodiments, by determining the principal component channel and the residual component channel based on the left and right channels in the stereo audio, it is possible to simply and effectively convert the stereo audio into the principal component channel and the residual component channel, reducing the computational complexity and time overhead. It can also ensure that the original stereo audio can be completely restored during the inverse transformation process, avoiding the loss of audio information.

[0108] An optional implementation manner is to perform encoding and decoding processing on the above mono audio using a mono network to obtain the corresponding encoding and decoding bitstreams, including:

[0109] S701, perform signal processing on the above mono audio to obtain multiple audio segments;

[0110] S702, continuously perform dimensionality reduction processing on the above multiple audio segments using multiple encoder networks in the above mono network to obtain multi-dimensional vectors of the above multiple audio segments;

[0111] S703. Use the quantization network in the above-mentioned mono network to retrieve the codeword in the codebook corresponding to the above-mentioned multi-dimensional vector, so that in subsequent processing, the index value of the above-mentioned codeword is used to replace the above-mentioned multi-dimensional vector as the encoded code stream of the above-mentioned mono audio;

[0112] S704. Use the inverse quantization network in the above-mentioned mono network to retrieve the corresponding above-mentioned multi-dimensional vector from the above-mentioned codebook according to the index value of the above-mentioned codeword;

[0113] S705. Use multiple decoder networks in the above-mentioned mono network to continuously perform dimension elevation processing on the above-mentioned multi-dimensional vector to obtain the decoded code stream of the above-mentioned mono audio.

[0114] In the above optional embodiment, by using the mono network to perform encoding and decoding processing on the mono audio, efficient and accurate encoding and decoding processing of the mono audio can be realized, improving the encoding and decoding quality and efficiency. The dimension reduction and elevation processing of the audio segment can also be realized by using multiple encoder networks and decoder networks in the mono network, reducing the data volume and computational complexity. The quantization and inverse quantization processing of the multi-dimensional vector of the audio segment can also be realized by using the quantization network and inverse quantization network in the mono network, improving the encoding and decoding accuracy and stability.

[0115] An optional implementation manner. Before using the quantization network in the above-mentioned mono network to retrieve the codeword in the codebook corresponding to the above-mentioned multi-dimensional vector, the above method further includes: if the above-mentioned multi-dimensional vector does not match the codebook tensor of the above-mentioned quantization network, then use the mapping network in the above-mentioned mono network to process the above-mentioned multi-dimensional vector to match the codebook tensor of the above-mentioned quantization network.

[0116] By performing mapping processing before quantization processing in the embodiments of the present application, unified and standardized processing can be realized between different types or scales of audio segments, improving the adaptability and flexibility of the method to different input audio segments, and avoiding encoding and decoding errors or distortions caused by the mismatch between the multi-dimensional vector and the codebook tensor.

[0117] Optionally, in the embodiments of the present application, the purpose of the mapping network is to make the tensor output by the encoder network match the codebook tensor of the quantization network; if a codebook tensor that matches the tensor output by the encoder network is used, in this case, the mapping network can be not used.

[0118] Optionally, in the embodiments of the present application, there is a codebook in the quantization network, and the purpose is to find the codeword closest to the input vector from the codebook and replace the input vector with the index of the closest codeword in subsequent processing.

[0119] Optionally, in the embodiments of the present application, inverse quantization is the process of finding the corresponding codeword from the codebook using the index of the codeword.

[0120] In an alternative embodiment, a monophonic audio training set is constructed. Considering the generality of the training set, monophonic audio generated from audio with different numbers of channels is added to the training set. A schematic diagram of an alternative training network structure in the embodiments of the present application is as Figure 7 shown. For example, the specific encoding and decoding processing procedures include:

[0121] 1. The training samples are audio with a monophonic sampling rate of 48 kHz, specifically, monophonic audio generated from audio with different numbers of channels.

[0122] 2. The encoder network, decoder network, and mapping network are all composed of convolutional networks; the purpose of the mapping network is to make the tensor output by the encoder network match the codebook tensor of the quantization network; it is also possible to use a codebook tensor that matches the output tensor of the encoder network. In this case, the mapping network can be not used.

[0123] 3. The time-domain dimensionality reduction multiples of encoder networks 1, 2, 3, and 4 are 3, 4, 5, and 5 respectively; the time-domain dimensionality increase multiple of the decoder network is the same as that of the corresponding encoder network.

[0124] 4. The codebook of the quantization network is (1024, 64), that is, it is composed of 1024 vectors with a length of 64.

[0125] 5. The role of the mapping network is to transform the output of the encoder network to be equal to the length of the codeword (for example, the length is 64). Therefore, the encoded code stream is the index value of the codeword corresponding to the vector output by the quantization network.

[0126] An optional training process: The monophonic audio in the training set is processed, truncated into xi:(1,9600), that is, a monophonic segment with 9600 sample points; it is reduced to a multi-dimensional vector of (64,3200) through the encoding network 1; it is reduced to a multi-dimensional vector of (128,800) through the encoding network 2; it is reduced to a multi-dimensional vector of (256,160) through the encoding network 3; it is reduced to a multi-dimensional vector of (512,32) through the encoding network 4; it is transformed into (64,32) through the mapping network, that is, 32 vectors with a length of 64; then, the index of the nearest codeword is found from the codebook in the quantization network, and 32 indexes are found for the 32 vectors with a length of 64; each index consists of 10 bits (10 bits cover the index range of the codebook); the 32 vectors with a length of 64 are transformed into 32 10-bit vectors = 320 bits, and 32 10-bit index values are output; according to the index values, the corresponding vectors are found from the codebook (inverse quantization) and output to the decoder network 4. Then, it is upsampled to a multi-dimensional vector of (256,160) through the decoder network 4; it is upsampled to a multi-dimensional vector of (128,800) through the decoder network 3; it is upsampled to a multi-dimensional vector of (64,3200) through the decoder network 2; it is upsampled to a multi-dimensional vector of (1,9600) through the decoder network 1; finally, a vector with 9600 monophonic sample points is output, that is, yi:(1,9600).

[0127] In the embodiment of the present application, the loss function Loss of the above training process may include two parts, Loss = Loss1 + Loss2:

[0128] Loss function 1: The loss function 1 of the input vector xi(1,9600) and the output vector yi(1,9600), that is

[0129] Loss function 2: The loss function 2 of the vector ei in the codebook corresponding to the actual vector and the actual vector gi, that is

[0130] This loss function is the objective function, and the gradient descent method is used to adjust the model parameters (each encoding network, each decoding network, mapping network, quantization network, etc.), that is: use Adjust the model parameters, where θ represents the model parameters to be adjusted, Loss represents the reconstruction loss, is the partial derivative symbol.

[0131] In an optional example, the encoding process provided by the embodiment of the present application is as follows Figure 8 An optional monophonic audio (by Figure 5Taking the schematic diagram of encoding processing for multi-channel audio processing as an example, the sample length is 96,000 sample points for stereo (2, 96,000), and the two channels of the stereo are respectively labeled as L and R; the stereo audio is subjected to the stereo transformation processing as shown in Figure 5 to obtain two mono channels L' and R', and the dimensions of the mono channels L' and R' are both (1, 96,000), where L' = (L + R) / 2 and R' = (L - R) / 2; L' is passed through encoder network 1, encoder network 2, encoder network 3, and encoder network 4, and L' is successively reduced in dimension to (64, 32,000), (128, 8,000), (256, 1,600), (512, 320).

[0132] After that, (512, 320) is transformed to (64, 320) through the mapping network; for 320 vectors of length 64, the indices of the closest codewords are found from the codebook respectively, and 320 10-bit indices are output; R' is similarly processed through the above steps to obtain another set of 320 10-bit indices; the two sets of indices in the above steps are combined and output, and 320x10 + 320x10 bit is the encoded bitstream of the stereo audio.

[0133] Optionally, in one example, the decoding process in the embodiments of the present application is as shown in Figure 9a : the encoded bitstream is the encoded bitstream of the stereo 320*10 + 320x10 bit; the first set of 320x10 bit encoded bitstream is taken out, and 320 vectors (64, 320) are found from the codebook of the quantization network according to 320 indices, that is, inverse quantization; (64, 320) is successively dimension-increased through decoder network 4, decoder network 3, decoder network 2, and decoder network 1 to obtain new vectors (256, 1,600), (128, 8,000), (64, 32,000), and finally the first vector L' (1, 96,000) is obtained; similarly, the second set of 320*10 bit encoded bitstream is taken out and processed through the above to obtain the second vector R' (1, 96,000); the stereo inverse transformation processing as shown in Figure 9b is performed on L' and R' to obtain the encoded and decoded bitstream L'' of the left-channel audio and the encoded and decoded bitstream R'' of the right-channel audio: L'' = L' + R', R'' = L' - R'; both L'' and R'' are vectors of (1, 96,000); L'' and R'' form a new vector (2, 96,000), which is the decoded bitstream of the stereo audio.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0135] According to one or more embodiments of the present application, there is provided an encoding and decoding processing device for audio with different numbers of channels. Figure 10 As shown in the structural block diagram of an encoding and decoding processing device for audio with different numbers of channels provided by an embodiment of the present application, Figure 10 as shown, the above device includes:

[0136] An acquisition module 601, configured to acquire monophonic audio and / or stereophonic audio in the audio to be processed;

[0137] A monophonic audio processing module 602, configured to perform encoding and decoding processing on the above monophonic audio by using a monophonic network to obtain a corresponding encoded and decoded bitstream;

[0138] A stereophonic audio processing module 603, configured to perform stereophonic transformation processing on the above stereophonic audio based on the left-channel audio and the right-channel audio in the above stereophonic audio to obtain a principal component channel audio and a residual component channel audio, and respectively perform encoding and decoding processing on the above principal component channel audio and the above residual component channel audio by using a monophonic network to finally obtain encoded and decoded bitstreams corresponding to the above left-channel audio and the above right-channel audio respectively.

[0139] According to one or more embodiments of the present application, the above acquisition module includes:

[0140] A first acquisition unit, configured to acquire the natural channel attribute of the audio to be processed if the audio to be processed is a multi-channel audio;

[0141] A splitting unit, configured to split the audio to be processed into the above stereophonic audio and / or the above monophonic audio according to the natural channel attribute of the audio to be processed.

[0142] An optional implementation manner, the above stereophonic audio processing module includes:

[0143] A first processing unit, configured to respectively perform encoding and decoding processing on the above principal component channel audio and the above residual component channel audio by using a monophonic network to obtain encoded and decoded bitstreams corresponding to the above principal component channel audio and the above residual component channel audio respectively;

[0144] A second processing unit, configured to perform a two-channel inverse transform process on the respective codec bitstreams corresponding to the above-mentioned main component channel audio and the above-mentioned residual component channel audio, to obtain the codec bitstreams corresponding to the above-mentioned left-channel audio and the above-mentioned right-channel audio respectively.

[0145] An optional implementation manner, the first processing unit includes:

[0146] A second obtaining unit, configured to obtain the correlation between the left-channel audio and the right-channel audio in the above-mentioned two-channel audio;

[0147] A determining unit, configured to determine the respective bitrates of the above-mentioned main component channel audio and the above-mentioned residual component channel audio according to the correlation between the left-channel audio and the right-channel audio in the above-mentioned two-channel audio;

[0148] A third processing unit, configured to perform codec processing on the above-mentioned main component channel audio and the above-mentioned residual component channel audio according to the respective bitrates of the above-mentioned main component channel audio and the above-mentioned residual component channel audio.

[0149] An optional implementation manner, the above-mentioned two-channel audio processing module includes:

[0150] A fourth processing unit, configured to perform a sample-by-sample addition process on the left-channel audio and the right-channel audio in the above-mentioned two-channel audio to obtain the above-mentioned main component channel audio;

[0151] A fifth processing unit, configured to perform a sample-by-sample subtraction process on the left-channel audio and the right-channel audio in the above-mentioned two-channel audio to obtain the above-mentioned residual component channel audio.

[0152] An optional implementation manner, the above-mentioned mono-channel audio processing module includes:

[0153] A sixth processing unit, configured to perform signal processing on the above-mentioned mono-channel audio to obtain a plurality of audio segments;

[0154] A dimensionality reduction processing unit, configured to continuously perform dimensionality reduction processing on the above-mentioned plurality of audio segments by using a plurality of encoder networks in the above-mentioned mono-channel network to obtain multi-dimensional vectors of the above-mentioned plurality of audio segments;

[0155] A seventh processing unit, configured to retrieve, by using a quantization network in the above-mentioned mono-channel network, a codeword corresponding to the above-mentioned multi-dimensional vector in a codebook, so that in subsequent processing, the index value of the above-mentioned codeword is used to replace the above-mentioned multi-dimensional vector as the encoded bitstream of the above-mentioned mono-channel audio;

[0156] A retrieval processing unit, configured to retrieve, by using an inverse quantization network in the above-mentioned mono-channel network, the corresponding above-mentioned multi-dimensional vector from the above-mentioned codebook according to the index value of the above-mentioned codeword;

[0157] A dimensionality increase processing unit, configured to continuously perform dimensionality increase processing on the above-mentioned multi-dimensional vector by using multiple decoder networks in the above-mentioned mono network to obtain a decoded bitstream of the above-mentioned mono audio.

[0158] An optional implementation manner, the above-mentioned apparatus further includes:

[0159] An eighth processing unit, configured to, if the above-mentioned multi-dimensional vector does not match the codebook tensor of the above-mentioned quantization network, use the mapping network in the above-mentioned mono network to process the above-mentioned multi-dimensional vector to match the codebook tensor of the above-mentioned quantization network.

[0160] In an exemplary embodiment, an embodiment of the present application further provides an electronic device, including: a processor, and a memory connected to the above-mentioned processor;

[0161] The above-mentioned memory stores computer-executable instructions;

[0162] The above-mentioned processor executes the computer-executable instructions stored in the above-mentioned memory to implement the method as described in any one of the above.

[0163] In an exemplary embodiment, an embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the above.

[0164] In an exemplary embodiment, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method as described in any one of the above.

[0165] To implement the above-mentioned embodiments, an embodiment of the present application further provides an electronic device. Refer to Figure 11 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present application. The electronic device 700 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, messaging devices, game consoles, medical devices, fitness devices, personal digital assistants (Personal Digital Assistant, abbreviated as PDA), tablet computers (Portable Android Device, abbreviated as PAD), portable multimedia players (Portable Media Player, abbreviated as PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0166] AsFigure 11 As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0167] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0168] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present application are executed.

[0169] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0170] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0171] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0172] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0174] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0175] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0176] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0177] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present application are pointed out by the following claims.

[0178] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is limited only by the appended claims.

Claims

1. A method for encoding and decoding processing of audio with different numbers of channels, characterized in that, the method includes: obtaining mono-channel audio and / or stereo audio in the audio to be processed; for the mono-channel audio, using a mono-channel network to perform encoding and decoding processing on the mono-channel audio to obtain a corresponding encoding and decoding bitstream; for the stereo audio, performing stereo transformation processing on the left-channel audio and the right-channel audio in the stereo audio to obtain a principal component channel audio and a residual component channel audio, and respectively using a mono-channel network to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio to finally obtain encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

2. The method according to claim 1, characterized in that, obtaining mono-channel audio and / or stereo audio in the audio to be processed includes: if the audio to be processed is multi-channel audio, obtaining the natural attributes of the channels of the audio to be processed; according to the natural attributes of the channels of the audio to be processed, splitting the audio to be processed to obtain the stereo audio and / or the mono-channel audio.

3. The method according to claim 1, characterized in that, respectively using a mono-channel network to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio to finally obtain encoding and decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively, includes: respectively using a mono-channel network to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio to obtain encoding and decoding bitstreams corresponding to the principal component channel audio and the residual component channel audio respectively; performing stereo inverse transformation processing on the decoding bitstreams corresponding to the principal component channel audio and the residual component channel audio respectively to obtain decoding bitstreams corresponding to the left-channel audio and the right-channel audio respectively.

4. The method according to any one of claims 1 to 3, characterized in that, respectively using a mono-channel network to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio, includes: obtaining the correlation between the left-channel audio and the right-channel audio in the stereo audio; according to the correlation between the left-channel audio and the right-channel audio in the stereo audio, determining the bitrates of the principal component channel audio and the residual component channel audio respectively; according to the bitrates of the principal component channel audio and the residual component channel audio respectively, performing encoding and decoding processing on the principal component channel audio and the residual component channel audio.

5. The method according to any one of claims 1 to 3, characterized in that, performing stereo transformation processing on the left-channel audio and the right-channel audio in the stereo audio to obtain a principal component channel audio and a residual component channel audio, includes: performing a sample-by-sample addition on the left-channel audio and the right-channel audio in the stereo audio, and then dividing by 2 to obtain the principal component channel audio; performing a sample-by-sample subtraction on the left-channel audio and the right-channel audio in the stereo audio, and then dividing by 2 to obtain the residual component channel audio.

6. The method according to any one of claims 1 to 3, wherein, performing binaural transformation processing on the left-channel audio and the right-channel audio in the binaural audio to obtain a principal component channel audio and a residual component channel audio, including: performing principal component analysis on the left-channel audio and the right-channel audio in the binaural audio to obtain an angle parameter for principal component transformation; based on the principal component transformation angle parameter, respectively performing transformation processing on the left-channel audio and the right-channel audio in the binaural audio to obtain the principal component channel audio and the residual component channel audio.

7. The method according to claim 1, wherein, using a mono network to perform encoding and decoding processing on the mono audio to obtain a corresponding encoding and decoding code stream, including: performing signal processing on the mono audio to obtain a plurality of audio segments; using a plurality of encoder networks in the mono network to continuously perform dimensionality reduction processing on the plurality of audio segments to obtain multi-dimensional vectors of the plurality of audio segments; using a quantization network in the mono network to retrieve a codeword corresponding to the multi-dimensional vector in a code table, so that in subsequent processing, the index value of the codeword is used to replace the multi-dimensional vector as the encoding code stream of the mono audio; using an inverse quantization network in the mono network to retrieve the corresponding multi-dimensional vector from the code table according to the index value of the codeword; using a plurality of decoder networks in the mono network to continuously perform dimensionality increase processing on the multi-dimensional vector to obtain the decoding code stream of the mono audio.

8. The method according to claim 7, wherein, before using the quantization network in the mono network to retrieve the codeword corresponding to the multi-dimensional vector in the code table, the method further includes: if the multi-dimensional vector does not match the code table tensor of the quantization network, using a mapping network in the mono network to process the multi-dimensional vector to match the code table tensor of the quantization network.

9. An encoding and decoding processing device for audio with different numbers of channels, wherein, the device includes: an acquisition module, configured to acquire mono audio and / or binaural audio in the audio to be processed; a mono audio processing module, configured to, for the mono audio, use a mono network to perform encoding and decoding processing on the mono audio to obtain a corresponding encoding and decoding code stream; a binaural audio processing module, configured to, for the binaural audio, perform binaural transformation processing on the left-channel audio and the right-channel audio in the binaural audio to obtain a principal component channel audio and a residual component channel audio, and respectively use a mono network to perform encoding and decoding processing on the principal component channel audio and the residual component channel audio, so as to finally obtain encoding and decoding code streams corresponding to the left-channel audio and the right-channel audio respectively.

10. An electronic device, wherein, including: a processor, and a memory connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 8.