Audio encoding / decoding negotiation method and terminal device
By obtaining the audio codec configuration information and hardware parameters of the terminal device, and negotiating the appropriate codec parameters, the problems of poor audio codec performance and resource waste in the existing technology are solved, resulting in a better user experience and resource utilization.
Patent Information
- Application Number
- PCT/CN2025/082517
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-14
- Publication Date
- 2025-11-27
AI Technical Summary
Existing audio codec negotiation methods, when selecting highly complex speech codecs, result in stable call connections but poor audio codec performance and waste of resources.
By obtaining the audio codec configuration information of both parties' terminal devices, and combining it with hardware parameters to determine the codec negotiation result, codec parameters that are compatible with the other party's capabilities are selected for negotiation to ensure the best user experience is provided under hardware constraints.
It enables more flexible call configuration while meeting hardware constraints, improves audio encoding and decoding performance, and avoids resource waste.
Smart Images

Figure CN2025082517_27112025_PF_FP_ABST
Abstract
Description
Method for negotiating audio codec and terminal device
[0001] The present application claims priority from the Chinese patent application No. 202410355475.6 filed on March 25, 2024, and entitled "Method for negotiating audio codec and terminal device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of audio codec, and in particular, to a method for negotiating audio codec and a terminal device. BACKGROUND
[0003] Compared with a speech codec using a single channel, an immersive speech codec supports more types of audio signals, supports a larger range of code rates, and includes rendering characteristics in addition to codec characteristics. For example, in the audio codec standard being developed by the 3rd Generation Partnership Project (3GPP), different audio codec standards support multiple types of signals, different types of signals support different numbers of channels and code rates, and the rendering characteristics of different types of signals support rendering of various types of signals to binaural and standard loudspeaker arrays. Due to supporting more signal channel numbers, a larger range of code rates, and additional rendering characteristics, the immersive speech codec has a very significant increase in maximum computational complexity and storage complexity compared to the traditional speech codec.
[0004] In order to support a speech codec with high complexity, in an existing method for negotiating audio codec, the speech codec is classified according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result can only select a level with lower complexity for the call to ensure that the call connection is stable. Although this method meets the speech requirements of the call connection, it has the problem of poor audio codec negotiation effect. SUMMARY
[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a method for negotiating audio codec, characterized in that the method is applied to a first terminal device, and the method comprises:
[0007] obtaining audio codec configuration information of a second terminal device, the audio codec configuration information of the second terminal device being used to indicate the audio codec capability of the second terminal device;
[0008] determining audio encoding / decoding configuration information of the first terminal device according to hardware parameters of the first terminal device, the hardware parameters of the first terminal device comprising at least one of the following: a parameter of an audio acquisition module of the first terminal device, a parameter of an audio playback module of the first terminal device;
[0009] determining a coding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the coding / decoding negotiation result comprising: a first coding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first coding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device;
[0010] sending the coding / decoding negotiation result of the first terminal device to the second terminal device.
[0011] In some embodiments of the present application, the first coding / decoding parameter comprises: a first coding / decoding mode performed by the first terminal device when encoding / decoding the audio signal to be processed, and a first coding / decoding rate corresponding to the first coding / decoding mode.
[0012] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the hardware parameters of the first terminal device comprises:
[0013] when the hardware parameters of the first terminal device satisfy a preset first hardware configuration condition, determining the first audio encoding / decoding configuration information according to the hardware parameters of the first terminal device;
[0014] when the hardware parameters indicated by the battery endurance information do not satisfy the first hardware configuration condition, determining the second audio encoding / decoding configuration information according to the hardware parameters of the first terminal device;
[0015] wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0016] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the hardware parameters of the first terminal device comprises:
[0017] determining the audio encoding / decoding configuration information of the first terminal device according to the hardware parameters of the first terminal device and the audio encoding / decoding capability of the first terminal device.
[0018] In some embodiments of the present application, the hardware parameters further comprise at least one of the following: a parameter of an audio coding / decoding chip of the first terminal device, a parameter of a battery of the first terminal device, and a parameter of an off-chip memory of the first terminal device other than the audio coding / decoding chip.
[0019] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0020] The battery endurance information of the first terminal device is obtained according to the battery parameters of the first terminal device;
[0021] The audio encoding / decoding configuration information of the first terminal device is updated according to the battery endurance information, to obtain the updated audio encoding / decoding configuration information of the first terminal device.
[0022] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0023] The audio signal type and the sampling rate supported by the audio acquisition module of the first terminal device are determined according to the parameters of the audio acquisition module of the first terminal device;
[0024] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the sampling rate supported by the audio acquisition module of the first terminal device.
[0025] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0026] The audio signal type and the code rate supported by the audio codec chip of the first terminal device are determined according to the parameters of the audio codec chip of the first terminal device;
[0027] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the code rate supported by the audio codec chip of the first terminal device.
[0028] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0029] The audio signal type and the code rate supported by the off-chip memory of the first terminal device are determined according to the parameters of the off-chip memory of the first terminal device, except for the audio codec chip;
[0030] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the code rate supported by the off-chip memory of the first terminal device.
[0031] In some embodiments of the present application, the method further includes:
[0032] The audio service standard type supported by the audio codec chip of the second terminal device is obtained;
[0033] When the audio service standard type supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device is the preset first audio service standard type, the step of obtaining the audio codec configuration information of the second terminal device is triggered.
[0034] In some embodiments of the present application, the audio codec configuration information of the second terminal device includes: a codec mode supported by the second terminal device, and a rate set corresponding to the codec mode supported by the second terminal device.
[0035] The audio codec configuration information of the first terminal device includes: a codec mode supported by the first terminal device, and a rate set corresponding to the codec mode supported by the first terminal device.
[0036] In some embodiments of the present application, the codec negotiation result of the first terminal device is determined according to the audio codec configuration information of the first terminal device and the audio codec configuration information of the second terminal device, including:
[0037] The audio codec negotiation order is determined according to the audio codec negotiation priority policy of the first terminal device, and the audio codec negotiation order includes: performing audio encoding negotiation first or performing audio decoding negotiation first.
[0038] The audio encoding negotiation includes: determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio codec configuration information of the first terminal device and the audio decoding capability of the second terminal device.
[0039] The audio decoding negotiation includes: determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio codec configuration information of the first terminal device and the audio encoding capability of the second terminal device.
[0040] In some embodiments of the present application, the encoding negotiation result of the first terminal device is determined according to the audio encoding capability indicated by the audio codec configuration information of the first terminal device and the audio decoding capability of the second terminal device, including:
[0041] When the highest encoding mode supported by the first terminal device is the first audio mode, and the decoding mode supported by the second terminal device includes the first audio mode, the encoding negotiation result of the first terminal device includes the first audio mode.
[0042] The first code rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device is determined.
[0043] The encoding negotiation result of the first terminal device includes the maximum code rate in the first code rate intersection.
[0044] In some embodiments of the application, the decoding negotiation result of the first terminal device is determined according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device, comprising:
[0045] According to the highest decoding mode supported by the first terminal device is the second audio mode, and the encoding mode supported by the second terminal device includes the second audio mode, the decoding negotiation result of the first terminal device includes the second audio mode.
[0046] Determine the second code rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device.
[0047] Determine the decoding negotiation result of the first terminal device includes the maximum code rate in the second code rate intersection.
[0048] In some embodiments of the application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0049] In some embodiments of the application, the first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), high order ambisonic (HOA), or first order ambisonic (FOA).
[0050] The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6kbps, 13.2kbps, 32kbps, 48kbps.
[0051] In some embodiments of the application, the method further comprises:
[0052] Send an audio negotiation request to the second terminal device.
[0053] In some embodiments of the application, the audio negotiation request includes: the audio encoding / decoding configuration information of the first terminal device.
[0054] In some embodiments of the application, the method further comprises:
[0055] Receive session description protocol information from the second terminal device;
[0056] Determine the audio encoding / decoding configuration information of the second terminal device from the session description protocol information.
[0057] In a second aspect, the embodiments of the present application further provide a terminal device, specifically a first terminal device, and the method comprises the following steps.
[0058] an obtaining module, configured to obtain audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device;
[0059] a determining module, configured to determine audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device comprising at least one of the following: a parameter of an audio acquisition module of the first terminal device, a parameter of an audio playback module of the first terminal device;
[0060] a negotiating module, configured to determine an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result comprising: a first encoding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device;
[0061] a sending module, configured to send the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0062] In the second aspect of the present application, the constituent modules of the first terminal device can also perform the steps described in the foregoing first aspect and various possible implementation manners, for details, refer to the foregoing description of the first aspect and various possible implementation manners.
[0063] In a third aspect, the embodiments of the present application provide a computer readable storage medium, and the computer readable storage medium stores instructions, when the instructions run on a computer, the computer executes the method in the foregoing first aspect.
[0064] In a fourth aspect, the embodiments of the present application provide a computer program product containing instructions, when the instructions run on a computer, the computer executes the method in the foregoing first aspect.
[0065] In a fifth aspect, the embodiments of the present application provide a communication apparatus, which can include a terminal device or a chip and the like entity, and the communication apparatus comprises a processor, a memory, the memory is used to store instructions, and the processor is used to execute the instructions in the memory, so that the communication apparatus executes the method in any one of the foregoing first aspect.
[0066] In a sixth aspect, the present application provides a chip system, which comprises a processor for supporting a terminal device to implement the functions involved in the above aspects, such as transmitting or processing the data and / or information involved in the above methods. In a possible design, the chip system further comprises a memory, and the memory is configured to store necessary program instructions and data of the terminal device. The chip system can be composed of a chip, or can comprise a chip and other discrete devices.
[0067] In a seventh aspect, an embodiment of the present application provides a chip, which comprises one or more interface circuits and one or more processors. The interface circuit is configured to receive a signal from a memory of an electronic device, and transmit the signal to the processor, and the signal comprises computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device performs the method in the first aspect or any possible implementation manner of the first aspect.
[0068] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0069] First, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device. Then, the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the hardware parameters of the first terminal device, and the hardware parameters of the first terminal device at least comprise the battery parameter of the first terminal device. Next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the encoding / decoding negotiation result comprises the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device. Finally, the first terminal device transmits the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiments of the present application, the two terminal devices for audio negotiation each have audio encoding / decoding configuration information, and the audio encoding / decoding configuration information describes the encoding / decoding capability that can be supported by the terminal under the constraint of the hardware parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiments of the present application, the optimal user experience can be provided for the two parties as much as possible under the condition of meeting the hardware parameter constraints of the two parties of the call. BRIEF DESCRIPTION OF DRAWINGS
[0070] FIG. 1 is a schematic diagram of the composition structure of an audio processing system provided by an embodiment of the present application;
[0071] FIG. 2a is a schematic diagram of an audio encoder and an audio decoder applied to a terminal device provided by an embodiment of the present application;
[0072] Fig. 2b is a schematic diagram of an audio encoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0073] Fig. 2c is a schematic diagram of an audio decoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0074] Fig. 3a is a schematic diagram of a multi-channel encoder and a multi-channel decoder applied to a terminal device according to an embodiment of the present application;
[0075] Fig. 3b is a schematic diagram of a multi-channel encoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0076] Fig. 3c is a schematic diagram of a multi-channel decoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0077] Fig. 4 is a schematic diagram of a negotiation method of audio encoding / decoding according to an embodiment of the present application;
[0078] Fig. 5 is a schematic diagram of a negotiation process of audio negotiation by two terminals through SDP according to an embodiment of the present application;
[0079] Fig. 6 is a schematic diagram of a negotiation process of audio encoding / decoding mode and code rate by two terminals according to an embodiment of the present application;
[0080] Fig. 7 is a schematic diagram of a negotiation result of encoding / decoding after negotiation by two terminals according to an embodiment of the present application;
[0081] Fig. 8 is a schematic diagram of another negotiation process of audio encoding / decoding mode and code rate by two terminals according to an embodiment of the present application;
[0082] Fig. 9 is a schematic diagram of a structure of a first terminal device according to an embodiment of the present application;
[0083] Fig. 10 is a schematic diagram of a structure of a second terminal device according to an embodiment of the present application;
[0084] Fig. 11 is a schematic diagram of another structure of a first terminal device according to an embodiment of the present application;
[0085] Fig. 12 is a schematic diagram of another structure of a second terminal device according to an embodiment of the present application. DETAILED DESCRIPTION
[0086] Embodiments of the present application will be described below with reference to the accompanying drawings.
[0087] The terms "first", "second", and the like in the description and in the claims of the present application and above-described drawings do not necessarily have an ordinal number meaning but are used to distinguish similar objects from each other. It should be understood that the terms so used are interchangeable under appropriate circumstances and are merely employed as a form of description to facilitate the comprehension of the embodiments of the present application. Furthermore, the terms "comprise", "to comprise", "comprising", "include", "to include", "including", "contain", "to contain", "containing", "characterized by" and any variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, system, product or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, system, product or apparatus.
[0088] The code rate of traditional monophonic speech codec is between several kbps and tens or 128 kbps. Immersive speech codec supports more signal types, supports a larger range of code rates, and contains rendering features in addition to codec features compared with monophonic speech codec. Taking the 3GPP ongoing Immersive Voice and Audio Services (IVAS) speech / audio codec standard as an example, the signal types supported by IVAS speech / audio codec include monophonic, stereo, multi-channel, multi-object, higher order ambisonics (HOA) signal or first order ambisonics (FOA) signal, MASA, etc. Among them, the multi-channel signal supports up to 7.1.4 channel format, i.e. 12 channels, and the HOA signal supports up to 3rd order HOA, i.e. 16 channels, and the supported code rate ranges from the lowest single-channel 5.9 kbps to 768 kbps for 3rd order HOA. Rendering supports rendering of various signal types to binaural and standard loudspeaker arrays. More signal channels, larger code rate range, and additional rendering features make the maximum computational complexity and storage complexity of immersive speech codec significantly higher than that of traditional speech codec, which will lead to a significant increase in the difficulty of landing chips of such codecs.
[0089] In order to solve the problem of high complexity of the speech codec, the speech codec is classified according to complexity, but when two terminals supporting different levels establish a call, the current negotiation result is to select a lower complexity level for the call to ensure stable call connection. This method meets the speech requirements of the call connection, but has the problem of poor audio codec negotiation effect. In addition, this method meets the speech requirements of the call connection, but the lower level terminal cannot decode all types of code streams, which causes the coding and decoding limitation of the audio signal of the higher level terminal, and has the problem of poor audio signal coding and decoding quality. In addition, when the levels of the two terminals are not equal, the higher level terminal also has the problem of waste of coding and decoding capability.
[0090] Based on the above description, the embodiment of the present application provides an audio coding / decoding negotiation method, which can perform accurate Codec negotiation when establishing a call, and select the most appropriate coding and decoding combination according to the actual capability of the communication platform of each party, so as to ensure optimal call experience and not waste power consumption.
[0091] The embodiment of the present application provides an audio coding technology, in particular, provides a three-dimensional audio coding technology for three-dimensional audio signals, and specifically provides an encoding technology for representing three-dimensional audio signals by using fewer channels, so as to improve the traditional audio coding system. Audio coding (or commonly referred to as coding) includes two parts of audio encoding and audio decoding. Audio encoding is performed at the source side, including processing (for example, compressing) the original audio to reduce the amount of data required to represent the audio, so as to more efficiently store and / or transmit. Audio decoding is performed at the destination side, including inverse processing relative to the encoder, to reconstruct the original audio. The encoding part and the decoding part are also collectively referred to as coding. The implementation of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0092] The technical solution of the embodiment of the present application can be applied to various audio processing systems. As shown in FIG. 1, it is a schematic diagram of the composition structure of the audio processing system provided by the embodiment of the present application. The audio processing system 100 can include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 can be used to generate a code stream, and then the audio encoding code stream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the code stream, then perform the audio decoding function of the audio decoding device 102, and finally obtain the reconstructed signal.
[0093] In the embodiments of the present application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices and core network devices requiring transcoding, for example, the audio encoding device can be an audio encoder of the terminal device, the wireless device or the core network device. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices and core network devices requiring transcoding, for example, the audio decoding device can be an audio decoder of the terminal device, the wireless device or the core network device. For example, the audio encoder can include a wireless access network, a media gateway of a core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc., and the audio encoder can also be an audio encoder applied to a virtual reality (VR) streaming service.
[0094] In the embodiments of the present application, taking the audio encoding module (audio encoding and audio decoding) applied to a virtual reality streaming (VR streaming) service as an example, the processing flow of the end-to-end audio signal includes: the audio signal A is preprocessed (audio PReprocessing) after being acquired by the acquisition module, the preprocessing operation includes filtering out the low frequency part of the signal, which can be divided by 20Hz or 50Hz, and extracting the orientation information of the signal, then the encoding processing (audio encoding) and the packaging (file / segment encapsulation) are performed, and then the delivery (delivery) is performed to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), performs binaural rendering (audio rendering) processing on the decoded signal, and maps the rendered signal to the listener's headphones (headphones), which can be independent headphones or headphones on glasses devices.
[0095] As shown in FIG. 2a, a schematic diagram of the audio encoder and the audio decoder provided by the embodiment of the present application applied to a terminal device is shown. For each terminal device, an audio encoder, a channel encoder, an audio decoder and a channel decoder can be included. Specifically, the channel encoder is used to perform channel encoding on the audio signal, and the channel decoder is used to perform channel decoding on the audio signal. For example, the first terminal device 20 can include a first audio encoder 201, a first channel encoder 202, a first audio decoder 203 and a first channel decoder 204. The second terminal device 21 can include a second audio decoder 211, a second channel decoder 212, a second audio encoder 213 and a second channel encoder 214. The first terminal device 20 is connected to a first network communication device 22 via a wireless or wired connection, the first network communication device 22 and a second network communication device 23 are connected via a digital channel, and the second terminal device 21 is connected to the second network communication device 23 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device.
[0096] In the audio communication, the terminal device as a sending end first performs audio acquisition, encodes the acquired audio signal, and then performs channel encoding before transmitting the signal in a digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding according to the received signal to obtain a code stream, and then performs audio decoding to restore the audio signal, which is played back by the terminal device as the receiving end.
[0097] As shown in FIG. 2b, a schematic diagram of the audio encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 includes a channel decoder 251, another audio decoder 252, the audio encoder 253 provided by the embodiment of the present application, and a channel encoder 254. The another audio decoder 252 refers to an audio decoder other than the audio decoder. In the wireless device or the core network device 25, the signal entering the device is first decoded by the channel decoder 251, and then decoded by the another audio decoder 252, and then encoded by the audio encoder 253 provided by the embodiment of the present application, and finally encoded by the channel encoder 254 before being transmitted out. The another audio decoder 252 is used to decode the code stream decoded by the channel decoder 251.
[0098] As shown in FIG. 2c, the audio decoder provided by the embodiments of the present application is applied to a wireless device or a core network device. The wireless device or the core network device 25 comprises a channel decoder 251, the audio decoder 255 provided by the embodiments of the present application, another audio encoder 256, and a channel encoder 254, wherein the another audio encoder 256 refers to an audio encoder other than the audio encoder. In the wireless device or the core network device 25, the signal entering the device is firstly channel-decoded by the channel decoder 251, then the received audio encoded code stream is decoded by the audio decoder 255, then the audio is encoded by the another audio encoder 256, and finally the audio signal is channel-encoded by the channel encoder 254, and the channel-encoded audio signal is transmitted out. In the wireless device or the core network device, if transcoding is needed, corresponding audio encoding processing is needed. The wireless device refers to a radio frequency related device in communication, and the core network device refers to a core network related device in communication.
[0099] In some embodiments of the present application, the audio encoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices, for example, the audio encoding apparatus can be a multi-channel encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices, for example, the audio decoding apparatus can be a multi-channel decoder of the terminal device, the wireless device, or the core network device.
[0100] As shown in FIG. 3a, a schematic diagram of the multi-channel encoder and the multi-channel decoder provided by the embodiment of the present application applied to terminal devices is shown, and each terminal device can include a multi-channel encoder, a channel encoder, a multi-channel decoder and a channel decoder. The multi-channel encoder can perform the audio encoding method provided by the embodiment of the present application, and the multi-channel decoder can perform the audio decoding method provided by the embodiment of the present application. Specifically, the channel encoder is configured to perform channel encoding on the multi-channel signal, and the channel decoder is configured to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 can include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303 and a first channel decoder 304. The second terminal device 31 can include a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313 and a second channel encoder 314. The first terminal device 30 is connected to a first network communication device 32 via a wireless or wired connection, the first network communication device 32 and a second network communication device 33 are connected via a digital channel, and the second terminal device 31 is connected to the second network communication device 33 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device. In the audio communication, the terminal device as a sending end performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and transmits the signal in the digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a multi-channel signal encoding stream, and then performs multi-channel decoding to restore the multi-channel signal, and the terminal device as the receiving end plays back the multi-channel signal.
[0101] As shown in FIG. 3b, a schematic diagram of the multi-channel encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, another audio decoder 352, a multi-channel encoder 353 and a channel encoder 354. The above-mentioned wireless device or the core network device 35 is similar to the wireless device or the core network device 35 shown in FIG. 2b, and thus will not be described herein.
[0102] As shown in FIG. 3c, a schematic diagram of the multi-channel decoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, a multi-channel decoder 355, another audio encoder 356 and a channel encoder 354. The above-mentioned wireless device or the core network device 35 is similar to the wireless device or the core network device 35 shown in FIG. 2c, and thus will not be described herein.
[0103] The audio encoding process can be part of a multi-channel encoder, and the audio decoding process can be part of a multi-channel decoder. For example, multi-channel encoding of a captured multi-channel signal can include processing the captured multi-channel signal to obtain an audio signal, and encoding the obtained audio signal according to the method provided in the embodiments of the present application. The decoding end decodes the multi-channel signal encoding stream to obtain an audio signal, and recovers the multi-channel signal after upmix processing. Therefore, the embodiments of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In a wireless device or a core network device, if transcoding is required, corresponding multi-channel encoding processing needs to be performed.
[0104] The embodiments of the present application involve two or more terminals that need to communicate, for example, two terminals that communicate, or multiple terminals that communicate. In the following embodiments, two terminals that communicate are taken as an example for description, for example, a calling terminal and a called terminal. The calling terminal can be the encoding end as described above, and the called terminal can be the decoding end. The called terminal can be the encoding end as described above, and the calling terminal can be the decoding end. When two terminals in a mobile network establish a communication, because the voice Codec types supported by the two terminals can be different, and the code rate, sampling rate, etc. supported by the same voice Codec can also be different, the voice Codec needs to be negotiated before the communication is established, to determine which voice Codec and which encoding and decoding configuration are used for the communication. The multiple codecs can include EVS, AMR-WB, and IVAS. The transmission rate can also be referred to as the encoding and decoding rate.
[0105] The negotiation is performed using the Session Description Protocol (SDP). The negotiation process includes negotiation of four important parameters: load type, sampling frequency, rate, and packet length. Specifically:
[0106] The load type is a value defined in a standard protocol.
[0107] The higher the sampling frequency, the more sampling points, and the closer the voice quality to the real voice signal.
[0108] The rate is related to the bandwidth occupancy rate. The higher the rate, the higher the bandwidth occupancy rate.
[0109] The packet length refers to the voice duration contained in each voice packet. The larger the packet length, the larger the packet delay, but the stronger the anti-jitter capability and the higher the bandwidth utilization rate.
[0110] The parameters involved in the negotiation are: codec type and rate, sampling frequency and packet length, which are fixed for each codec type. In the SDP negotiation, the codec type is negotiated first, and then the rate. For AMR-WB and EVS, adaptive rate is supported, and the adaptive range supported is selected during negotiation.
[0111] The SDP template is as follows:
[0112] a=rtpmap:100 AMR / 8000 / / 100 is the load type, AMR represents the codec type, and 8000 represents the sampling frequency
[0113] a=fmtp:100 mode-set=0,2,4,7;mode-change-neighbor=1;mode-change-period=2 / / mode-set represents the rate set, the rate adjustment period and the adjustment mode, etc.
[0114] a=ptime:20 / / packet length.
[0115] First, a method for negotiating audio encoding / decoding provided by an embodiment of the present application is introduced. The method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in the embodiment of the present application can refer to audio encoding, or audio decoding, or audio encoding and decoding. Hereinafter, the audio encoding / decoding refers to the above-mentioned encoding and decoding processes.
[0116] As shown in FIG. 4, the method for negotiating audio encoding / decoding mainly includes the following steps:
[0117] 401. The second terminal device obtains the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device at least include the battery parameters of the second terminal device.
[0118] In the embodiment of the present application, the second terminal device determines the hardware parameters of the second terminal device, and the hardware parameters of the second terminal device can include hardware information of various terminal devices. The types of hardware and hardware configuration parameters involved in the embodiment of the present application are not limited.
[0119] The hardware parameters of the first terminal device include at least one of the following: parameters of an audio acquisition module of the first terminal device, and parameters of an audio playback module of the first terminal device.
[0120] The audio acquisition module can be a microphone, and the audio acquisition module can be an audio acquisition module built in the first terminal device or an audio acquisition module connected to the first terminal device, which is not limited herein.
[0121] The audio playback module can be a loudspeaker, and the audio playback module can be an audio playback module built in the first terminal device or an audio playback module connected to the first terminal device, which is not limited herein.
[0122] The battery parameter of the first terminal device, so that the first terminal device can perform audio codec negotiation according to the battery hardware of the first terminal device, for example, the battery parameter can include the capacity, the power, the temperature, the working mode of the battery, and the like.
[0123] 402、The second terminal device determines the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0124] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, so that the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device, for example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.
[0125] The audio encoding / decoding capability of the second terminal device refers to the capability of the second terminal device itself, for example, the audio encoding / decoding capability can include the computing power, the memory, and the like of the second terminal device.
[0126] For example, the audio encoding / decoding configuration information can be a codec mode configuration table.
[0127] 403、The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0128] The second terminal device can send the audio encoding / decoding configuration information of the second terminal device to the first terminal device, so that the first terminal device can perform audio encoding / decoding negotiation.
[0129] It is not limited that the first terminal device can also send the audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device can perform audio encoding / decoding negotiation.
[0130] 411、The first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0131] 412、the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to hardware parameters of the first terminal device, the hardware parameters of the first terminal device including at least one of the following: parameters of an audio acquisition module of the first terminal device, parameters of an audio playback module of the first terminal device.
[0132] In the embodiment of the application, the first terminal device determines the hardware parameters of the first terminal device, which can include hardware information of various terminal devices. The types of hardware and hardware configuration parameters involved in the embodiment of the application are not limited.
[0133] The hardware parameters of the first terminal device include at least one of the following: parameters of an audio acquisition module of the first terminal device, parameters of an audio playback module of the first terminal device.
[0134] 413、the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0135] The first terminal device can obtain the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device respectively, and can perform encoding / decoding negotiation to obtain the encoding / decoding negotiation result of the first terminal device.
[0136] The encoding / decoding negotiation result includes: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0137] 414、the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0138] 404、the first terminal device receives the encoding / decoding negotiation result from the second terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0139] In the embodiment of the application, the first terminal device and the second terminal device can perform static audio negotiation according to their respective hardware parameters. For example, negotiation is performed based on the hardware capability, processor and bottom chip of the terminal device, or based on the microphone and memory of the terminal device.
[0140] In some embodiments of the present application, the first coding / decoding parameter includes a first coding / decoding mode performed by the first terminal device when coding the audio signal to be processed and a first coding / decoding rate corresponding to the first coding / decoding mode.
[0141] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0142] When the hardware parameter of the first terminal device meets a preset first hardware configuration condition, the first audio coding / decoding configuration information is determined according to the hardware parameter of the first terminal device;
[0143] When the hardware parameter of the first terminal device does not meet the first hardware configuration condition, the second audio coding / decoding configuration information is determined according to the hardware parameter of the first terminal device;
[0144] The audio coding / decoding capability corresponding to the first audio coding / decoding configuration information is higher than the audio coding / decoding capability corresponding to the second audio coding / decoding configuration information.
[0145] In the embodiments of the present application, different hardware parameters of the first terminal device indicate different audio coding / decoding capabilities, thereby completing the audio coding / decoding negotiation based on the hardware parameter.
[0146] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0147] The audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device and the audio coding / decoding capability of the first terminal device.
[0148] In the embodiments of the present application, when the first terminal device determines the audio coding / decoding configuration information, the hardware parameter of the first terminal device is used, and the audio coding / decoding capability of the first terminal device is also used, so that the audio coding / decoding configuration information of the first terminal device can be more accurately determined.
[0149] In some embodiments of the present application, the hardware parameter further includes at least one of the following: a parameter of an audio coding / decoding chip of the first terminal device, a parameter of an audio coding / decoding chip of the first terminal device, a battery parameter of the second terminal device, and a parameter of an off-chip memory other than the audio coding / decoding chip in the second terminal device.
[0150] The audio coding / decoding chip of the first terminal device refers to a chip for coding and decoding an audio signal.
[0151] The off-chip memory of the first terminal device other than the audio codec chip refers to other memory of the first terminal device other than the audio codec chip, which can be the off-chip memory of the first terminal device, and the off-chip memory can also be a hardware parameter of the first terminal device.
[0152] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0153] The battery endurance information of the first terminal device is obtained according to the battery parameter of the first terminal device.
[0154] The audio encoding / decoding configuration information of the first terminal device is updated according to the battery endurance information, to obtain the updated audio encoding / decoding configuration information of the first terminal device.
[0155] In the embodiments of the present application, different battery parameters of the first terminal device indicate different battery endurance information, for example, when the battery endurance information indicates more power, stronger audio encoding / decoding capability can be used, and when the battery endurance information indicates that the power decreases, the audio encoding / decoding capability can be reduced, thereby completing the audio encoding / decoding negotiation based on the hardware parameter.
[0156] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0157] The audio signal type and the collection rate supported by the audio collection module of the first terminal device are determined according to the parameter of the audio collection module of the first terminal device.
[0158] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the collection rate supported by the audio collection module of the first terminal device.
[0159] In the embodiments of the present application, different audio collection modules of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, thereby completing the audio encoding / decoding negotiation based on the hardware parameter.
[0160] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0161] The audio signal type and the code rate supported by the audio codec chip of the first terminal device are determined according to the parameter of the audio codec chip of the first terminal device.
[0162] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the code rate supported by the audio codec chip of the first terminal device.
[0163] In the embodiments of the present application, different audio codec chips of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, so as to complete the audio encoding / decoding negotiation based on the hardware parameters.
[0164] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0165] The audio signal types and code rates supported by the off-chip memory of the first terminal device are determined according to the parameters of the off-chip memory of the first terminal device except the audio codec chip.
[0166] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal types and code rates supported by the off-chip memory of the first terminal device.
[0167] In the embodiments of the present application, different off-chip memories of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, so as to complete the audio encoding / decoding negotiation based on the hardware parameters.
[0168] In some embodiments of the present application, the method further includes:
[0169] The audio service standard types supported by the audio codec chip of the second terminal device are obtained.
[0170] When the audio service standard types supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device are both the preset first audio service standard type, the step of obtaining the audio encoding / decoding configuration information of the second terminal device is triggered.
[0171] The audio service standard types can include at least one of EVS / IVAS / AMR, etc. The audio service standard types are negotiated before the audio negotiation, so that more accurate audio negotiation can be achieved.
[0172] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding mode supported by the second terminal device, and the rate set corresponding to the encoding / decoding mode supported by the second terminal device.
[0173] The audio encoding / decoding configuration information of the first terminal device includes: the encoding / decoding mode supported by the first terminal device, and the rate set corresponding to the encoding / decoding mode supported by the first terminal device.
[0174] For example, the encoding / decoding mode can include MONO / STEREO / FOA / MASA, and the values of the rates are not equal, and the specific values depend on the corresponding configuration information.
[0175] In some embodiments of the present application, determining the coding / decoding negotiation result of the first terminal device according to the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device comprises:
[0176] Determining the audio coding / decoding negotiation order according to the audio coding / decoding negotiation priority policy of the first terminal device, the audio coding / decoding negotiation order comprises: performing the audio coding negotiation before the audio decoding negotiation, or performing the audio decoding negotiation before the audio coding negotiation.
[0177] The audio coding negotiation comprises: determining the coding negotiation result of the first terminal device according to the audio coding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device.
[0178] The audio decoding negotiation comprises: determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio coding capability of the second terminal device.
[0179] The order of the audio coding negotiation and the audio decoding negotiation in the embodiments of the present application is not limited, and a more flexible audio negotiation mode can be realized.
[0180] In some embodiments of the present application, determining the coding negotiation result of the first terminal device according to the audio coding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device comprises:
[0181] According to the first audio mode supported by the highest coding mode of the first terminal device, and the decoding mode supported by the second terminal device includes the first audio mode, determining the coding negotiation result of the first terminal device includes the first audio mode.
[0182] Determining the first code rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device.
[0183] Determining the coding negotiation result of the first terminal device includes the maximum code rate in the first code rate intersection.
[0184] In the embodiments of the present application, the coding negotiation result includes the maximum code rate in the first code rate intersection, according to the principle of maximizing experience, the coding code rate with higher experience effect of the opposite terminal is preferentially selected, and the call quality is improved.
[0185] In some embodiments of the present application, determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio coding capability of the second terminal device comprises:
[0186] The highest decoding mode supported by the first terminal device is the second audio mode, and the encoding mode supported by the second terminal device includes the second audio mode, and the decoding negotiation result of the first terminal device includes the second audio mode;
[0187] The second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device are determined to have a second code rate intersection;
[0188] The decoding negotiation result of the first terminal device includes the maximum code rate in the second code rate intersection.
[0189] In the embodiments of the present application, the decoding negotiation result includes the maximum code rate in the second code rate intersection, and according to the principle of maximizing experience, the other terminal is preferentially selected to have a higher encoding code rate, thereby improving the call quality.
[0190] In some embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0191] In the embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device, thereby fully utilizing the maximum code rate of the first terminal device, maximizing the use of the audio encoding / decoding capability of the first terminal device, and improving the call quality.
[0192] In some embodiments of the present application, the first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), high-order ambisonic (HOA), or first-order ambisonic (FOA).
[0193] The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0194] In some embodiments of the present application, the method further includes:
[0195] The first terminal device sends an audio negotiation request to the second terminal device.
[0196] The first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains the hardware parameters of the second terminal device according to the trigger of the audio negotiation request.
[0197] In some embodiments of the present application, the audio negotiation request includes audio encoding / decoding configuration information of the first terminal device.
[0198] The capability negotiation request sent by the first terminal device to the second terminal device carries audio encoding / decoding configuration information of the first terminal device, and the second terminal device can preliminarily select, according to the received audio encoding / decoding configuration information of the first terminal device, to obtain one or more audio encoding / decoding configuration information of the second terminal device after preliminary screening, and then the first terminal device performs audio negotiation, thereby further improving the efficiency of audio negotiation.
[0199] Further comprising: the capability negotiation request sent by the first terminal device to the second terminal device, which can also not include the audio encoding / decoding configuration information of the first terminal device, which is not limited here.
[0200] The foregoing embodiments illustrate the method performed by the first terminal device, and the method performed by the second terminal device is described as follows. The configuration mode of the hardware parameters of the second terminal device is similar to the configuration mode of the hardware parameters of the first terminal device, and the specific process of determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device is not described in detail.
[0201] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device comprises:
[0202] Determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0203] In some embodiments of the present application, the hardware parameters further comprise at least one of the following: parameters of an audio encoding / decoding chip of the second terminal device, parameters of a battery of the second terminal device, and parameters of an off-chip memory of the second terminal device other than the audio encoding / decoding chip.
[0204] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0205] Obtaining battery endurance information of the second terminal device according to the battery parameters of the second terminal device;
[0206] Updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0207] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0208] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the audio collection module of the second terminal device;
[0209] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the audio collection module of the second terminal device;
[0210] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0211] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0212] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the audio encoding / decoding chip of the second terminal device;
[0213] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the audio encoding / decoding chip of the second terminal device;
[0214] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0215] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0216] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the off-chip memory of the second terminal device except the audio encoding / decoding chip;
[0217] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the off-chip memory of the second terminal device;
[0218] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0219] In some embodiments of the present application, the method further comprises:
[0220] The audio service standard type supported by the audio encoding / decoding chip of the first terminal device is acquired;
[0221] When the audio service standard type supported by the audio codec chip of the first terminal device and the audio codec chip of the second terminal device is the preset first audio service standard type, the foregoing step of obtaining the hardware parameter of the second terminal device is triggered to be performed.
[0222] In some embodiments of the present application, the audio codec configuration information of the second terminal device includes: a codec mode supported by the second terminal device, and a rate set corresponding to the codec mode supported by the second terminal device.
[0223] In some embodiments of the present application, the first codec parameter includes at least two of the following: a mono (MONO), a stereo (STEREO), a metadata assisted spatial audio (MASA), a higher order ambisonics (HOA), or a first order ambisonics (FOA).
[0224] The rate set corresponding to the codec mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0225] In some embodiments of the present application, the method further includes:
[0226] Receiving an audio negotiation request from the first terminal device.
[0227] In some embodiments of the present application, the audio negotiation request includes: the audio codec configuration information of the first terminal device.
[0228] The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device.
[0229] The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device.
[0230] Next, an actual application scenario example is used for illustration.
[0231] The method provided by the embodiments of the present application is applicable to IVAS encoding and decoding, and different signal types and corresponding rates can be used, and further negotiation is required. The embodiments of the present application provide a negotiation method for IVAS codec mode and rate.
[0232] Each IVAS-enabled terminal has its own IVAS codec configuration table, which describes all the IVAS encoding-decoding combinations that can be supported by the terminal under its hardware constraints. By passing the codec configuration table in the negotiation process, the two parties of the call can negotiate the call configuration more flexibly, such as the type of encoded signal, the code rate, etc. In addition, it is also possible to support the two parties of the call to use different signal types and different code rates for the call. The present application can provide the best user experience for both parties as much as possible under the condition of meeting the hardware constraints of both parties.
[0233] Each IVAS-enabled mobile phone will maintain two tables according to the capabilities of the mobile phone (computing power, memory, etc.). One table is the IVAS codec mode configuration table, as shown in Table 1 below:
[0234] The IVAS codec mode configuration table can represent all possible combinations of IVAS encoding modes and decoding modes that the mobile phone can support, including the signal type and code rate of encoding or decoding. As an example in Table 1 above, the mobile phone supports three types of decoded signal types, which are MONO, STEREO and FOA, among which the maximum code rate supported by MONO decoding is 9.6kbps, the maximum code rate supported by STEREO decoding is 32kbps, and the maximum code rate supported by FOA decoding is 48kbps. When the mobile phone decodes the MONO signal, it can support three types of encoded signal types, which are MONO, STEREO and FOA, and the maximum code rates supported by encoding these three signals are 9.6kbps, 32kbps and 48kbps respectively. When the mobile phone decodes the STEREO signal, it can still support three types of encoded signal types, which are MONO, STEREO and FOA, but since the decoding overhead of STEREO is larger than that of MONO, the computing power left for encoding will decrease accordingly, and therefore the encoding code rate of FOA cannot be supported to 48kbps, but only to 24kbps at most (the higher the code rate, the larger the overhead). When the mobile phone decodes the FOA signal, since the decoding overhead of FOA is even larger, the remaining computing power cannot support the encoding of FOA, and therefore the mobile phone can only support two types of encoded signal types, which are MONO and STEREO, and the maximum code rates supported are 9.6kbps and 32kbps respectively.
[0235] As can be seen from the above example, the combination of encoding-decoding modes is limited by the capabilities of the terminal / chip, and if the encoding overhead is large, the decoding overhead will decrease, and vice versa.
[0236] The other table is the codec code rate set table, as shown in Table 2 below:
[0237] The coding rate set table can represent all the coding rates supported by IVAS for each signal. In combination with the maximum rate supported by each coding / decoding signal in the coding mode configuration table and the coding rate set table, the rate range supported by each coding / decoding signal in the coding mode configuration table can be obtained.
[0238] Each mobile phone has a default IVAS coding and decoding configuration table when it is shipped. Due to factors such as the chip computing power of the mobile phone, the IVAS coding and decoding configuration table of different mobile phone platforms may be different.
[0239] Since the type of IVAS signal that the mobile phone can encode is also related to some external factors, for example, only when the mobile phone has FOA or HOA acquisition equipment can it support the encoding of FOA / HOA, for example, only when the voice service contains multi-channel or object audio / sound effect signals can it support the encoding of multi-channel or objects, for example, only when the mobile phone supports the complete MASA solution can it support the encoding of MASA signals, for example, when the mobile phone is low in power, it may limit the coding and decoding of some high complexity signals, and the mobile phone needs to be able to dynamically update its own IVAS coding and decoding configuration table to reflect the current actual ability of the mobile phone.
[0240] As shown in FIG. 5, taking the audio negotiation process of terminal UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) as an example, the following is described by way of example:
[0241] Before A and B voice calls, if IVAS voice coding and decoding is used for communication, media negotiation is needed, which includes:
[0242] 1. Audio coding and decoding negotiation (EVS / IVAS / AMR).
[0243] 2. When the audio coding and decoding negotiation is IVAS, the coding and decoding mode (MONO / STEREO / FOA / MASA) and the rate (each coding and decoding mode has a corresponding rate range) need to be further negotiated. The coding mode and the decoding mode can be negotiated to be different.
[0244] Negotiation process:
[0245] 1. When UE-A initiates a call INVITE, if it supports IVAS voice coding and decoding, it sends its own supported IVAS coding and decoding configuration table to B in the IVAS SDP information, including the list of A-supported coding mode and decoding mode combinations, the maximum rate supported by each mode, and the rate set supported by each coding mode.
[0246] 2. UE-B processes as follows according to its own ability:
[0247] 1) First, codec negotiation is carried out, such as supporting IVAS, and preliminary negotiation is IVAS;
[0248] 2) Continue to negotiate the codec mode and rate. According to the local strategy, it is decided whether to negotiate according to the encoding ability or the decoding ability, taking the priority of the encoding ability negotiation as an example:
[0249] ① According to the highest encoding mode supported by UE-B and whether UE-A supports the decoding mode, the encoding mode of UE-B is determined;
[0250] ② According to the intersection of the code rate supported by the calling and called parties for the encoding mode, the encoding code rate of UE-B after negotiation is determined. If there is no intersection, the encoding mode of UE-B needs to be re-negotiated;
[0251] ③ According to the decoding mode corresponding to the encoding mode list supported by UE-B, combined with whether UE-A supports the encoding mode, the decoding mode of UE-B after negotiation is determined;
[0252] ④ According to the intersection of the code rate of the calling and called parties for the decoding mode, the decoding code rate of UE-B after negotiation is determined.
[0253] If any of the codec modes negotiated by IVAS fails, IVAS will not be finally used.
[0254] As shown in FIGS. 6 and 7, the SDP example adds the following parameters in the offer a line of IVAS:
[0255] Code rate set:
[0256] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.
[0257] [mono,5.9 / 7.2 / 9.6] represents that the code rate set supported by mono is 5.9, 7.2 and 9.6.
[0258] Decoding mode set:
[0259] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.
[0260] The list represents that 3 sets are supported, each set corresponds to the following encoding mode set one by one, set 1 corresponds to the decoding mode [mono / 9.6], which represents that mono is supported and the highest code rate is 9.6.
[0261] Encoding mode set:
[0262] enc-mode-list = {[mono / 9.6, stereo / 32, foa / 48], [mono / 9.6, stereo / 32, foa / 24], [mono / 9.6, stereo / 32]}.
[0263] The list indicates that 3 sets are supported, which are one-to-one corresponding to the above decoding mode set. The encoding mode corresponding to set 1 is = { [mono / 9.6, stereo / 32, foa / 48], which indicates that the decoding mode is mono / stereo / foa from low to high, and the highest code rate is 9.6 / 32 / 48.
[0264] The following parameters are added in the answer a line of IVAS:
[0265] The negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]}.
[0266] The negotiated encoding mode and rate set: enc-mode = {[foa, 24 / 32]}.
[0267] An example of the negotiation process is shown in FIG. 8.
[0268] 1. Mobile phone A sends its IVAS encoding and decoding configuration table to mobile phone B, including the encoding and decoding modes supported by A and the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0269] 2. Mobile phone B decides according to the local strategy whether to negotiate according to its encoding ability or according to its decoding ability. Taking the encoding ability negotiation as an example:
[0270] ① According to the highest encoding mode supported by UE-B, which is FOA, and the decoding mode of UE-A containing FOA, it is determined that the encoding mode of UE-B after negotiation is FOA;
[0271] ② Determine the intersection of the code rates supported by the FOA mode of the main and called parties, and determine that the encoding code rate of UE-B after negotiation is 24 and 32;
[0272] ③ According to the decoding mode corresponding to the encoding mode column supported by UE-B, which is MONO, and the encoding mode of UE-A containing MONO, it is determined that the decoding mode of UE-B after negotiation is MONO;
[0273] ④ Determine the intersection of the code rates supported by the MONO mode of the main and called parties, and determine that the decoding code rate of UE-B after negotiation is 5.9, 7.2 and 9.6.
[0274] Another example of the negotiation process is shown in FIG. 8.
[0275] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A and the maximum rate combinations supported by each mode, and the rate set supported by each coding mode.
[0276] 2. Mobile phone B decides according to the local policy whether to negotiate according to its coding capability or according to its decoding capability, taking the coding capability negotiation as an example:
[0277] ① According to the highest decoding mode supported by UE-B, which is FOA, and the decoding mode of UE-A containing FOA, it is determined that the decoding mode of UE-B after negotiation is FOA;
[0278] ② The intersection of the code rates of the FOA mode supported by the main and called parties is determined, and it is determined that the coding code rate of UE-B after negotiation is 24 and 32, 48;
[0279] ③ According to the highest coding mode supported by the decoding mode list of UE-B, which is STEREO, and the decoding mode of UE-A containing STEREO, it is determined that the coding mode of UE-B after negotiation is STEREO;
[0280] ④ The intersection of the code rates of the STEREO mode supported by the main and called parties is determined, and it is determined that the decoding code rate of UE-B after negotiation is 13.2 and 24.
[0281] Based on the foregoing example, for real-time communication audio Codec with numerous coding modes and coding rates and large span of coding and decoding complexity, each terminal saves a copy of the codec configuration table for it. The configuration table describes the various combinations of coding and decoding modes / rates that the terminal can support. When two or more terminals establish a call to negotiate the Codec, they select their respective suitable coding configurations for coding according to their respective codec configuration tables.
[0282] Each terminal has a factory default codec configuration table. When establishing a call to negotiate the Codec, the terminal can update its codec configuration table according to other factors, and negotiate the Codec according to the updated codec configuration table. The updated codec configuration table does not overwrite the default factory codec configuration table.
[0283] In the codec configuration table, the supported decoding modes / code rates and / or coding modes / code rates are sorted by experience priority.
[0284] During Codec negotiation, one terminal sends its complete codec configuration table or part of the codec configuration table to the other terminal for negotiation, and the terminal makes a decision according to the entire or partial codec configuration tables of at least two terminals. Here, part of the coding mode and rate refers to a subset.
[0285] The decision of the negotiation can be based on the principle of maximizing experience, or other principles, such as a preset decision principle of the terminal making the decision of the negotiation, for example, preferentially selecting the other terminal to perform coding with a higher experience effect.
[0286] Next, an audio negotiation process for terminal power is illustrated by way of example.
[0287] When the terminal power is 100%, the configuration table of the terminal is an initial configuration table, for example, as shown in Table 3 below:
[0288] When the terminal power decreases, the configuration table needs to be adjusted to reduce the configuration in order to ensure the operation of other functions of the terminal, for example, when the power is only 10%, the configuration table is adjusted to be as shown in Table 4 below:
[0289] When the terminal only uses configuration 1 (default hardware configuration / factory configuration), the configuration table of the terminal is an initial configuration table, as shown in Table 3 above.
[0290] When the terminal uses configuration 2 (external FOA microphone), the configuration table is modified for some combinations of Table 3.
[0291] When the terminal uses configuration 3 (external HOA microphone), the configuration table is modified to replace combination 3 in Table 3 with HOA and the corresponding code rate.
[0292] When the terminal uses configuration 4 (external microphone), the terminal side device can convert the microphone signal into a Codec supported FOA / HOA / MASA signal, and the configuration table includes combination 3 as FOA and the corresponding code rate, combination 4 as HOA and the corresponding code rate, and combination 5 as MASA and the corresponding code rate.
[0293] It should be noted that, for each of the foregoing method embodiments, in order to simply describe, each is described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0294] In order to better implement the above-mentioned scheme of the embodiments of the present application, the following provides a related device for implementing the above-mentioned scheme.
[0295] Please refer to FIG. 9, a first terminal device 900 provided by the embodiments of the present application can include an acquisition module 901, a determination module 902, a negotiation module 903, and a sending module 904, wherein,
[0296] obtaining an audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device;
[0297] determining an audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device including at least one of a parameter of an audio acquisition module of the first terminal device and a parameter of an audio playback module of the first terminal device;
[0298] determining an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including a first encoding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device;
[0299] sending the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0300] As can be known from the foregoing embodiments, first, the first terminal device obtains an audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; then, the first terminal device determines an audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device including at least one of a parameter of an audio acquisition module of the first terminal device and a parameter of an audio playback module of the first terminal device; next, the first terminal device determines an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including a first encoding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiments of the present application, both of the two terminals for audio negotiation have audio encoding / decoding configuration information, which describes the encoding / decoding capability that can be supported by the terminal under the constraint of the hardware parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiments of the present application, the optimal user experience can be provided for the two parties as much as possible under the constraint of the hardware parameter of the two parties of the call.
[0301] Referring to FIG. 10, a second terminal device 1000 provided by an embodiment of the present application can include an obtaining module 1001, a determining module 1002, a sending module 1003, and a receiving module 1004, wherein
[0302] The obtaining module is configured to obtain a hardware parameter of the second terminal device, and the hardware parameter of the second terminal device includes at least one of a parameter of an audio acquisition module of the second terminal device and a parameter of an audio playback module of the second terminal device.
[0303] The determining module is configured to determine audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate an audio encoding / decoding capability of the second terminal device.
[0304] The sending module is configured to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0305] The receiving module is configured to receive a coding / decoding negotiation result from the first terminal device, and the coding / decoding negotiation result includes a first coding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, and the first coding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0306] It should be noted that the information interaction and execution process between the modules / units of the above apparatus are based on the same consideration as the method embodiments of the present application, and the technical effects brought by the same are the same as those of the method embodiments of the present application. For details, refer to the description in the foregoing method embodiments of the present application, which will not be repeated here.
[0307] The embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a program, and the program performs part or all of the steps recorded in the foregoing method embodiments.
[0308] Next, another first terminal device provided by an embodiment of the present application is introduced. Referring to FIG. 11, the first terminal device 1100 includes a receiver 1101, a transmitter 1102, a processor 1103, and a memory 1104 (wherein the number of the processor 1103 in the first terminal device 1100 can be one or more, and one processor is taken as an example in FIG. 11).
[0309] The receiver 1101, the transmitter 1102, the processor 1103, and the memory 1104 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 11.
[0310] The memory 1104 can include read-only memory and random access memory, and provide instructions and data to the processor 1103. A portion of the memory 1104 can also include non-volatile random access memory (NVRAM). The memory 1104 stores operating systems and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0311] The processor 1103 controls the operation of the first terminal device, and the processor 1103 can also be referred to as a central processing unit (CPU). In specific applications, various components of the first terminal device are coupled together through a bus system, which can include a data bus, a power bus, a control bus, a state signal bus, etc. However, for the sake of clarity, various buses are referred to as a bus system in the figure.
[0312] The method disclosed in the above embodiments of the present application can be applied in the processor 1103 or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1103 or an instruction in the form of software. The processor 1103 described above can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 1104, and the processor 1103 reads the information in the memory 1104 and combines the hardware to complete the steps of the above method.
[0313] The receiver 1101 can be configured to receive inputted digital or character information, and generate signal input related to the relevant settings and function control of the first terminal device. The transmitter 1102 can include a display device such as a display screen, and the transmitter 1102 can be configured to output digital or character information through an external interface.
[0314] In the embodiments of the present application, the processor 1103 is configured to execute the method performed by the first terminal device shown in FIG. 4.
[0315] Next, another second terminal device provided by the embodiments of the present application is introduced. Referring to FIG. 12, the second terminal device 1200 includes:
[0316] The receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 (wherein the number of the processor 1203 in the second terminal device 1200 can be one or more, and one processor is taken as an example in FIG. 12). In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 12.
[0317] The memory 1204 can include read-only memory and random access memory, and provide the processor 1203 with instructions and data. A portion of the memory 1204 can also include NVRAM. The memory 1204 stores an operating system and operation instructions, executable modules or data structures, or subsets thereof, or expanded sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0318] The processor 1203 controls the operation of the second terminal device, and the processor 1203 can also be referred to as a CPU. In specific applications, various components of the second terminal device are coupled together through a bus system, wherein the bus system can include a data bus in addition to a power bus, a control bus and a status signal bus, etc. However, in order to clearly illustrate, various buses are referred to as a bus system in the figure.
[0319] The method disclosed in the embodiments of the present application can be applied to the processor 1203 or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the above method can be completed by an integrated logic circuit or a software form of the instructions in the processor 1203. The processor 1203 can be a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed by the processor 1203. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code executed by the processor, or a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an electrically programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, or a removable disk. The storage medium is located in the storage 1204, and the processor 1203 reads information in the storage 1204 and combines the hardware to complete the steps of the above method.
[0320] In the embodiments of the present application, the processor 1203 is configured to execute the method performed by the second terminal device shown in Fig. 7.
[0321] In another possible design, when the first terminal device or the second terminal device is a chip in the terminal, the chip includes a processing unit and a communication unit. The processing unit can be a processor, and the communication unit can be an input / output interface, a pin, a circuit, or the like. The processing unit can execute computer-executable instructions stored in a storage unit, so that the chip in the terminal executes the method of any one of the first aspect. Alternatively, the storage unit is a storage unit in the chip, such as a register, a cache, or the like. The storage unit can also be a storage unit in the terminal and located outside the chip, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or the like.
[0322] The processor mentioned in any of the above can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling execution of programs of the method of the first aspect or the second aspect.
[0323] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate units can or can not be physically separate, and the units displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0324] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and the necessary general hardware, and of course can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and the specific hardware structure for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of software products, which are stored in readable storage media, such as computer floppy disks, U disks, mobile hard disks, ROM, RAM, magnetic or optical disks, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0325] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.
[0326] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A method of negotiation of audio encoding / decoding, characterized in that, The method is applied to a first terminal device, and the method comprises: obtaining audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; determining audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device comprising at least one of a parameter of an audio acquisition module of the first terminal device and a parameter of an audio playback module of the first terminal device; determining an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result comprising a first encoding / decoding parameter of the first terminal device for encoding / decoding a to-be-processed audio signal, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; sending the encoding / decoding negotiation result of the first terminal device to the second terminal device.
2. The method of claim 1, wherein, The first encoding / decoding parameter comprises a first encoding / decoding mode executed by the first terminal device when encoding / decoding the to-be-processed audio signal and a first encoding / decoding rate corresponding to the first encoding / decoding mode.
3. The method of claim 1, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device comprises: when the hardware parameter of the first terminal device satisfies a preset first hardware configuration condition, determining first audio encoding / decoding configuration information according to the hardware parameter of the first terminal device; when the battery parameter of the first terminal device does not satisfy the first hardware configuration condition, determining second audio encoding / decoding configuration information according to the hardware parameter of the first terminal device; wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
4. The method according to claim 1 or 2, characterized in that, The determining of the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device comprises: determining the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device and an audio encoding / decoding capability of the first terminal device.
5. The method according to any one of claims 1 to 4, characterized in that, The hardware parameter further comprises at least one of a parameter of an audio encoding / decoding chip of the first terminal device, a battery parameter of the first terminal device, and a parameter of a memory other than the audio encoding / decoding chip in the first terminal device.
6. The method of claim 5, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device comprises: obtaining battery endurance information of the first terminal device according to the battery parameter of the first terminal device; updating the audio encoding / decoding configuration information of the first terminal device according to the battery endurance information to obtain updated audio encoding / decoding configuration information of the first terminal device.
7. The method of claim 5, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device comprises: determining the audio signal type and the sampling rate supported by the audio acquisition module of the first terminal device according to the parameters of the audio acquisition module of the first terminal device; determining the audio codec configuration information of the first terminal device according to the audio signal type and the sampling rate supported by the audio acquisition module of the first terminal device.
8. The method of claim 5, wherein, The determining the audio codec configuration information of the first terminal device according to the hardware parameters of the first terminal device comprises: determining the audio signal type and the sampling rate supported by the audio codec chip of the first terminal device according to the parameters of the audio codec chip of the first terminal device; determining the audio codec configuration information of the first terminal device according to the audio signal type and the sampling rate supported by the audio codec chip of the first terminal device.
9. The method of claim 5, wherein, The determining the audio codec configuration information of the first terminal device according to the hardware parameters of the first terminal device comprises: determining the audio signal type and the sampling rate supported by the off-chip memory of the first terminal device except the audio codec chip according to the parameters of the off-chip memory of the first terminal device; determining the audio codec configuration information of the first terminal device according to the audio signal type and the sampling rate supported by the off-chip memory of the first terminal device.
10. The method according to any one of claims 1 to 9, characterized in that, The method further comprises: obtaining the audio service standard type supported by the audio codec chip of the second terminal device; when the audio service standard type supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device is the preset first audio service standard type, triggering the execution of the aforementioned step of obtaining the audio codec configuration information of the second terminal device.
11. The method according to any one of claims 1 to 10, characterized in that, The audio codec configuration information of the second terminal device comprises the codec mode supported by the second terminal device and the rate set corresponding to the codec mode supported by the second terminal device. The audio codec configuration information of the first terminal device comprises the codec mode supported by the first terminal device and the rate set corresponding to the codec mode supported by the first terminal device.
12. The method according to any one of claims 1 to 11, characterized in that, The determining the codec negotiation result of the first terminal device according to the audio codec configuration information of the first terminal device and the audio codec configuration information of the second terminal device comprises: determining the audio codec negotiation order according to the audio codec negotiation priority strategy of the first terminal device, wherein the audio codec negotiation order comprises performing the audio encoding negotiation before the audio decoding negotiation or performing the audio decoding negotiation before the audio encoding negotiation; wherein the performing the audio encoding negotiation comprises determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio codec configuration information of the first terminal device and the audio decoding capability of the second terminal device; the performing the audio decoding negotiation comprises determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio codec configuration information of the first terminal device and the audio encoding capability of the second terminal device.
13. The method of claim 12, wherein, The audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device determine the encoding negotiation result of the first terminal device, including: The first audio mode is determined as the encoding negotiation result of the first terminal device according to that the highest encoding mode supported by the first terminal device is the first audio mode and the decoding mode supported by the second terminal device includes the first audio mode; A first code rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device is determined; The maximum code rate in the first code rate intersection is determined as the encoding negotiation result of the first terminal device.
14. The method of claim 12, wherein, The audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device determine the decoding negotiation result of the first terminal device, including: The second audio mode is determined as the decoding negotiation result of the first terminal device according to that the highest decoding mode supported by the first terminal device is the second audio mode and the encoding mode supported by the second terminal device includes the second audio mode; A second code rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device is determined; The maximum code rate in the second code rate intersection is determined as the decoding negotiation result of the first terminal device.
15. The method of claim 12, wherein, The sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode performed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
16. The method according to any one of claims 1 to 15, characterized in that, The first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), higher order ambisonics (HOA), or first order ambisonics (FOA). The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
17. The method of any one of claims 1 to 16, wherein, The method further includes: sending an audio negotiation request to the second terminal device.
18. The method of claim 17, wherein, The audio negotiation request includes the audio encoding / decoding configuration information of the first terminal device.
19. A terminal device, comprising: The terminal device is specifically a first terminal device, and the method includes: an obtaining module configured to obtain audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate audio encoding / decoding capability of the second terminal device; a determining module configured to determine the audio encoding / decoding configuration information of the first terminal device according to hardware parameters of the first terminal device, the hardware parameters of the first terminal device including at least one of the following: a parameter of an audio acquisition module of the first terminal device or a parameter of an audio playback module of the first terminal device. The negotiation module is configured to determine the coding / decoding negotiation result of the first terminal device according to the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device, wherein the coding / decoding negotiation result comprises first coding / decoding parameters of the first terminal device for coding / decoding audio signals to be processed, and the first coding / decoding parameters are adapted to the audio coding / decoding capability of the second terminal device. The sending module is configured to send the coding / decoding negotiation result of the first terminal device to the second terminal device.
20. A terminal device, comprising: The terminal device comprises at least one processor, which is coupled with a memory and reads and executes instructions in the memory to implement the method in any one of claims 1 to 18.
21. The terminal device of claim 20, wherein, The terminal device further comprises the memory.
22. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to carry out the method of any one of claims 1 to 18.
23. A computer readable storage medium comprising the coding / decoding negotiation result generated by the method of any one of claims 1 to 18.