Audio encoding / decoding negotiation method and terminal device
By obtaining the hardware parameters of the terminal device to determine the audio codec configuration information and conducting precise codec negotiation, the problems of poor audio codec negotiation effect and power waste in the existing technology are solved, and a better user experience and codec effect are achieved.
Patent Information
- Application Number
- PCT/CN2025/082516
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-14
- Publication Date
- 2025-11-27
AI Technical Summary
Existing audio codec negotiation methods result in poor audio codec negotiation performance when selecting a lower complexity level for a call, failing to meet the codec requirements of higher-level terminals and also causing power consumption waste.
By obtaining the hardware parameters of both parties' terminal devices, the audio codec configuration information of each party is determined, and precise codec negotiation is carried out to select codec parameters that are suitable for each party's capabilities, ensuring that the best user experience is provided under hardware constraints.
It achieves stable call connection while improving audio encoding and decoding performance, avoiding power waste, and providing a better user experience.
Smart Images

Figure CN2025082516_27112025_PF_FP_ABST
Abstract
Description
Method for negotiating audio encoding / decoding and terminal device
[0001] The present application claims priority from the Chinese patent application No. 202410361453.0 filed on March 25, 2024, and entitled "Method for negotiating audio encoding / decoding and terminal device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of audio encoding and decoding, and in particular to a method for negotiating audio encoding / decoding and a terminal device. BACKGROUND
[0003] Compared with a speech codec using a single channel, an immersive speech codec supports more types of audio signals, supports a larger range of code rates, and includes rendering characteristics in addition to codec characteristics. For example, in the ongoing audio codec standard of the 3rd Generation Partnership Project (3GPP), different audio codec standards support multiple types of signals, different types of signals support different numbers of channels and code rates, and the rendering characteristics of different types of signals support rendering of various types of signals to binaural and standard loudspeaker arrays. Due to supporting more signal channel numbers, a larger range of code rates, and additional rendering characteristics, the immersive speech codec has a very significant increase in maximum computational complexity and storage complexity compared to the traditional speech codec.
[0004] In order to support a speech codec with high complexity, in an existing method for negotiating audio encoding / decoding, the speech codec is classified according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result can only select a lower complexity level for the call to ensure that the call connection is stable. Although this method meets the speech requirements of the call connection, it has the problem of poor audio encoding / decoding negotiation effect. SUMMARY
[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a method for negotiating audio encoding / decoding, applied to a second terminal device, the method comprising:
[0007] obtaining a hardware parameter of the second terminal device, the hardware parameter of the second terminal device comprising at least one of a parameter of an audio acquisition module of the second terminal device and a parameter of an audio playback module of the second terminal device;
[0008] determining audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device;
[0009] sending the audio encoding / decoding configuration information of the second terminal device to the first terminal device;
[0010] receiving a result of the encoding / decoding negotiation from the first terminal device, the result of the encoding / decoding negotiation including first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0011] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device includes:
[0012] determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0013] In some embodiments of the present application, the hardware parameter of the second terminal device further includes at least one of the following: a parameter of an audio encoding / decoding chip of the second terminal device, a battery parameter of the second terminal device, and a parameter of an off-chip memory of the second terminal device other than the audio encoding / decoding chip.
[0014] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0015] obtaining battery endurance information of the second terminal device according to the battery parameter of the second terminal device;
[0016] updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0017] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0018] determining an audio signal type and a collection rate supported by an audio collection module of the second terminal device according to a parameter of the audio collection module of the second terminal device;
[0019] determining the audio encoding / decoding capability of the second terminal device according to the audio signal type and the audio acquisition rate supported by the audio acquisition module of the second terminal device;
[0020] determining the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability of the second terminal device.
[0021] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0022] determining the audio signal type and the audio acquisition rate supported by the audio acquisition module of the second terminal device according to the parameter of the audio codec chip of the second terminal device;
[0023] determining the audio encoding / decoding capability of the second terminal device according to the audio signal type and the audio acquisition rate supported by the audio codec chip of the second terminal device;
[0024] determining the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability of the second terminal device.
[0025] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0026] determining the audio signal type and the audio acquisition rate supported by the audio acquisition module of the second terminal device according to the parameter of the off-chip memory of the second terminal device other than the audio codec chip;
[0027] determining the audio encoding / decoding capability of the second terminal device according to the audio signal type and the audio acquisition rate supported by the off-chip memory of the second terminal device;
[0028] determining the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability of the second terminal device.
[0029] In some embodiments of the present application, the method further comprises:
[0030] acquiring the audio service standard type supported by the audio codec chip of the first terminal device;
[0031] when the audio service standard type supported by the audio codec chip of the first terminal device and the audio codec chip of the second terminal device is respectively the preset first audio service standard type, triggering the execution of the aforementioned step of acquiring the hardware parameter of the second terminal device.
[0032] In some embodiments of the present application, the audio coding / decoding configuration information of the second terminal device comprises: a coding mode supported by the second terminal device, and a rate set corresponding to the coding mode supported by the second terminal device.
[0033] In some embodiments of the present application, the first coding / decoding parameter comprises at least two of: a mono (MONO), a stereo (STEREO), a metadata assisted spatial audio (MASA), a higher order ambisonic (HOA), or a first order ambisonic (FOA).
[0034] The rate set corresponding to the coding mode comprises at least two of: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0035] In some embodiments of the present application, the method further comprises:
[0036] receiving an audio negotiation request from the first terminal device.
[0037] In some embodiments of the present application, the audio negotiation request comprises: audio coding / decoding configuration information of the first terminal device.
[0038] The method further comprises: screening the audio coding / decoding configuration information of the second terminal device according to the audio coding / decoding configuration information of the first terminal device to obtain candidate audio coding / decoding configuration information of the second terminal device.
[0039] The sending of the audio coding / decoding configuration information of the second terminal device to the first terminal device comprises: sending the candidate audio coding / decoding configuration information of the second terminal device to the first terminal device.
[0040] In a second aspect, embodiments of the present application provide a terminal device, specifically a second terminal device, and the method comprises:
[0041] The obtaining module is configured to obtain hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device comprise at least one of: a parameter of an audio acquisition module of the second terminal device, and a parameter of an audio playback module of the second terminal device.
[0042] The determining module is configured to determine, according to the hardware parameters of the second terminal device, audio coding / decoding configuration information of the second terminal device, wherein the audio coding / decoding configuration information of the second terminal device is used to indicate an audio coding / decoding capability of the second terminal device.
[0043] The sending module is configured to send the audio coding / decoding configuration information of the second terminal device to a first terminal device.
[0044] receive a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result comprising: first coding / decoding parameters of the first terminal device for coding / decoding an audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.
[0045] In a second aspect of the present application, the composing module of the first terminal device can further perform the steps described in the foregoing first aspect and various possible implementation manners, for details, refer to the foregoing description of the first aspect and various possible implementation manners.
[0046] In a third aspect, the embodiments of the present application provide a computer readable storage medium, the computer readable storage medium stores instructions, when the instructions run on a computer, cause the computer to execute the method in the foregoing first aspect.
[0047] In a fourth aspect, the embodiments of the present application provide a computer program product containing instructions, when the instructions run on a computer, cause the computer to execute the method in the foregoing first aspect.
[0048] In a fifth aspect, the embodiments of the present application provide a communication apparatus, which can include a terminal device or a chip and other entities, the communication apparatus includes: a processor, a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the communication apparatus executes the method in any one of the foregoing first aspect.
[0049] In a sixth aspect, the embodiments of the present application provide a chip system, which includes a processor, used to support the terminal device to realize the functions involved in the foregoing aspects, for example, sending or processing the data and / or information involved in the foregoing method. In a possible design, the chip system further includes a memory, the memory is used to save the necessary program instructions and data of the terminal device. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0050] In a seventh aspect, the embodiments of the present application provide a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive signals from the memory of the electronic device, and send signals to the processor, the signals include computer instructions stored in the memory; when the processor executes the computer instructions, causes the electronic device to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0051] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0052] Firstly, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the hardware parameter of the first terminal device, the hardware parameter of the first terminal device at least including the battery parameter of the first terminal device; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiment of the application, the two terminal devices for audio negotiation each have the audio encoding / decoding configuration information, which describes the encoding / decoding capability that the terminal can support under the constraint of the hardware parameter. By transferring the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiment of the application, the optimal user experience can be provided for the two parties as far as possible under the condition of satisfying the hardware parameter constraint of the two parties of the call. BRIEF DESCRIPTION OF DRAWINGS
[0053] Fig. 1 is a schematic diagram of the component structure of the audio processing system provided by the embodiment of the application;
[0054] Fig. 2a is a schematic diagram of the application of the audio encoder and the audio decoder to the terminal device provided by the embodiment of the application;
[0055] Fig. 2b is a schematic diagram of the application of the audio encoder to the wireless device or the core network device provided by the embodiment of the application;
[0056] Fig. 2c is a schematic diagram of the application of the audio decoder to the wireless device or the core network device provided by the embodiment of the application;
[0057] Fig. 3a is a schematic diagram of the application of the multi-channel encoder and the multi-channel decoder to the terminal device provided by the embodiment of the application;
[0058] Fig. 3b is a schematic diagram of the application of the multi-channel encoder to the wireless device or the core network device provided by the embodiment of the application;
[0059] Fig. 3c is a schematic diagram of the application of the multi-channel decoder to the wireless device or the core network device provided by the embodiment of the application;
[0060] Fig. 4 is a schematic diagram of the negotiation method of the audio encoding / decoding provided by the embodiment of the application;
[0061] Fig. 5 is a diagram of a negotiation process of audio negotiation between two terminals through SDP according to an embodiment of the present application;
[0062] Fig. 6 is a diagram of a negotiation process of audio codec mode and bit rate between two terminals according to an embodiment of the present application;
[0063] Fig. 7 is a diagram of a negotiation result of audio codec mode and bit rate between two terminals according to an embodiment of the present application;
[0064] Fig. 8 is a diagram of another negotiation process of audio codec mode and bit rate between two terminals according to an embodiment of the present application;
[0065] Fig. 9 is a diagram of a structure of a first terminal device according to an embodiment of the present application;
[0066] Fig. 10 is a diagram of a structure of a second terminal device according to an embodiment of the present application;
[0067] Fig. 11 is a diagram of another structure of a first terminal device according to an embodiment of the present application;
[0068] Fig. 12 is a diagram of another structure of a second terminal device according to an embodiment of the present application. DETAILED DESCRIPTION
[0069] Embodiments of the present application will be described below with reference to the accompanying drawings.
[0070] The terms "first", "second", and the like in the description and in the claims of the present application and in the above description of the drawings are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and are for descriptive purposes notwithstanding that all the descriptions set forth pursuant to those terms start with the word "a" or "an". Furthermore, the term "comprising" or "having" as used in the specification includes an open-ended description of the components to be included in the process, method, system, product or apparatus, rather than a limitation on the components to be included in the process, method, system, product or apparatus. It is not intended to exclude any components not explicitly listed.
[0071] The code rate of traditional single-channel speech codec is between several kbps and tens or 128 kbps. Compared with the single-channel speech codec, the immersive speech codec supports more signal types, supports a larger range of code rates, and additionally includes rendering characteristics in addition to the codec characteristics. Taking the 3GPP ongoing immersive voice and audio service (IVAS) speech / audio codec standard as an example, the signal types supported by the IVAS speech / audio codec include single-channel, stereo, multi-channel, multi-object, higher order ambisonics (HOA) signal or first order ambisonics (FOA) signal, MASA, etc. Among them, the multi-channel signal supports up to 7.1.4 channel format, i.e., 12 channels, and the HOA signal supports up to 3-order HOA, i.e., 16 channels, and the supported code rate can be as high as 768 kbps from the lowest single-channel 5.9 kbps. The rendering supports rendering of various signal types to binaural and standard loudspeaker arrays. More signal channels, a larger range of code rates, and additional rendering characteristics significantly increase the maximum computational complexity and storage complexity of the immersive speech codec compared with the traditional speech codec, which will result in a greater difficulty in landing the chip of this type of codec.
[0072] In order to solve the problem of high complexity of the speech codec, the speech codec is divided into levels according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result is to select the lower level for the call to ensure stable call connection. This method meets the speech demand of the call connection, but has the problem of poor audio codec negotiation effect. In addition, this method meets the speech demand of the call connection, but the terminal of the lower level cannot decode all types of code streams, which causes the coding and decoding limitation of the audio signal of the terminal of the higher level, and has the problem of poor quality of the audio signal coding and decoding. In addition, when the levels of the two terminals are not equal, the terminal of the higher level also has the problem of waste of coding and decoding capability.
[0073] Based on the above description, the embodiment of the present application provides an audio coding / decoding negotiation method, which can perform accurate codec negotiation when establishing a call, select the most appropriate codec combination according to the actual capability of the communication platform of each party, ensure the optimal call experience, and ensure that the power consumption is not wasted.
[0074] The embodiments of the present application provide an audio coding technology, in particular, provide a three-dimensional audio coding technology for a three-dimensional audio signal, and specifically provide an encoding technology for representing a three-dimensional audio signal by using fewer channels to improve a conventional audio coding system. Audio coding (or commonly referred to as encoding) includes two parts of audio encoding and audio decoding. Audio encoding is performed at a source side, including processing (for example, compression) of original audio to reduce the amount of data required to represent the audio, thereby more efficiently storing and / or transmitting. Audio decoding is performed at a destination side, including inverse processing relative to the encoder to reconstruct the original audio. The encoding part and the decoding part are also collectively referred to as encoding. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0075] The technical solutions of the embodiments of the present application can be applied to various audio processing systems. As shown in FIG. 1, an audio processing system provided by the embodiments of the present application includes an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 can be used to generate a bitstream, and then the audio encoding bitstream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the bitstream and then perform the audio decoding function of the audio decoding device 102, and finally obtain the reconstructed signal.
[0076] In the embodiments of the present application, the audio encoding device can be applied to various terminal devices that need audio communication, wireless devices that need transcoding, and core network devices. For example, the audio encoding device can be an audio encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding device can be applied to various terminal devices that need audio communication, wireless devices that need transcoding, and core network devices. For example, the audio decoding device can be an audio decoder of the terminal device, the wireless device, or the core network device. For example, the audio encoder can include a radio access network, a media gateway of a core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc. The audio encoder can also be an audio encoder applied to a virtual reality (VR) streaming service.
[0077] In the application embodiment, taking the audio encoding module and audio decoding module suitable for virtual reality streaming service as an example, the processing flow of the audio signal in an end-to-end manner includes: the audio signal A is preprocessed (audio PReprocessing) after being acquired, the preprocessing operation includes filtering out the low-frequency part in the signal, which can be divided by 20 Hz or 50 Hz, extracting the azimuth information in the signal, and then performing encoding processing (audio encoding) and packaging (file / segment encapsulation) and then delivering (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), and then decodes (audio decoding), performs binaural rendering (audio rendering) processing on the decoded signal, and maps the processed signal to the listener's headphones (headphones), which can be independent headphones or headphones on glasses devices.
[0078] As shown in FIG. 2a, the audio encoder and the audio decoder provided by the application embodiment are applied to a schematic diagram of a terminal device. Each terminal device can include an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is configured to perform channel encoding on the audio signal, and the channel decoder is configured to perform channel decoding on the audio signal. For example, the first terminal device 20 can include a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 can include a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a first network communication device 22 via a wireless or wired connection, the first network communication device 22 and a second network communication device 23 are connected via a digital channel, and the second terminal device 21 is connected to the second network communication device 23 via a wireless or wired connection. The wireless or wired network communication device described above can be a signal transmission device, such as a communication base station or a data exchange device.
[0079] In audio communication, the terminal device as a sending end first acquires audio, encodes the acquired audio signal, and then performs channel encoding and transmits the signal in a digital channel through a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a code stream, and then restores the audio signal through audio decoding, and the terminal device as the receiving end plays back the audio.
[0080] As shown in FIG. 2b, a schematic diagram of the audio encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 comprises a channel decoder 251, another audio decoder 252, the audio encoder provided by the embodiment of the present application 253, and a channel encoder 254, wherein the another audio decoder 252 refers to an audio decoder other than the audio decoder. In the wireless device or the core network device 25, the signal entering the device is firstly channel-decoded by the channel decoder 251, then audio-decoded by the another audio decoder 252, then audio-encoded by the audio encoder provided by the embodiment of the present application 253, and finally channel-encoded by the channel encoder 254, and the channel-encoded audio signal is transmitted out. The another audio decoder 252 is used to audio-decode the code stream decoded by the channel decoder 251.
[0081] As shown in FIG. 2c, a schematic diagram of the audio decoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 comprises a channel decoder 251, the audio decoder provided by the embodiment of the present application 255, another audio encoder 256, and a channel encoder 254, wherein the another audio encoder 256 refers to an audio encoder other than the audio encoder. In the wireless device or the core network device 25, the signal entering the device is firstly channel-decoded by the channel decoder 251, then decoded by the audio decoder 255, then audio-encoded by the another audio encoder 256, and finally channel-encoded by the channel encoder 254, and the channel-encoded audio signal is transmitted out. In the wireless device or the core network device, if transcoding is needed, corresponding audio encoding is needed. The wireless device refers to a radio frequency related device in communication, and the core network device refers to a core network related device in communication.
[0082] In some embodiments of the present application, the audio encoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices having transcoding needs, for example, the audio encoding apparatus can be a multi-channel encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices having transcoding needs, for example, the audio decoding apparatus can be a multi-channel decoder of the terminal device, the wireless device, or the core network device.
[0083] As shown in FIG. 3a, a schematic diagram of the multi-channel encoder and the multi-channel decoder provided by the embodiment of the present application applied to terminal devices is shown, and each terminal device can include a multi-channel encoder, a channel encoder, a multi-channel decoder and a channel decoder. The multi-channel encoder can perform the audio encoding method provided by the embodiment of the present application, and the multi-channel decoder can perform the audio decoding method provided by the embodiment of the present application. Specifically, the channel encoder is configured to perform channel encoding on the multi-channel signal, and the channel decoder is configured to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 can include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303 and a first channel decoder 304. The second terminal device 31 can include a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313 and a second channel encoder 314. The first terminal device 30 is connected to a first network communication device 32 via a wireless or wired connection, the first network communication device 32 and a second network communication device 33 are connected via a digital channel, and the second terminal device 31 is connected to the second network communication device 33 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device. In the audio communication, the terminal device as a sending end performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and transmits the signal in the digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a multi-channel signal encoding stream, and then performs multi-channel decoding to restore the multi-channel signal, and the terminal device as the receiving end plays back the multi-channel signal.
[0084] As shown in FIG. 3b, a schematic diagram of the multi-channel encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, another audio decoder 352, a multi-channel encoder 353 and a channel encoder 354. The above-mentioned FIG. 2b is similar, and thus will not be described herein.
[0085] As shown in FIG. 3c, a schematic diagram of the multi-channel decoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, a multi-channel decoder 355, another audio encoder 356 and a channel encoder 354. The above-mentioned FIG. 2c is similar, and thus will not be described herein.
[0086] The audio encoding process can be part of a multi-channel encoder, and the audio decoding process can be part of a multi-channel decoder. For example, multi-channel encoding of a captured multi-channel signal can include processing the captured multi-channel signal to obtain an audio signal, and encoding the obtained audio signal according to the method provided in the embodiments of the present application. The decoding end decodes the multi-channel signal code stream to obtain an audio signal, and recovers the multi-channel signal after upmix processing. Therefore, the embodiments of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In a wireless device or a core network device, if transcoding is required, corresponding multi-channel encoding processing needs to be performed.
[0087] The embodiments of the present application involve two or more terminals that need to communicate, for example, two terminals that communicate, or multiple terminals that communicate. In the following embodiments, two terminals that communicate are taken as an example for description, for example, a calling terminal and a called terminal. The calling terminal can be the encoding end as described above, and the called terminal can be the decoding end. The called terminal can be the encoding end as described above, and the calling terminal can be the decoding end. When two terminals in a mobile network establish a communication, because the voice Codec types supported by the two terminals can be different, and the code rate, sampling rate, etc. supported by the same voice Codec can also be different, the voice Codec needs to be negotiated before the communication is established, to determine which voice Codec and which encoding and decoding configuration are used for the communication. The multiple codecs can include EVS, AMR-WB, and IVAS. The transmission rate can also be referred to as the encoding and decoding rate.
[0088] The negotiation is performed using the Session Description Protocol (SDP). The negotiation process includes negotiation of four important parameters: load type, sampling frequency, rate, and packet length. Specifically:
[0089] The load type is a value defined in a standard protocol.
[0090] The higher the sampling frequency, the more sampling points, and the closer the voice quality to the real voice signal.
[0091] The rate is related to the bandwidth occupancy rate. The higher the rate, the higher the bandwidth occupancy rate.
[0092] The packet length refers to the voice duration contained in each voice packet. The larger the packet length, the larger the packet delay, but the stronger the anti-jitter capability and the higher the bandwidth utilization rate.
[0093] The parameters involved in the negotiation are: codec type and rate, sampling frequency and packet length, which are fixed for each codec type. In the SDP negotiation, the codec type is negotiated first, and then the rate. For AMR-WB and EVS, adaptive rate is supported, and the adaptive range supported is selected during negotiation.
[0094] The SDP template is as follows:
[0095] a=rtpmap:100 AMR / 8000 / / 100 is the load type, AMR represents the codec type, and 8000 represents the sampling frequency
[0096] a=fmtp:100 mode-set=0,2,4,7;mode-change-neighbor=1;mode-change-period=2 / / mode-set represents the rate set, the rate adjustment period and the adjustment mode, etc.
[0097] a=ptime:20 / / packet length.
[0098] First, a method for negotiating audio encoding / decoding provided by an embodiment of the present application is introduced. The method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in the embodiment of the present application can refer to audio encoding, or audio decoding, or audio encoding and decoding. Hereinafter, the audio encoding / decoding refers to the above-mentioned encoding and decoding processes.
[0099] As shown in FIG. 4, the method for negotiating audio encoding / decoding mainly includes the following steps:
[0100] 401. The second terminal device obtains the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device at least include the battery parameters of the second terminal device.
[0101] In the embodiment of the present application, the second terminal device determines the hardware parameters of the second terminal device, and the hardware parameters of the second terminal device can include hardware information of various terminal devices. The types of hardware and hardware configuration parameters involved in the embodiment of the present application are not limited.
[0102] The hardware parameters of the first terminal device include at least one of the following: parameters of an audio acquisition module of the first terminal device and parameters of an audio playback module of the first terminal device.
[0103] The audio acquisition module can be a microphone, and the audio acquisition module can be an audio acquisition module built in the first terminal device or an audio acquisition module connected to the first terminal device, which is not limited herein.
[0104] The audio playback module can be a loudspeaker, and the audio playback module can be an audio playback module built in the first terminal device or an audio playback module connected to the first terminal device, which is not limited herein.
[0105] The battery parameter of the first terminal device, so that the first terminal device can perform audio codec negotiation according to the battery hardware of the first terminal device, for example, the battery parameter can include the capacity, the power, the temperature, the working mode of the battery, and the like.
[0106] 402、The second terminal device determines the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0107] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, so that the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device, for example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.
[0108] The audio encoding / decoding capability of the second terminal device refers to the capability of the second terminal device itself, for example, the audio encoding / decoding capability can include the computing power, the memory, and the like of the second terminal device.
[0109] For example, the audio encoding / decoding configuration information can be a codec mode configuration table.
[0110] 403、The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0111] The second terminal device can send the audio encoding / decoding configuration information of the second terminal device to the first terminal device, so that the first terminal device can perform audio encoding / decoding negotiation.
[0112] It is not limited that in the embodiment of the application, the first terminal device can also send the audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device can perform audio encoding / decoding negotiation.
[0113] 411、The first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0114] 412、the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to hardware parameters of the first terminal device, the hardware parameters of the first terminal device including at least one of the following: parameters of an audio acquisition module of the first terminal device, parameters of an audio playback module of the first terminal device.
[0115] In the embodiment of the application, the first terminal device determines the hardware parameters of the first terminal device, which can include hardware information of various terminal devices. The types of hardware and hardware configuration parameters involved in the embodiment of the application are not limited.
[0116] The hardware parameters of the first terminal device include at least one of the following: parameters of an audio acquisition module of the first terminal device, parameters of an audio playback module of the first terminal device.
[0117] 413、the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0118] The first terminal device can obtain the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device respectively, and can perform encoding / decoding negotiation to obtain the encoding / decoding negotiation result of the first terminal device.
[0119] The encoding / decoding negotiation result includes: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0120] 414、the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0121] 404、the first terminal device receives the encoding / decoding negotiation result from the second terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0122] In the embodiment of the application, the first terminal device and the second terminal device can perform static audio negotiation according to their respective hardware parameters. For example, negotiation is performed based on the hardware capability, processor and bottom chip of the terminal device, or based on the microphone and memory of the terminal device.
[0123] In some embodiments of the present application, the first coding / decoding parameter comprises a first coding mode performed by the first terminal device when coding the audio signal to be processed and a first coding rate corresponding to the first coding mode.
[0124] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, comprising:
[0125] When the hardware parameter of the first terminal device satisfies a preset first hardware configuration condition, the first audio coding / decoding configuration information is determined according to the hardware parameter of the first terminal device;
[0126] When the hardware parameter of the first terminal device does not satisfy the first hardware configuration condition, the second audio coding / decoding configuration information is determined according to the hardware parameter of the first terminal device;
[0127] The audio coding / decoding capability corresponding to the first audio coding / decoding configuration information is higher than the audio coding / decoding capability corresponding to the second audio coding / decoding configuration information.
[0128] In the embodiments of the present application, different hardware parameters of the first terminal device indicate different audio coding / decoding capabilities, thereby completing the audio coding / decoding negotiation based on the hardware parameter.
[0129] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, comprising:
[0130] The audio coding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device and the audio coding / decoding capability of the first terminal device.
[0131] In the embodiments of the present application, when the first terminal device determines the audio coding / decoding configuration information, the hardware parameter of the first terminal device is used, and the audio coding / decoding capability of the first terminal device is also used, so that the audio coding / decoding configuration information of the first terminal device can be more accurately determined.
[0132] In some embodiments of the present application, the hardware parameter further comprises at least one of the following: a parameter of an audio coding / decoding chip of the first terminal device, a parameter of an audio coding / decoding chip of the first terminal device, a battery parameter of the second terminal device, and a parameter of an off-chip memory other than the audio coding / decoding chip in the second terminal device.
[0133] The audio coding / decoding chip of the first terminal device refers to a chip for coding and decoding an audio signal.
[0134] The off-chip memory of the first terminal device other than the audio codec chip refers to other memory of the first terminal device other than the audio codec chip, which can be the off-chip memory of the first terminal device, and the off-chip memory can also be a hardware parameter of the first terminal device.
[0135] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0136] The battery endurance information of the first terminal device is obtained according to the battery parameter of the first terminal device.
[0137] The audio encoding / decoding configuration information of the first terminal device is updated according to the battery endurance information, to obtain the updated audio encoding / decoding configuration information of the first terminal device.
[0138] In the embodiments of the present application, different battery parameters of the first terminal device indicate different battery endurance information, for example, when the battery endurance information indicates more power, stronger audio encoding / decoding capability can be used, and when the battery endurance information indicates that the power decreases, the audio encoding / decoding capability can be reduced, thereby completing the audio encoding / decoding negotiation based on the hardware parameter.
[0139] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0140] The audio signal type and the collection rate supported by the audio collection module of the first terminal device are determined according to the parameter of the audio collection module of the first terminal device.
[0141] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the collection rate supported by the audio collection module of the first terminal device.
[0142] In the embodiments of the present application, different audio collection modules of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, thereby completing the audio encoding / decoding negotiation based on the hardware parameter.
[0143] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameter of the first terminal device, including:
[0144] The audio signal type and the code rate supported by the audio codec chip of the first terminal device are determined according to the parameter of the audio codec chip of the first terminal device.
[0145] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal type and the code rate supported by the audio codec chip of the first terminal device.
[0146] In the embodiments of the present application, different audio codec chips of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, so as to complete the audio encoding / decoding negotiation based on the hardware parameters.
[0147] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the hardware parameters of the first terminal device, including:
[0148] The audio signal types and code rates supported by the off-chip memory of the first terminal device are determined according to the parameters of the off-chip memory of the first terminal device except the audio codec chip.
[0149] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio signal types and code rates supported by the off-chip memory of the first terminal device.
[0150] In the embodiments of the present application, different off-chip memories of the first terminal device support corresponding audio signal types and collection rates, and indicate different audio encoding / decoding capabilities, so as to complete the audio encoding / decoding negotiation based on the hardware parameters.
[0151] In some embodiments of the present application, the method further includes:
[0152] The audio service standard types supported by the audio codec chip of the second terminal device are obtained.
[0153] When the audio service standard types supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device are both the preset first audio service standard type, the step of obtaining the audio encoding / decoding configuration information of the second terminal device is triggered.
[0154] The audio service standard types can include at least one of EVS / IVAS / AMR, etc. The audio service standard types are negotiated before the audio negotiation, so that more accurate audio negotiation can be achieved.
[0155] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding mode supported by the second terminal device, and the rate set corresponding to the encoding / decoding mode supported by the second terminal device.
[0156] The audio encoding / decoding configuration information of the first terminal device includes: the encoding / decoding mode supported by the first terminal device, and the rate set corresponding to the encoding / decoding mode supported by the first terminal device.
[0157] For example, the encoding / decoding mode can include MONO / STEREO / FOA / MASA, and the values of the rates are not equal, and the specific values depend on the corresponding configuration information.
[0158] In some embodiments of the present application, the encoding / decoding negotiation result of the first terminal device is determined according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, including:
[0159] The audio encoding / decoding negotiation order is determined according to the audio encoding / decoding negotiation priority policy of the first terminal device, and the audio encoding / decoding negotiation order includes: performing audio encoding negotiation first or performing audio decoding negotiation first.
[0160] The audio encoding negotiation includes: determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device.
[0161] The audio decoding negotiation includes: determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device.
[0162] The order of the audio encoding negotiation and the audio decoding negotiation is not limited in the embodiments of the present application, and a more flexible audio negotiation mode can be realized.
[0163] In some embodiments of the present application, the encoding negotiation result of the first terminal device is determined according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device, including:
[0164] The encoding negotiation result of the first terminal device includes the first audio mode when the highest encoding mode supported by the first terminal device is the first audio mode and the decoding mode supported by the second terminal device includes the first audio mode.
[0165] The first bit rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device is determined.
[0166] The encoding negotiation result of the first terminal device includes the maximum bit rate in the first bit rate intersection.
[0167] The encoding negotiation result of the first terminal device includes the maximum bit rate in the first bit rate intersection in the embodiments of the present application, which is based on the principle of maximizing experience, and the encoding bit rate of the other terminal with higher experience effect is preferentially selected to improve the call quality.
[0168] In some embodiments of the present application, the decoding negotiation result of the first terminal device is determined according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device, including:
[0169] The highest decoding mode supported by the first terminal device is the second audio mode, and the encoding mode supported by the second terminal device includes the second audio mode, and the decoding negotiation result of the first terminal device includes the second audio mode;
[0170] The second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device are determined to have a second code rate intersection;
[0171] The decoding negotiation result of the first terminal device includes the maximum code rate in the second code rate intersection.
[0172] In the embodiments of the present application, the decoding negotiation result includes the maximum code rate in the second code rate intersection, and according to the principle of maximizing experience, the other terminal is preferentially selected to have a higher encoding code rate, thereby improving the call quality.
[0173] In some embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0174] In the embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device, thereby fully utilizing the maximum code rate of the first terminal device, maximizing the use of the audio encoding / decoding capability of the first terminal device, and improving the call quality.
[0175] In some embodiments of the present application, the first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), high-order ambisonic (HOA), or first-order ambisonic (FOA).
[0176] The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0177] In some embodiments of the present application, the method further includes:
[0178] The first terminal device sends an audio negotiation request to the second terminal device.
[0179] The first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains the hardware parameters of the second terminal device according to the trigger of the audio negotiation request.
[0180] In some embodiments of the present application, the audio negotiation request includes audio encoding / decoding configuration information of the first terminal device.
[0181] The capability negotiation request sent by the first terminal device to the second terminal device carries audio encoding / decoding configuration information of the first terminal device, and the second terminal device can preliminarily select, according to the received audio encoding / decoding configuration information of the first terminal device, to obtain one or more audio encoding / decoding configuration information of the second terminal device after preliminary screening, and then the first terminal device performs audio negotiation, thereby further improving the efficiency of audio negotiation.
[0182] Further comprising: the capability negotiation request sent by the first terminal device to the second terminal device, which can also not include the audio encoding / decoding configuration information of the first terminal device, and is not limited herein.
[0183] The foregoing embodiments illustrate the method performed by the first terminal device, and the method performed by the second terminal device is described as follows. The configuration manner of the hardware parameters of the second terminal device is similar to the configuration manner of the hardware parameters of the first terminal device, and the specific process of determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device is not described in detail.
[0184] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device comprises:
[0185] Determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0186] In some embodiments of the present application, the hardware parameters further comprise at least one of the following: parameters of an audio encoding / decoding chip of the second terminal device, parameters of a battery of the second terminal device, and parameters of an off-chip memory of the second terminal device except the audio encoding / decoding chip.
[0187] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0188] Obtaining battery endurance information of the second terminal device according to the battery parameters of the second terminal device;
[0189] Updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0190] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0191] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the audio collection module of the second terminal device;
[0192] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the audio collection module of the second terminal device;
[0193] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0194] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0195] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the audio encoding / decoding chip of the second terminal device;
[0196] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the audio encoding / decoding chip of the second terminal device;
[0197] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0198] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device, comprising:
[0199] The audio signal type and the collection rate supported by the audio collection module of the second terminal device are determined according to the parameters of the off-chip memory other than the audio encoding / decoding chip in the second terminal device;
[0200] The audio encoding / decoding capability of the second terminal device is determined according to the audio signal type and the collection rate supported by the off-chip memory in the second terminal device;
[0201] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability of the second terminal device.
[0202] In some embodiments of the present application, the method further comprises:
[0203] The audio service standard type supported by the audio encoding / decoding chip of the first terminal device is acquired;
[0204] When the audio service standard type supported by the audio codec chip of the first terminal device and the audio codec chip of the second terminal device is the preset first audio service standard type, the foregoing step of obtaining the hardware parameter of the second terminal device is triggered to be performed.
[0205] In some embodiments of the present application, the audio codec configuration information of the second terminal device includes: a codec mode supported by the second terminal device, and a rate set corresponding to the codec mode supported by the second terminal device.
[0206] In some embodiments of the present application, the first codec parameter includes at least two of the following: a mono (MONO), a stereo (STEREO), a metadata assisted spatial audio (MASA), a higher order ambisonics (HOA), or a first order ambisonics (FOA).
[0207] The rate set corresponding to the codec mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0208] In some embodiments of the present application, the method further includes:
[0209] Receiving an audio negotiation request from the first terminal device.
[0210] In some embodiments of the present application, the audio negotiation request includes: the audio codec configuration information of the first terminal device.
[0211] The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device.
[0212] The method further includes: sending the audio codec configuration information of the second terminal device to the first terminal device, including: sending the candidate audio codec configuration information of the second terminal device to the first terminal device.
[0213] Next, an actual application scenario example is used for illustration.
[0214] The method provided by the embodiments of the present application is suitable for IVAS encoding and decoding, and different signal types and corresponding rates can be used, and further negotiation is required. The embodiments of the present application provide a negotiation method for IVAS codec mode and rate.
[0215] Each IVAS-enabled terminal has its own IVAS codec configuration table, which describes all the IVAS encoding-decoding combinations that can be supported by the terminal under its hardware constraints. By passing the codec configuration table in the negotiation process, the two parties of the call can negotiate the call configuration more flexibly, such as the type of encoded signal, the code rate, etc. In addition, it is also possible to support the two parties of the call to use different signal types and different code rates for the call. The present application can provide the best user experience for both parties as much as possible under the condition of meeting the hardware constraints of both parties.
[0216] Each IVAS-enabled mobile phone will maintain two tables according to the capabilities of the mobile phone (computing power, memory, etc.). One table is the IVAS codec mode configuration table, as shown in Table 1 below:
[0217] The IVAS codec mode configuration table can represent all possible combinations of IVAS encoding modes and decoding modes that the mobile phone can support, including the signal type and code rate of encoding or decoding. As an example in Table 1 above, the mobile phone supports three types of decoded signal types, which are MONO, STEREO, and FOA, among which the maximum code rate supported by MONO decoding is 9.6kbps, the maximum code rate supported by STEREO decoding is 32kbps, and the maximum code rate supported by FOA decoding is 48kbps. When the mobile phone decodes a MONO signal, it can support three types of encoded signal types, which are MONO, STEREO, and FOA, and the maximum code rates supported by encoding these three signals are 9.6kbps, 32kbps, and 48kbps, respectively. When the mobile phone decodes a STEREO signal, it can still support three types of encoded signal types, which are MONO, STEREO, and FOA, but since the decoding overhead of STEREO is larger than that of MONO, the computing power left for encoding will decrease, and therefore the encoding code rate of FOA cannot be supported to 48kbps, but only to 24kbps at most (the higher the code rate, the larger the overhead). When the mobile phone decodes a FOA signal, since the decoding overhead of FOA is even larger, the remaining computing power cannot support the encoding of FOA, and therefore the mobile phone can only support two types of encoded signal types, which are MONO and STEREO, and the maximum code rates supported are 9.6kbps and 32kbps, respectively.
[0218] As can be seen from the above example, the combination of encoding-decoding modes is limited by the capabilities of the terminal / chip, and if the encoding overhead is large, the decoding overhead will decrease, and vice versa.
[0219] The other table is the codec code rate set table, as shown in Table 2 below:
[0220] The coding rate set table can represent all the coding rates supported by IVAS for each signal. In combination with the maximum rate supported by each coding / decoding signal in the coding mode configuration table and the coding rate set table, the rate range supported by each coding / decoding signal in the coding mode configuration table can be obtained.
[0221] Each mobile phone has a default IVAS coding and decoding configuration table when it is shipped. Due to factors such as the chip computing power of the mobile phone, the IVAS coding and decoding configuration table of different mobile phone platforms may be different.
[0222] Since the type of IVAS signal that the mobile phone can encode is also related to some external factors, for example, only when the mobile phone has FOA or HOA acquisition equipment can it support the encoding of FOA / HOA, for example, only when the voice service contains multi-channel or object audio / sound effect signals can it support the encoding of multi-channel or objects, for example, only when the mobile phone supports the complete MASA solution can it support the encoding of MASA signals, for example, when the mobile phone is low in power, it may limit the coding and decoding of some high complexity signals, and so on, which requires the mobile phone to dynamically update its own IVAS coding and decoding configuration table to reflect the current actual ability of the mobile phone.
[0223] As shown in FIG. 5, taking the audio negotiation process of terminal UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) as an example, an example is described:
[0224] Before A and B voice calls, if IVAS voice coding and decoding is used for communication, media negotiation is needed, which includes:
[0225] 1. Audio coding and decoding negotiation (EVS / IVAS / AMR).
[0226] 2. When the audio coding and decoding negotiation is IVAS, the coding and decoding mode (MONO / STEREO / FOA / MASA) and the rate (each coding and decoding mode has a corresponding rate range) need to be further negotiated. The coding mode and the decoding mode can be negotiated to be different.
[0227] Negotiation process:
[0228] 1. When UE-A initiates a call INVITE, if it supports IVAS voice coding and decoding, it sends its own supported IVAS coding and decoding configuration table to B in the SDP information of IVAS, including the list of A supported coding mode and decoding mode combinations, the maximum rate supported by each mode, and the rate set supported by each coding mode.
[0229] 2. UE-B processes as follows according to its own ability:
[0230] 1) First, codec negotiation is carried out, such as supporting IVAS, and preliminary negotiation is IVAS;
[0231] 2) Continue to negotiate the codec mode and rate. According to the local strategy, it is decided whether to negotiate according to the encoding ability or the decoding ability, taking the priority of the encoding ability negotiation as an example:
[0232] ① According to the highest encoding mode supported by UE-B and whether UE-A supports the decoding mode, the encoding mode of UE-B is determined;
[0233] ② According to the intersection of the code rate supported by the calling and called parties for the encoding mode, the encoding code rate of UE-B after negotiation is determined. If there is no intersection, the encoding mode of UE-B needs to be re-negotiated;
[0234] ③ According to the decoding mode corresponding to the encoding mode list supported by UE-B, combined with whether UE-A supports the encoding mode, the decoding mode of UE-B after negotiation is determined;
[0235] ④ According to the intersection of the code rate of the calling and called parties for the decoding mode, the decoding code rate of UE-B after negotiation is determined.
[0236] If any of the codec modes negotiated by IVAS fails, IVAS will not be finally used.
[0237] As shown in FIGS. 6 and 7, the SDP example adds the following parameters in the offer a line of IVAS:
[0238] Code rate set:
[0239] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.
[0240] [mono,5.9 / 7.2 / 9.6] represents that the code rate set supported by mono is 5.9, 7.2 and 9.6.
[0241] Decoding mode set:
[0242] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.
[0243] The list represents that 3 sets are supported, each set corresponds to the following encoding mode set one by one, set 1 corresponds to the decoding mode [mono / 9.6], which represents that mono is supported and the highest code rate is 9.6.
[0244] Encoding mode set:
[0245] enc-mode-list = {[mono / 9.6, stereo / 32, foa / 48], [mono / 9.6,
[0246] stereo / 32, foa / 24], [mono / 9.6, stereo / 32]}.
[0247] This list indicates that 3 sets are supported, which are one-to-one corresponding to the above decoding mode set, and the encoding mode corresponding to set 1 is = { [mono / 9.6, stereo / 32, foa / 48] indicates that the decoding mode is mono / stereo / foa from low to high, and the highest code rate is 9.6 / 32 / 48.
[0248] Add the following parameters in the answer a line of IVAS:
[0249] The negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]};
[0250] The negotiated encoding mode and rate set: enc-mode = {[foa, 24 / 32]}.
[0251] An example of the negotiation process:
[0252] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the encoding and decoding mode supported by A and the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0253] 2. Mobile phone B decides according to the local strategy whether to negotiate according to its own encoding ability or according to its own decoding ability, taking the encoding ability negotiation as an example:
[0254] ① According to the highest encoding mode supported by UE-B, which is FOA, and the decoding mode of UE-A contains FOA, it is determined that the encoding mode of UE-B after negotiation is FOA;
[0255] ② Determine the intersection of the code rates supported by the main and called FOA modes, and determine that the encoding code rate of UE-B after negotiation is 24 and 32;
[0256] ③ According to the decoding mode corresponding to the encoding mode column supported by UE-B, which is MONO, and the encoding mode of UE-A contains MONO, it is determined that the decoding mode of UE-B after negotiation is MONO;
[0257] ④ Determine the intersection of the code rates supported by the main and called MONO modes, and determine that the decoding code rate of UE-B after negotiation is 5.9, 7.2 and 9.6.
[0258] As shown in FIG. 8, another example of the negotiation process:
[0259] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A and the maximum rate combinations supported by each mode, and the rate set supported by each coding mode.
[0260] 2. Mobile phone B decides according to its local policy whether to negotiate according to its coding capability or according to its decoding capability, taking the coding capability negotiation as an example:
[0261] ① According to the highest decoding mode supported by UE-B being FOA and the decoding mode of UE-A containing FOA, it is determined that the decoding mode of UE-B after negotiation is FOA;
[0262] ② The intersection of the code rates of the FOA mode supported by the main and called parties is determined, and it is determined that the coding code rate of UE-B after negotiation is 24 and 32, 48;
[0263] ③ According to the highest coding mode supported by the decoding mode column of UE-B being STEREO and the decoding mode of UE-A containing STEREO, it is determined that the coding mode of UE-B after negotiation is STEREO;
[0264] ④ The intersection of the code rates of the STEREO mode supported by the main and called parties is determined, and it is determined that the decoding code rate of UE-B after negotiation is 13.2 and 24.
[0265] As can be known from the foregoing example, for real-time communication audio Codec with numerous coding modes and coding rates and large span of coding and decoding complexity, each terminal saves a codec configuration table for it. The configuration table describes various combinations of coding and decoding modes / rates that the terminal can support. When two or more terminals establish a call to negotiate the Codec, according to their respective codec configuration tables, they select their respective coding configurations for coding.
[0266] Each terminal has a default codec configuration table from the factory. When a call is established to negotiate the Codec, the terminal can update its codec configuration table according to other factors, and negotiate the Codec according to the updated codec configuration table. The updated codec configuration table does not overwrite the default factory codec configuration table.
[0267] In the codec configuration table, the supported decoding modes / code rates and / or coding modes / code rates are sorted by experience priority.
[0268] During Codec negotiation, one terminal sends its complete codec configuration table or part of the codec configuration table to the other terminal for negotiation, and the terminal makes a decision according to the entire or partial codec configuration tables of at least two terminals. Here, part of the coding mode and rate refers to a subset.
[0269] The decision of the negotiation can be based on the principle of maximizing experience, or other principles, such as a preset decision principle in the terminal making the decision of the negotiation, for example, preferentially selecting the other terminal to perform coding with a higher experience effect.
[0270] Next, an audio negotiation process for terminal power is described by way of example.
[0271] When the terminal power is 100%, the configuration table of the terminal is an initial configuration table, for example, as shown in Table 3 below:
[0272] When the terminal power decreases, the configuration table needs to be adjusted to reduce the configuration to ensure the operation of other functions of the terminal, for example, when the power is only 10%, the configuration table is adjusted to be as shown in Table 4 below:
[0273] When the terminal only uses configuration 1 (default hardware configuration / factory configuration), the configuration table of the terminal is an initial configuration table, as shown in Table 3 above.
[0274] When the terminal uses configuration 2 (external FOA microphone), the configuration table is modified for some combinations of Table 3.
[0275] When the terminal uses configuration 3 (external HOA microphone), the configuration table is modified to replace combination 3 in Table 3 with HOA and the corresponding code rate.
[0276] When the terminal uses configuration 4 (external microphone), the terminal-side device can convert the microphone signal into a Codec-supported FOA / HOA / MASA signal, and the configuration table includes combination 3 as FOA and the corresponding code rate, combination 4 as HOA and the corresponding code rate, and combination 5 as MASA and the corresponding code rate.
[0277] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0278] In order to better implement the above-mentioned scheme of the present application, the following provides a related device for implementing the above-mentioned scheme.
[0279] Referring to FIG. 9, a first terminal device 900 provided by an embodiment of the present application can include an acquisition module 901, a determination module 902, a negotiation module 903, and a sending module 904, wherein,
[0280] obtaining an audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device;
[0281] determining an audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device including at least one of a parameter of an audio acquisition module of the first terminal device and a parameter of an audio playback module of the first terminal device;
[0282] determining an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including a first encoding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device;
[0283] sending the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0284] As can be known from the foregoing embodiments, first, the first terminal device obtains an audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; then, the first terminal device determines an audio encoding / decoding configuration information of the first terminal device according to a hardware parameter of the first terminal device, the hardware parameter of the first terminal device including at least one of a parameter of an audio acquisition module of the first terminal device and a parameter of an audio playback module of the first terminal device; next, the first terminal device determines an encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including a first encoding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiments of the present application, both of the two terminals for audio negotiation have audio encoding / decoding configuration information, which describes the encoding / decoding capability that can be supported by the terminal under the constraint of the hardware parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiments of the present application, the optimal user experience can be provided for the two parties as much as possible under the constraint of the hardware parameter of the two parties of the call.
[0285] Referring to FIG. 10, a second terminal device 1000 provided by an embodiment of the present application can include an obtaining module 1001, a determining module 1002, a sending module 1003, and a receiving module 1004, wherein
[0286] The obtaining module is configured to obtain a hardware parameter of the second terminal device, and the hardware parameter of the second terminal device includes at least one of a parameter of an audio acquisition module of the second terminal device and a parameter of an audio playback module of the second terminal device.
[0287] The determining module is configured to determine audio encoding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate an audio encoding / decoding capability of the second terminal device.
[0288] The sending module is configured to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0289] The receiving module is configured to receive a coding / decoding negotiation result from the first terminal device, and the coding / decoding negotiation result includes a first coding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, and the first coding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0290] It should be noted that the information interaction and execution process between the modules / units of the above apparatus are based on the same consideration as the method embodiments of the present application, and the technical effects brought by the same are the same as those of the method embodiments of the present application. For details, refer to the description in the foregoing method embodiments of the present application, which will not be repeated here.
[0291] The embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a program, and the program performs part or all of the steps recorded in the foregoing method embodiments.
[0292] Next, another first terminal device provided by an embodiment of the present application is introduced. Referring to FIG. 11, the first terminal device 1100 includes a receiver 1101, a transmitter 1102, a processor 1103, and a memory 1104 (wherein the number of the processor 1103 in the first terminal device 1100 can be one or more, and one processor is taken as an example in FIG. 11).
[0293] The receiver 1101, the transmitter 1102, the processor 1103, and the memory 1104 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 11.
[0294] The memory 1104 can include read-only memory and random access memory, and provide instructions and data to the processor 1103. A portion of the memory 1104 can also include non-volatile random access memory (NVRAM). The memory 1104 stores operating systems and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0295] The processor 1103 controls the operation of the first terminal device, and the processor 1103 can also be referred to as a central processing unit (CPU). In specific applications, various components of the first terminal device are coupled together through a bus system, which can include a data bus, a power bus, a control bus, a state signal bus, etc. However, for the sake of clarity, various buses are referred to as a bus system in the figure.
[0296] The method disclosed in the above embodiments of the present application can be applied in the processor 1103 or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1103 or an instruction in the form of software. The processor 1103 described above can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 1104, and the processor 1103 reads the information in the memory 1104 and combines the hardware to complete the steps of the above method.
[0297] The receiver 1101 can be configured to receive inputted digital or character information, and generate signal input related to the relevant settings and function control of the first terminal device. The transmitter 1102 can include a display device such as a display screen, and the transmitter 1102 can be configured to output digital or character information through an external interface.
[0298] In the embodiments of the present application, the processor 1103 is configured to execute the method performed by the first terminal device shown in FIG. 4.
[0299] Next, another second terminal device provided by the embodiments of the present application is introduced. Referring to FIG. 12, the second terminal device 1200 includes:
[0300] The receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 (wherein the number of the processor 1203 in the second terminal device 1200 can be one or more, and one processor is taken as an example in FIG. 12). In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 12.
[0301] The memory 1204 can include read-only memory and random access memory, and provide the processor 1203 with instructions and data. A portion of the memory 1204 can also include NVRAM. The memory 1204 stores an operating system and operation instructions, executable modules or data structures, or subsets thereof, or expanded sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0302] The processor 1203 controls the operation of the second terminal device, and the processor 1203 can also be referred to as a CPU. In specific applications, various components of the second terminal device are coupled together through a bus system, wherein the bus system can include a data bus in addition to a power bus, a control bus and a status signal bus, etc. However, in order to clearly illustrate, various buses are referred to as a bus system in the figure.
[0303] The method disclosed in the embodiments of the present application can be applied to the processor 1203 or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the above method can be completed by using an integrated logic circuit or a software form of the instructions in the processor 1203. The processor 1203 can be a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed by the processor 1203. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code executed by the processor, or a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an electrically programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, or a removable disk. The storage medium is located in the storage 1204, and the processor 1203 reads information in the storage 1204 and combines the hardware to complete the steps of the above method.
[0304] In the embodiments of the present application, the processor 1203 is configured to execute the method performed by the second terminal device shown in Fig. 7.
[0305] In another possible design, when the first terminal device or the second terminal device is a chip in the terminal, the chip includes a processing unit and a communication unit. The processing unit can be a processor, and the communication unit can be an input / output interface, a pin, a circuit or the like. The processing unit can execute computer-executable instructions stored in a storage unit, so that the chip in the terminal executes the method of any one of the first aspect. Alternatively, the storage unit is a storage unit in the chip, such as a register, a cache or the like. The storage unit can also be a storage unit in the terminal located outside the chip, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or the like.
[0306] The processor mentioned in any of the above can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of programs of the method of the first aspect or the second aspect.
[0307] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate units can or can not be physically separate, and the units displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0308] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in various embodiments of the present application.
[0309] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.
[0310] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. An audio encoding / decoding negotiation method, characterized by, The method is applied to a second terminal device, and the method comprises: obtaining hardware parameters of the second terminal device, the hardware parameters of the second terminal device comprising at least one of the following: parameters of an audio acquisition module of the second terminal device, parameters of an audio playback module of the second terminal device; determining audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; sending the audio encoding / decoding configuration information of the second terminal device to a first terminal device; receiving a result of encoding / decoding negotiation from the first terminal device, the result of encoding / decoding negotiation comprising: first encoding / decoding parameters of the first terminal device for encoding / decoding an audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
2. The method of claim 1, wherein, The method further comprises: determining the audio encoding / decoding configuration information of the second terminal device according to the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device.
3. The method according to claim 1 or 2, characterized in that, The hardware parameters of the second terminal device further comprise at least one of the following: parameters of an audio encoding / decoding chip of the second terminal device, battery parameters of the second terminal device, and parameters of an off-chip memory of the second terminal device other than the audio encoding / decoding chip.
4. The method of claim 2, wherein, The method further comprises: obtaining battery endurance information of the second terminal device according to the battery parameters of the second terminal device; updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
5. The method according to claim 2 or 3, characterized in that, The method further comprises: determining an audio signal type and an acquisition rate supported by the audio acquisition module of the second terminal device according to the parameters of the audio acquisition module of the second terminal device; determining the audio encoding / decoding capability of the second terminal device according to the audio signal type and the acquisition rate supported by the audio acquisition module of the second terminal device; determining the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability of the second terminal device.
6. The method according to claim 2 or 3, characterized in that, The method further comprises: determining an audio signal type and an acquisition rate supported by the audio acquisition module of the second terminal device according to the parameters of the audio encoding / decoding chip of the second terminal device; determining the audio coding / decoding capability of the second terminal device according to the audio signal type and the sampling rate supported by the audio coding chip of the second terminal device; determining the audio coding / decoding configuration information of the second terminal device according to the audio coding / decoding capability of the second terminal device.
7. The method of claim 2 or 3, wherein, The determining the audio coding / decoding configuration information of the second terminal device according to the hardware parameter of the second terminal device and the audio coding / decoding capability of the second terminal device comprises: determining the audio signal type and the sampling rate supported by the audio acquisition module of the second terminal device according to the parameter of the off-chip memory of the second terminal device other than the audio coding chip; determining the audio coding / decoding capability of the second terminal device according to the audio signal type and the sampling rate supported by the off-chip memory of the second terminal device; determining the audio coding / decoding configuration information of the second terminal device according to the audio coding / decoding capability of the second terminal device.
8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: obtaining the audio service standard type supported by the audio coding chip of the first terminal device; when the audio service standard type supported by the audio coding chip of the first terminal device and the audio coding chip of the second terminal device is the preset first audio service standard type respectively, triggering the execution of the aforementioned step of obtaining the hardware parameter of the second terminal device.
9. The method according to any one of claims 1 to 8, characterized in that, The audio coding / decoding configuration information of the second terminal device comprises the coding mode supported by the second terminal device and the rate set corresponding to the coding mode supported by the second terminal device.
10. The method according to any one of claims 1 to 10, characterized in that, The first coding / decoding parameter comprises at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), higher order ambisonics (HOA), or first order ambisonics (FOA); The rate set corresponding to the coding mode comprises at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
11. The method according to any one of claims 1 to 10, characterized in that, The method further comprises: receiving the audio negotiation request from the first terminal device.
12. The method of claim 11, wherein, The audio negotiation request comprises the audio coding / decoding configuration information of the first terminal device. The method further comprises: screening the audio coding / decoding configuration information of the second terminal device according to the audio coding / decoding configuration information of the first terminal device to obtain the candidate audio coding / decoding configuration information of the second terminal device. The sending of the audio coding / decoding configuration information of the second terminal device to the first terminal device comprises the sending of the candidate audio coding / decoding configuration information of the second terminal device to the first terminal device.
13. A terminal device, comprising: The terminal device is specifically the second terminal device, and the method comprises: an obtaining module, configured to obtain the hardware parameter of the second terminal device, the hardware parameter of the second terminal device comprising at least one of the following: the parameter of the audio acquisition module of the second terminal device and the parameter of the audio playback module of the second terminal device; determining, by a determining module, audio encoding / decoding configuration information of the second terminal device according to a hardware parameter of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; sending, by a sending module, the audio encoding / decoding configuration information of the second terminal device to the first terminal device; receiving, by a receiving module, a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result including a first coding / decoding parameter used by the first terminal device to encode / decode the audio signal to be processed, the first coding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device.
14. A terminal device, comprising: The terminal device includes at least one processor, which is used to be coupled with a memory, read and execute instructions in the memory, so as to implement the method in any one of claims 1 to 12.
15. The terminal device according to claim 14, characterized by The terminal device further includes the memory.
16. A computer readable storage medium, including instructions which, when run on a computer, cause the computer to perform the method in any one of claims 1 to 12.