Negotiation method for audio encoding / decoding and terminal device
By acquiring the hardware parameters of the terminal device, especially the battery parameters, the audio codec configuration information is determined, enabling precise codec negotiation in the immersive voice codec. This solves the problems of poor codec performance and resource waste, and improves the user experience.
Patent Information
- Application Number
- PCT/CN2025/082507
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-14
- Publication Date
- 2026-01-15
AI Technical Summary
Existing audio codec negotiation methods cannot select the optimal codec level when supporting highly complex immersive speech codecs, resulting in stable call connections but poor codec performance and wasted resources.
By acquiring the hardware parameters of the terminal device, especially the battery parameters, the audio codec configuration information is determined and transmitted during the negotiation process, so that both terminals can select the most suitable codec parameters according to their respective hardware capabilities, thus achieving accurate codec negotiation.
While meeting the constraints of their respective hardware parameters, it provides more flexible call configurations, improves user experience, and avoids problems such as resource waste and poor codec quality.
Smart Images

Figure CN2025082507_15012026_PF_FP_ABST
Abstract
Description
An audio encoding / decoding negotiation method and terminal device
[0001] This application claims priority to Chinese Patent Application No. 202410350992.4, filed on March 25, 2024, entitled "A Negotiation Method and Terminal Device for Audio Encoding / Decoding", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of audio encoding and decoding, and more particularly to an audio encoding / decoding negotiation method and terminal device. Background Technology
[0003] Compared to mono speech codecs, immersive speech codecs support more audio signal types, a wider bitrate range, and include rendering features in addition to encoding / decoding capabilities. Taking the ongoing audio codec standards of the 3rd Generation Partnership Project (3GPP) as an example, different audio codec standards support a variety of signal types, each with different numbers of channels and bitrates. The rendering features for different signal types support rendering for various signal types, from binaural to standard speaker arrays. Because immersive speech codecs support more signal channels, a wider bitrate range, and additional rendering features, their maximum computational and storage complexity is significantly increased compared to traditional speech codecs.
[0004] To support more complex speech codecs, one existing audio codec negotiation method categorizes speech codecs according to their complexity. However, when two terminals supporting different complexity levels establish a call, the current negotiation result can only select the lower complexity level to ensure a stable call connection. While this method meets the voice requirements of the call connection, it suffers from poor audio codec negotiation performance. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides the following technical solutions:
[0006] In a first aspect, embodiments of this application provide an audio encoding / decoding negotiation method, the method being applied to a second terminal device, the method comprising:
[0007] Obtain the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device include at least the battery parameters of the second terminal device;
[0008] The audio encoding / decoding configuration information of the second terminal device is determined based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0009] Send the audio encoding / decoding configuration information of the second terminal device to the first terminal device;
[0010] The first terminal device receives an encoding / decoding negotiation result, which includes: first encoding / decoding parameters for encoding and decoding the audio signal to be processed by the first terminal device, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.
[0011] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on its hardware parameters includes:
[0012] The audio encoding / decoding configuration information of the second terminal device is determined based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device.
[0013] In some embodiments of this application, the hardware parameters further include at least one of the following: parameters of the audio codec chip of the second terminal device, parameters of the audio acquisition module of the second terminal device, parameters of the audio playback module of the first terminal device, and parameters of the off-chip memory of the second terminal device other than the audio codec chip.
[0014] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes:
[0015] The battery life information of the second terminal device is obtained based on the battery parameters of the second terminal device;
[0016] The audio encoding / decoding configuration information of the second terminal device is updated based on the battery life information and the audio encoding / decoding capability of the second terminal device to obtain the updated audio encoding / decoding configuration information of the second terminal device.
[0017] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes:
[0018] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio acquisition module of the second terminal device.
[0019] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device;
[0020] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0021] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes:
[0022] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio codec chip of the second terminal device.
[0023] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio codec chip of the second terminal device;
[0024] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0025] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes:
[0026] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the off-chip memory (excluding the audio codec chip) in the second terminal device.
[0027] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal type and acquisition rate supported by the off-chip memory in the second terminal device;
[0028] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0029] In some embodiments of this application, the method further includes:
[0030] Obtain the audio service standard types supported by the audio codec chip of the first terminal device;
[0031] When the audio codec chip of the first terminal device and the audio codec chip of the second terminal device each support the preset first audio service standard type, the aforementioned step of obtaining the hardware parameters of the second terminal device is triggered.
[0032] In some embodiments of this application, the audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding modes supported by the second terminal device, and the set of rates corresponding to the encoding / decoding modes supported by the second terminal device.
[0033] In some embodiments of this application, the first encoding / decoding parameters include at least two of the following: mono MONO, stereo STEREO, metadata-assisted spatial audio MASA, high-order stereo reverb HOA, or first-order stereo reverb FOA.
[0034] The set of rates corresponding to the encoding / decoding modes includes at least two of the following: 9.6kbps, 13.2kbps, 32kbps, and 48kbps.
[0035] In some embodiments of this application, the method further includes:
[0036] Receive an audio negotiation request from the first terminal device.
[0037] In some embodiments of this application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device;
[0038] The method further includes: filtering the audio encoding / decoding configuration information of the second terminal device based on the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device;
[0039] Sending the audio encoding / decoding configuration information of the second terminal device to the first terminal device includes: sending the candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0040] Secondly, embodiments of this application provide a terminal device, specifically a second terminal device, and the method includes:
[0041] The acquisition module is used to acquire the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device include at least the battery parameters of the second terminal device.
[0042] The determining module is used to determine the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0043] The sending module is used to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device;
[0044] The receiving module is configured to receive an encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0045] In the second aspect of this application, the constituent modules of the second terminal device may also perform the steps described in the first aspect and various possible implementations, as detailed in the foregoing description of the first aspect and various possible implementations.
[0046] Thirdly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect above.
[0047] Fourthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect above.
[0048] Fifthly, embodiments of this application provide a communication device, which may include entities such as terminal devices or chips. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the preceding first aspects.
[0049] Sixthly, this application provides a chip system including a processor for supporting a terminal device in implementing the functions involved in the above aspects, such as transmitting or processing data and / or information involved in the above methods. In one possible design, the chip system further includes a memory for storing necessary program instructions and data for the terminal device. This chip system may be composed of chips or may include chips and other discrete devices.
[0050] In a seventh aspect, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the electronic device to perform the method in the first aspect or any possible implementation of the first aspect.
[0051] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0052] First, the first terminal device obtains the audio encoding / decoding configuration information of the second terminal device. This configuration information indicates the audio encoding / decoding capabilities of the second terminal device. Then, the first terminal device determines its own audio encoding / decoding configuration information based on its hardware parameters, which include at least its battery parameters. Next, the first terminal device determines its encoding / decoding negotiation result based on both the first and second terminal device's configuration information. This negotiation result includes the first encoding / decoding parameters used by the first terminal device to encode and decode the audio signal to be processed. These parameters are adapted to the audio encoding / decoding capabilities of the second terminal device. Finally, the first terminal device sends its own encoding / decoding negotiation result to the second terminal device. In this embodiment, both terminals involved in audio negotiation have their own audio encoding / decoding configuration information, which describes the encoding / decoding capabilities that the terminal can support under hardware parameter constraints. By transmitting audio encoding / decoding configuration information during the negotiation process, the two parties in the call can negotiate the call configuration more flexibly. In this embodiment of the application, the best user experience can be provided to both parties as much as possible while satisfying the hardware parameter constraints of each party. Attached Figure Description
[0053] Figure 1 is a schematic diagram of the composition structure of the audio processing system provided in an embodiment of this application;
[0054] Figure 2a is a schematic diagram of the audio encoder and audio decoder provided in the embodiments of this application applied to a terminal device;
[0055] Figure 2b is a schematic diagram of the audio encoder provided in the embodiments of this application applied to wireless devices or core network devices;
[0056] Figure 2c is a schematic diagram of the audio decoder provided in the embodiment of this application being applied to a wireless device or a core network device;
[0057] Figure 3a is a schematic diagram of the multi-channel encoder and multi-channel decoder provided in the embodiments of this application applied to a terminal device;
[0058] Figure 3b is a schematic diagram of the multi-channel encoder provided in the embodiments of this application applied to wireless devices or core network devices;
[0059] Figure 3c is a schematic diagram of the multi-channel decoder provided in the embodiment of this application applied to a wireless device or a core network device;
[0060] Figure 4 is a schematic diagram of an audio encoding / decoding negotiation method provided in an embodiment of this application;
[0061] Figure 5 is a schematic diagram of a negotiation process for audio negotiation between two terminals via SDP, provided in an embodiment of this application.
[0062] Figure 6 is a schematic diagram of the negotiation process between two terminals regarding audio encoding / decoding mode and bitrate, provided in an embodiment of this application.
[0063] Figure 7 is a schematic diagram of an encoding / decoding negotiation result generated after negotiation between two terminals according to an embodiment of this application;
[0064] Figure 8 is a schematic diagram of another negotiation process between the two terminals regarding the audio codec mode and bit rate provided in an embodiment of this application;
[0065] Figure 9 is a schematic diagram of the composition structure of a first terminal device provided in an embodiment of this application;
[0066] Figure 10 is a schematic diagram of the composition structure of a second terminal device provided in an embodiment of this application;
[0067] Figure 11 is a schematic diagram of the composition structure of another first terminal device provided in an embodiment of this application;
[0068] Figure 12 is a schematic diagram of the composition structure of another second terminal device provided in an embodiment of this application. Detailed Implementation
[0069] The embodiments of this application will now be described with reference to the accompanying drawings.
[0070] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0071] Traditional mono speech codecs have bitrates ranging from a few kbps to tens or even 128 kbps. Immersive speech codecs, compared to mono codecs, support more signal types, have a wider bitrate range, and include additional rendering features in addition to encoding and decoding capabilities. Taking the 3GPP Immersive Voice and Audio Services (IVAS) speech / audio codec standard as an example, IVAS speech / audio codecs support signal types including mono, stereo, multi-channel, multi-object, higher-order ambisonic (HOA) signals, first-order ambisonic (FOA) signals, MASA, etc. The multi-channel signal supports a maximum of 7.1.4 channel format, meaning 12 channels, while the HOA signal supports up to 3rd order HOA, meaning 16 channels. Supported bitrates range from a minimum of 5.9kbps per channel to a maximum of 768kbps for 3rd order HOA. Rendering supports various signal types, from binaural to standard speaker arrays. The increased number of signal channels, wider bitrate range, and additional rendering features significantly increase the maximum computational and storage complexity of immersive voice codecs compared to traditional voice codecs, greatly increasing the difficulty of chip deployment for this type of codec.
[0072] To address the issue of high complexity in voice codecs, voice codecs are categorized into levels based on complexity. However, when two terminals supporting different levels establish a call, the current negotiation result only allows the lower complexity level to be selected to ensure a stable call connection. While this method satisfies the voice requirements of the call connection, it suffers from poor audio codec negotiation performance. Furthermore, although this method meets the voice requirements, the lower-level terminal cannot decode all types of bitstreams, imposing audio signal encoding and decoding limitations on the higher-level terminal, resulting in poor audio signal encoding and decoding quality. Additionally, when the two terminals have unequal complexity levels, the higher-level terminal also suffers from wasted encoding and decoding capabilities.
[0073] Based on the above description, this application provides an audio codec negotiation method, which can perform precise codec negotiation when establishing a call, and select the most appropriate codec combination according to the actual capabilities of the two communicating platforms, so as to ensure the best call experience and avoid wasting power consumption.
[0074] This application provides an audio encoding technique, particularly a three-dimensional audio encoding technique for three-dimensional audio signals. Specifically, it provides an encoding technique that uses fewer channels to represent three-dimensional audio signals, thereby improving traditional audio encoding systems. Audio encoding (or commonly referred to as encoding) includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and includes processing (e.g., compressing) the raw audio to reduce the amount of data required to represent the audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are also collectively referred to as encoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0075] The technical solutions of this application embodiment can be applied to various audio processing systems. As shown in Figure 1, this is a schematic diagram of the composition structure of the audio processing system provided in this application embodiment. The audio processing system 100 may include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 can generate a bitstream, which can then be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the bitstream, execute its audio decoding function, and finally obtain the reconstructed signal.
[0076] In the embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be an audio encoder for the aforementioned terminal devices, wireless devices, or core network devices. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be an audio decoder for the aforementioned terminal devices, wireless devices, or core network devices. For example, the audio encoder can include a wireless access network, a core network media gateway, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The audio encoder can also be an audio encoder used in virtual reality (VR) streaming services.
[0077] In the application embodiment, taking the audio encoding module (audio encoding and audio decoding) applicable to virtual reality streaming (VR streaming) services as an example, the end-to-end audio signal processing flow includes: after the audio signal A passes through the acquisition module, it undergoes a preprocessing operation (audioPReprocessing). The preprocessing operation includes filtering out the low-frequency part in the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the directional information in the signal, and then performing encoding processing (audio encoding), packaging (file / segment encapsulation), and then sending (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), and performs binaural rendering processing on the decoded signal. The rendered signal is mapped onto the listener's headphones, which can be independent headphones or headphones on the glasses device.
[0078] Figure 2a shows a schematic diagram of the audio encoder and audio decoder provided in this application embodiment applied to a terminal device. Each terminal device may include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used to perform channel encoding on the audio signal, and the channel decoder is used to perform channel decoding on the audio signal. For example, the first terminal device 20 may include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 may include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and a wireless or wired second network communication device 23 are connected via a digital channel. The second terminal device 21 is connected to the wireless or wired second network communication device 23. The aforementioned wireless or wired network communication device can refer to signal transmission devices, such as communication base stations, data exchange devices, etc.
[0079] In audio communication, the transmitting terminal device first acquires audio, encodes the acquired audio signal, and then performs channel coding before transmitting it over a digital channel via a wireless network or core network. The receiving terminal device, acting as the receiver, decodes the received signal to obtain the bitstream, then recovers the audio signal through audio decoding for playback.
[0080] Figure 2b illustrates the application of the audio encoder provided in this embodiment of the application in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, the audio encoder 253 provided in this embodiment of the application, and a channel encoder 254. The other audio decoders 252 refer to audio decoders other than the standard audio decoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, then audio-decoded by the other audio decoder 252, then audio-encoded by the audio encoder 253 provided in this embodiment of the application, and finally channel-encoded by the channel encoder 254. After channel encoding, the signal is transmitted out. The other audio decoders 252 perform audio decoding on the bitstream decoded by the channel decoder 251.
[0081] Figure 2c shows a schematic diagram of the audio decoder provided in this application embodiment applied to a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in this application embodiment, other audio encoders 256, and a channel encoder 254. The other audio encoders 256 refer to audio encoders other than the audio encoder. Within the wireless device or core network device 25, the channel decoder 251 first performs channel decoding on the signal entering the device. Then, the audio decoder 255 decodes the received audio encoded bitstream. Next, the other audio encoders 256 perform audio encoding. Finally, the channel encoder 254 performs channel encoding on the audio signal before transmission. If transcoding is required in the wireless device or core network device, corresponding audio encoding processing is necessary. The wireless device refers to radio frequency related devices in communication, and the core network device refers to core network related devices in communication.
[0082] In some embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be a multi-channel encoder of the aforementioned terminal device, wireless device, or core network device. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be a multi-channel decoder of the aforementioned terminal device, wireless device, or core network device.
[0083] Figure 3a shows a schematic diagram of the multi-channel encoder and multi-channel decoder provided in this application embodiment applied to a terminal device. Each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in this application embodiment, and the multi-channel decoder can execute the audio decoding method provided in this application embodiment. Specifically, the channel encoder is used to perform channel encoding on the multi-channel signal, and the channel decoder is used to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. The second terminal device 31 may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and a wireless or wired second network communication device 33 are connected via a digital channel. The second terminal device 31 is connected to the wireless or wired second network communication device 33. The aforementioned wireless or wired network communication equipment can broadly refer to signal transmission equipment, such as communication base stations and data switching equipment. In audio communication, the transmitting terminal device performs multi-channel encoding on the acquired multi-channel signal, followed by channel encoding, and then transmits it through a wireless network or core network in a digital channel. The receiving terminal device performs channel decoding on the received signal to obtain the multi-channel signal encoded bitstream, and then recovers the multi-channel signal through multi-channel decoding for playback by the receiving terminal device.
[0084] As shown in Figure 3b, this is a schematic diagram of the multi-channel encoder provided in this application being applied to a wireless device or a core network device. The wireless device or core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, which are similar to those in Figure 2b above and will not be described again here.
[0085] As shown in Figure 3c, this is a schematic diagram of the multi-channel decoder provided in this application being applied to a wireless device or a core network device. The wireless device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354. Similar to Figure 2c above, it will not be described again here.
[0086] The audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the acquired multi-channel signal can involve processing the acquired multi-channel signal to obtain an audio signal, and then encoding the obtained audio signal according to the method provided in this application embodiment. The decoding end decodes the audio signal based on the multi-channel signal encoded bitstream, and recovers the multi-channel signal after upmixing. Therefore, this application embodiment can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding is required, corresponding multi-channel encoding processing is necessary.
[0087] This application involves two or more terminals that need to conduct a call. For example, it could be two terminals conducting a call, or it could be multiple terminals. Subsequent embodiments will use a two-party call as an example, such as a calling terminal and a called terminal. When the calling terminal is the aforementioned encoding end, the called terminal can be the decoding end; conversely, when the called terminal is the aforementioned encoding end, the calling terminal can be the decoding end. When two parties establish a call in a mobile network, since the voice codec types supported by the two terminals may be different, and the bitrate, sampling rate, etc., supported for the same voice codec may also be different, it is necessary to negotiate the voice codec before establishing the call to determine which voice codec and encoding / decoding configuration the two parties will use for the call. Multiple codecs can include EVS, AMR-WB, and IVAS. The transmission rate can also be called the encoding / decoding rate.
[0088] Negotiation is conducted using the Session Description Protocol (SDP). The negotiation process includes negotiating four important parameters: load type, sampling frequency, data rate, and packet duration. Specifically:
[0089] Load type is a value defined by the standard protocol.
[0090] The higher the sampling frequency, the more sampling points there are, and the closer the speech quality is to the real speech signal.
[0091] Speed is related to bandwidth utilization; the higher the speed, the higher the bandwidth utilization.
[0092] Packet duration refers to the duration of voice data contained in each voice packet. The longer the packet duration, the greater the packet latency, but the stronger the jitter resistance and the higher the bandwidth utilization.
[0093] The parameters involved in the negotiation are: codec type and rate. The sampling frequency and packet duration are fixed for each codec type. During SDP negotiation, the codec type is negotiated first, followed by the rate. For AMR-WB and EVS, adaptive rate is supported, and a supported adaptive range will be selected during negotiation.
[0094] SDP examples are as follows:
[0095] a = rtpmap:100AMR / 8000 / / 100 represents the payload type, AMR represents the codec type, and 8000 represents the sampling frequency.
[0096] a = fmtp:100 mode-set = 0, 2, 4, 7; mode-change-neighbor = 1; mode-change-period = 2 / / mode-set represents the rate set, rate adjustment period, and adjustment method, etc.
[0097] a = ptime:20 / / Packaging time.
[0098] This application first introduces an audio encoding / decoding negotiation method provided in its embodiments. This method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in this application embodiment can specifically refer to audio encoding, audio decoding, or both audio encoding and decoding. Hereafter, "audio encoding / decoding" will refer to the various encoding and decoding processes described above.
[0099] As shown in Figure 4, the negotiation methods for audio encoding / decoding mainly include the following:
[0100] 401. The second terminal device obtains the hardware parameters of the second terminal device, which include at least the battery parameters of the second terminal device.
[0101] In this embodiment, the second terminal device determines its hardware parameters, which may include hardware information from various terminal devices. This embodiment does not limit the types of hardware or hardware configuration parameters involved.
[0102] The hardware parameters of the second terminal device include at least the battery parameters of the second terminal device. Therefore, the second terminal device can negotiate audio codecs based on the battery hardware of the second terminal device. For example, the battery parameters may include the battery capacity, power, temperature, battery operating mode, etc.
[0103] 402. The second terminal device determines the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0104] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device. Thus, the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device. For example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.
[0105] The audio encoding / decoding capability of the second terminal device refers to the capabilities of the second terminal device itself, such as the computing power and memory of the second terminal device.
[0106] For example, audio encoding / decoding configuration information can specifically be an encoding / decoding mode configuration table.
[0107] 403. The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0108] The second terminal device can send its audio encoding / decoding configuration information to the first terminal device, thereby enabling the first terminal device to negotiate audio encoding / decoding.
[0109] Not limited to, in the embodiments of this application, the first terminal device may send audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device may negotiate audio encoding / decoding.
[0110] 411. The first terminal device obtains the audio encoding / decoding configuration information of the second terminal device, which is used to indicate the audio encoding / decoding capability of the second terminal device.
[0111] 412. The first terminal device determines the audio encoding / decoding configuration information of the first terminal device based on the hardware parameters of the first terminal device. The hardware parameters of the first terminal device include at least the battery parameters of the first terminal device.
[0112] In this embodiment, the first terminal device determines its hardware parameters, which may include hardware information from various terminal devices. This embodiment does not limit the types of hardware or hardware configuration parameters involved.
[0113] The hardware parameters of the first terminal device include at least the battery parameters of the first terminal device. Therefore, the first terminal device can negotiate audio codecs based on the battery hardware of the first terminal device. For example, the battery parameters may include the battery capacity, power, temperature, battery operating mode, etc.
[0114] 413. The first terminal device determines the encoding / decoding negotiation result based on the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device. The encoding / decoding negotiation result includes: the first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capability of the second terminal device.
[0115] The first terminal device can obtain the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device respectively. The first terminal device can perform encoding / decoding negotiation to obtain the encoding / decoding negotiation result of the first terminal device.
[0116] The encoding / decoding negotiation result includes: the first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, and the first encoding / decoding parameters being adapted to the audio encoding / decoding capabilities of the second terminal device.
[0117] 414. The first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0118] 404. The first terminal device receives the encoding / decoding negotiation result from the device, the encoding / decoding negotiation result including: the first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0119] In this embodiment, the first terminal device and the second terminal device can perform static audio negotiation based on their respective hardware parameters. For example, negotiation can be based on the hardware capabilities, processor, and bottom-end chip of the terminal device, or on the microphone and memory of the terminal device.
[0120] In some embodiments of this application, the first encoding / decoding parameters include: a first encoding / decoding mode executed by the first terminal device when encoding and decoding the audio signal to be processed, and a first encoding / decoding rate corresponding to the first encoding / decoding mode.
[0121] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0122] When the hardware parameters of the first terminal device meet the preset first hardware configuration conditions, the first audio encoding / decoding configuration information is determined according to the hardware parameters of the first terminal device.
[0123] When the hardware parameters of the first terminal device do not meet the first hardware configuration conditions, the second audio encoding / decoding configuration information is determined according to the hardware parameters of the first terminal device.
[0124] Among them, the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0125] In the embodiments of this application, different hardware parameters of the first terminal device indicate different audio encoding / decoding capabilities, thereby completing audio encoding / decoding negotiation based on hardware parameters.
[0126] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0127] The audio encoding / decoding configuration information of the first terminal device is determined based on the hardware parameters and audio encoding / decoding capabilities of the first terminal device.
[0128] In this embodiment of the application, when the first terminal device determines the audio encoding / decoding configuration information, in addition to using the hardware parameters of the first terminal device, it can also use the audio encoding / decoding capabilities of the first terminal device, thus enabling it to determine the audio encoding / decoding configuration information of the first terminal device more accurately.
[0129] In some embodiments of this application, the hardware parameters also include at least one of the following: parameters of the audio codec chip of the first terminal device, parameters of the audio acquisition module of the first terminal device, parameters of the audio playback module of the first terminal device, and parameters of the off-chip memory of the first terminal device other than the audio codec chip.
[0130] Specifically, the audio acquisition module can be a microphone. The audio acquisition module can be a built-in audio acquisition module of the first terminal device or an external audio acquisition module of the first terminal device. No limitation is made here.
[0131] The audio playback module can be a speaker, and it can be a built-in audio playback module of the first terminal device or an external audio playback module of the first terminal device. There is no limitation here.
[0132] The audio codec chip in the first terminal device refers to the chip that encodes and decodes audio signals.
[0133] The off-chip memory in the first terminal device, excluding the audio codec chip, refers to other memory in the first terminal device besides the audio codec chip. For example, it can be the off-chip memory of the first terminal device, and off-chip memory can also be used as a hardware parameter of the first terminal device.
[0134] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0135] The battery life information of the first terminal device is obtained based on the battery parameters of the first terminal device.
[0136] The audio encoding / decoding configuration information of the first terminal device is updated based on the battery life information to obtain the updated audio encoding / decoding configuration information of the first terminal device.
[0137] In this embodiment of the application, different battery parameters of the first terminal device indicate different battery life information. For example, when the battery life information indicates a higher level of power, a stronger audio encoding / decoding capability can be used, and when the battery life information indicates a lower level of power, the audio encoding / decoding capability can be reduced, thereby completing audio encoding / decoding negotiation based on hardware parameters.
[0138] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0139] The audio signal types and acquisition rates supported by the audio acquisition module of the first terminal device are determined based on the parameters of the audio acquisition module of the first terminal device.
[0140] The audio encoding / decoding configuration information of the first terminal device is determined based on the audio signal types and acquisition rates supported by the audio acquisition module of the first terminal device.
[0141] In this embodiment, different audio acquisition modules of the first terminal device support corresponding audio signal types and acquisition rates, indicating different audio encoding / decoding capabilities, thereby completing audio encoding / decoding negotiation based on hardware parameters.
[0142] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0143] The audio signal type and bit rate supported by the audio codec chip of the first terminal device are determined based on the parameters of the audio codec chip of the first terminal device.
[0144] The audio encoding / decoding configuration information of the first terminal device is determined based on the audio signal types and bit rates supported by the audio codec chip of the first terminal device.
[0145] In this embodiment of the application, different audio codec chips of the first terminal device support corresponding audio signal types and acquisition rates, indicating different audio encoding / decoding capabilities, thereby completing audio codec negotiation based on hardware parameters.
[0146] In some embodiments of this application, determining the audio encoding / decoding configuration information of the first terminal device based on its hardware parameters includes:
[0147] The audio signal type and bit rate supported by the off-chip memory of the first terminal device are determined based on the parameters of the off-chip memory other than the audio codec chip.
[0148] The audio encoding / decoding configuration information of the first terminal device is determined based on the audio signal type and bit rate supported by the off-chip memory of the first terminal device.
[0149] In this embodiment, different off-chip memories of the first terminal device support corresponding audio signal types and acquisition rates, indicating different audio encoding / decoding capabilities, thereby completing audio encoding / decoding negotiation based on hardware parameters.
[0150] In some embodiments of this application, the method further includes:
[0151] Obtain the audio service standard types supported by the audio codec chip of the second terminal device;
[0152] When the audio codec chip of the second terminal device and the audio codec chip of the first terminal device each support the preset first audio service standard type, the aforementioned steps are triggered: obtaining the audio codec configuration information of the second terminal device.
[0153] The audio service standard type can include at least one of the following: EVS / IVAS / AMR, etc. Negotiating the audio service standard type before audio negotiation can achieve more accurate audio negotiation.
[0154] In some embodiments of this application, the audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding modes supported by the second terminal device, and the set of rates corresponding to the encoding / decoding modes supported by the second terminal device;
[0155] The audio encoding / decoding configuration information of the first terminal device includes: the encoding / decoding modes supported by the first terminal device, and the set of rates corresponding to the encoding / decoding modes supported by the first terminal device.
[0156] For example, the encoding / decoding modes can include MONO / STEREO / FOA / MASA, and the bit rate values vary, depending on the corresponding configuration information.
[0157] In some embodiments of this application, determining the encoding / decoding negotiation result of the first terminal device based on the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device includes:
[0158] The audio codec negotiation order is determined according to the audio codec negotiation priority strategy of the first terminal device. The audio codec negotiation order includes: performing audio encoding negotiation first and then performing audio decoding negotiation, or performing audio decoding negotiation first and then performing audio encoding negotiation.
[0159] The audio encoding negotiation includes: determining the encoding negotiation result of the first terminal device based on the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device;
[0160] The audio decoding negotiation includes: determining the decoding negotiation result of the first terminal device based on the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device.
[0161] In this embodiment, the order of audio encoding negotiation and audio decoding negotiation is not limited, which can realize a more flexible audio negotiation method.
[0162] In some embodiments of this application, determining the encoding negotiation result of the first terminal device based on the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device includes:
[0163] Based on the fact that the highest encoding mode supported by the first terminal device is the first audio mode, and the decoding modes supported by the second terminal device include the first audio mode, it is determined that the encoding negotiation result of the first terminal device includes the first audio mode.
[0164] Determine the first bitrate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device;
[0165] The coding negotiation result of the first terminal device is determined to include the maximum code rate in the first code rate intersection.
[0166] In this embodiment, the encoding negotiation result includes the maximum bitrate in the first bitrate intersection. Based on the principle of maximizing experience, the encoding bitrate of the other party's terminal with a higher experience effect is selected first to improve call quality.
[0167] In some embodiments of this application, determining the decoding negotiation result of the first terminal device based on the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device includes:
[0168] Based on the fact that the highest decoding mode supported by the first terminal device is the second audio mode, and the encoding modes supported by the second terminal device include the second audio mode, it is determined that the decoding negotiation result of the first terminal device includes the second audio mode.
[0169] Determine the second bitrate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device;
[0170] The decoding negotiation result of the first terminal device is determined to include the maximum bit rate in the second bit rate intersection.
[0171] In this embodiment of the application, the decoding negotiation result includes the maximum bit rate in the second bit rate intersection. Based on the principle of maximizing experience, the encoding bit rate of the other party terminal with a higher experience effect is selected first to improve call quality.
[0172] In some embodiments of this application, the sum of the bitrate corresponding to the highest encoding mode supported by the first terminal device and the bitrate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum bitrate supported by the audio encoding / decoding capability of the first terminal device.
[0173] In this embodiment, the sum of the bitrate corresponding to the highest encoding mode supported by the first terminal device and the bitrate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum bitrate supported by the audio encoding / decoding capability of the first terminal device. This allows for full utilization of the maximum bitrate of the first terminal device, maximizing the use of the audio encoding / decoding capability of the first terminal device and improving call quality.
[0174] In some embodiments of this application, the first encoding / decoding parameters include at least two of the following: mono MONO, stereo STEREO, metadata-assisted spatial audio MASA, high-order stereo reverb HOA, or first-order stereo reverb FOA.
[0175] The set of rates corresponding to the encoding / decoding modes includes at least two of the following: 9.6kbps, 13.2kbps, 32kbps, and 48kbps.
[0176] In some embodiments of this application, the method further includes:
[0177] The first terminal device sends an audio negotiation request to the second terminal device.
[0178] In this process, the first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains its hardware parameters based on the triggering of the audio negotiation request.
[0179] In some embodiments of this application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device.
[0180] In this process, the first terminal device sends a capability negotiation request to the second terminal device. This request carries the audio codec / decoding configuration information of the first terminal device. The second terminal device can make a preliminary selection based on the received audio codec / decoding configuration information of the first terminal device. The obtained audio codec / decoding configuration information of the second terminal device can be one or more audio codec / decoding configuration information after the initial screening. Then, the first terminal device performs audio negotiation, which further improves the efficiency of audio negotiation.
[0181] Further, it includes: a capability negotiation request sent by the first terminal device to the second terminal device, but may not include the audio encoding / decoding configuration information of the first terminal device; this is not limited here.
[0182] The foregoing embodiments have illustrated the method executed by the first terminal device. The method executed by the second terminal device will be described next. The configuration method of the hardware parameters of the second terminal device is similar to that of the first terminal device. The specific process of determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device will not be described in detail thereafter.
[0183] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on its hardware parameters includes:
[0184] The audio encoding / decoding configuration information of the second terminal device is determined based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device.
[0185] In some embodiments of this application, the hardware parameters also include at least one of the following: parameters of the audio codec chip of the second terminal device, parameters of the audio acquisition module of the second terminal device, parameters of the audio playback module of the first terminal device, and parameters of the off-chip memory of the second terminal device other than the audio codec chip.
[0186] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0187] Obtain the battery life information of the second terminal device based on its battery parameters.
[0188] The audio encoding / decoding configuration information of the second terminal device is updated based on the battery life information and the audio encoding / decoding capability of the second terminal device to obtain the updated audio encoding / decoding configuration information of the second terminal device.
[0189] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0190] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio acquisition module of the second terminal device.
[0191] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device.
[0192] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0193] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0194] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio codec chip of the second terminal device.
[0195] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio codec chip of the second terminal device.
[0196] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0197] In some embodiments of this application, determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0198] The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the off-chip memory other than the audio codec chip in the second terminal device.
[0199] The audio encoding / decoding capability of the second terminal device is determined based on the audio signal type and acquisition rate supported by the off-chip memory in the second terminal device.
[0200] The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
[0201] In some embodiments of this application, the method further includes:
[0202] Obtain the audio service standard types supported by the audio codec chip of the first terminal device;
[0203] When the audio codec chip of the first terminal device and the audio codec chip of the second terminal device each support the preset first audio service standard type, the aforementioned step of obtaining the hardware parameters of the second terminal device is triggered.
[0204] In some embodiments of this application, the audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding modes supported by the second terminal device, and the set of rates corresponding to the encoding / decoding modes supported by the second terminal device;
[0205] In some embodiments of this application, the first encoding / decoding parameters include at least two of the following: mono MONO, stereo STEREO, metadata-assisted spatial audio MASA, high-order stereo reverb HOA, or first-order stereo reverb FOA.
[0206] The set of rates corresponding to the encoding / decoding modes includes at least two of the following: 9.6kbps, 13.2kbps, 32kbps, and 48kbps.
[0207] In some embodiments of this application, the method further includes:
[0208] Receive an audio negotiation request from the first terminal device.
[0209] In some embodiments of this application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device;
[0210] The method further includes: filtering the audio encoding / decoding configuration information of the second terminal device based on the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device;
[0211] Sending audio encoding / decoding configuration information of the second terminal device to the first terminal device includes: sending candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0212] The following will illustrate this with examples of real-world applications.
[0213] The method provided in this application is applicable to IVAS encoding and decoding. Different signal types and corresponding rates can be used, which require further negotiation. This application provides a method for negotiating IVAS encoding and decoding modes and rates.
[0214] Each IVAAS-enabled terminal possesses its own IVAAS codec configuration table, which describes all IVAAS encoding-decoding combinations that the terminal can support within its hardware constraints. By transmitting the codec configuration table during the negotiation process, both parties can more flexibly negotiate call configurations, such as the type of encoded signal and the bit rate. Furthermore, it supports both parties using different signal types and different bit rates. This invention provides the optimal user experience for both parties while satisfying their respective hardware constraints.
[0215] [Revised from Rule 26, October 2025] Each mobile phone supporting IVAS maintains two tables based on its own capabilities (computing power, memory, etc.). One table is the IVAS encoding / decoding mode configuration table, as shown in Table 1 below:
[0216] This IVAS encoding / decoding mode configuration table indicates all possible combinations of IVAS encoding and decoding modes that the phone can support, including the signal type and bit rate for encoding or decoding. As shown in the example in Table 1 above, this phone supports three decoded signal types: MONO, STEREO, and FOA. The maximum bit rate supported for MONO decoding is 9.6kbps, for STEREO decoding it is 32kbps, and for FOA decoding it is 48kbps. When the phone decodes a MONO signal, it can support encoding signal types of MONO, STEREO, and FOA, with maximum bit rates of 9.6kbps, 32kbps, and 48kbps respectively. When the phone decodes a STEREO signal, it still supports three encoded signal types: MONO, STEREO, and FOA. However, because decoding STEREO is more expensive than decoding MONO, the remaining computing power for encoding decreases. Therefore, the encoding bitrate for FOA cannot be supported up to 48kbps; it can only support a maximum of 24kbps (higher bitrates result in greater overhead). When the phone decodes a FOA signal, due to the even greater overhead, the remaining computing power is insufficient to support FOA encoding. At this point, the phone can only support MONO and STEREO, with maximum supported bitrates of 9.6kbps and 32kbps, respectively.
[0217] As can be seen from the above examples, the combination of encoding and decoding modes is limited by the capabilities of the terminal / chip. If the encoding overhead is large, the decoding overhead must be reduced, and vice versa.
[0218] The other table is the encoding / decoding bitrate set table, as shown in Table 2 below:
[0219] The codec rate set table represents all the codec rates supported by IVAS for each signal. By combining the maximum codec rate supported by each encoded / decoded signal in the codec mode configuration table with the codec rate set table, the codec rate range supported by each encoded / decoded signal in the codec mode configuration table can be obtained.
[0220] Every mobile phone comes with a default IVAS codec configuration table when it leaves the factory. However, depending on factors such as the computing power of the phone's chip, the IVAS codec configuration table may differ between different mobile phone platforms.
[0221] Because the types of IVAS signals that a mobile phone can encode are also related to some external factors, such as only supporting FOA / HOA encoding when the mobile phone has FOA or HOA acquisition equipment, only supporting multi-channel or object encoding when the voice service contains multi-channel or object audio / sound effect signals, only supporting full MASA signal encoding when the mobile phone battery is low, and the encoding and decoding of certain high-complexity signals may be limited when the mobile phone battery is low, the mobile phone needs to be able to dynamically update its IVAS encoding and decoding configuration table to reflect the current actual capabilities of the mobile phone.
[0222] Figure 5 illustrates the audio negotiation process between terminal UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) as an example:
[0223] Before a voice call between A and B, if IVAS voice codec is used for communication, media negotiation is required. Media negotiation includes:
[0224] 1. Audio codec negotiation (EVS / IVAS / AMR).
[0225] 2. After the audio codec is negotiated to IVAS, the codec mode (MONO / STEREO / FOA / MASA) and bitrate need to be further negotiated (each codec mode has a corresponding bitrate range). The encoding mode and decoding mode can be negotiated to be different.
[0226] Negotiation process:
[0227] 1. When UE-A initiates an INVITE call, if it supports IVAS voice codec, it sends its supported IVAS codec configuration table to B in the IVAS SDP information, including a list of encoding and decoding mode combinations supported by A, the maximum rate supported by each mode, and the rate set supported by each encoding mode.
[0228] 2. UE-B will handle the following based on its own capabilities:
[0229] 1) First, perform codec negotiation. If IVAS is supported, then initially negotiate to IVAS.
[0230] 2) Continue negotiating the encoding / decoding mode and rate. Based on local policy, decide whether to prioritize negotiation based on your encoding capabilities or your decoding capabilities. For example, let's take prioritizing encoding capabilities:
[0231] ① Determine the encoding mode of UE-B based on the highest encoding mode supported by UE-B and whether UE-A supports the decoding mode;
[0232] ② Determine the UE-B negotiated coding rate based on the intersection of the code rates supported by the calling and called parties for this coding mode. If there is no intersection, the UE-B coding mode needs to be renegotiated;
[0233] ③ Based on the decoding mode corresponding to the list of encoding modes supported by UE-B, and considering whether UE-A supports the encoding mode, determine the decoding mode negotiated by UE-B;
[0234] ④ Determine the decoding bitrate after UE-B negotiation based on the intersection of the bitrates of the calling and called parties in the decoding mode.
[0235] If any one of the codec modes negotiated by IVAS fails, IVAS will not be used in the end.
[0236] As shown in Figures 6 and 7, the SDP example adds the following parameter to the offer a line in IVAS:
[0237] Bitrate set:
[0238] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.
[0239] [mono,5.9 / 7.2 / 9.6] indicates that mono supports three bitrate sets: 5.9, 7.2, and 9.6.
[0240] Decoding mode set:
[0241] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.
[0242] This list indicates that three sets are supported. Each set corresponds one-to-one with the following encoding mode sets. Set 1 corresponds to the decoding mode [mono / 9.6], indicating that mono is supported and the highest bitrate is 9.6.
[0243] Encoding pattern set:
[0244] enc-mode-list={[mono / 9.6,stereo / 32,foa / 48],[mono / 9.6,stereo / 32,foa / 24],[mono / 9.6,stereo / 32]}.
[0245] This list indicates that three sets are supported, which correspond one-to-one with the decoding mode sets above. The encoding mode corresponding to set 1 is = {[mono / 9.6,stereo / 32,foa / 48], which means that the decoding modes are mono / stereo / foa in order from low to high, and the highest bitrates are 9.6 / 32 / 48 in order.
[0246] Add the following parameter to line a of the IVAS answer:
[0247] The negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]};
[0248] The negotiated encoding mode and rate set: enc-mode = {[foa, 24 / 32]}.
[0249] Example of a negotiation process:
[0250] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A, the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0251] 2. Mobile phone B decides whether to negotiate based on its own encoding capabilities or its own decoding capabilities according to its local policy. Taking encoding capability negotiation as an example:
[0252] ① Based on the fact that the highest encoding mode supported by UE-B is FOA, and the decoding mode of UE-A includes FOA, the encoding mode negotiated by UE-B is determined to be FOA;
[0253] ② Determine the intersection of the code rates of the FOA modes supported by the calling and called parties, and determine that the UE-B negotiated code rates are 24 and 32;
[0254] ③ Based on the fact that the decoding mode corresponding to the encoding mode column supported by UE-B is MONO, and the encoding mode of UE-A includes MONO, it is determined that the decoding mode negotiated by UE-B is MONO;
[0255] ④ Determine the intersection of the code rates of the MONO modes supported by the calling and called parties, and determine that the decoding code rates after UE-B negotiation are 5.9, 7.2 and 9.6.
[0256] Figure 8 shows another example of a negotiation process:
[0257] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A, the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0258] 2. Mobile phone B decides whether to negotiate based on its own encoding capabilities or its own decoding capabilities according to its local policy. Taking encoding capability negotiation as an example:
[0259] ① Based on the fact that the highest decoding mode supported by UE-B is FOA, and the decoding mode of UE-A includes FOA, it is determined that the decoding mode negotiated by UE-B is FOA;
[0260] ② Determine the intersection of the code rates of the FOA modes supported by the calling and called parties, and determine that the UE-B negotiated code rates are 24, 32, and 48;
[0261] ③ Based on the fact that the highest encoding mode corresponding to the decoding mode column supported by UE-B is STEREO, and the decoding mode of UE-A includes STEREO, it is determined that the encoding mode negotiated by UE-B is STEREO.
[0262] ④ Determine the intersection of the code rates of the STEREO modes supported by the calling and called parties, and determine that the decoding code rates after UE-B negotiation are 13.2 and 24.
[0263] As illustrated by the foregoing examples, for real-time communication audio codecs with numerous encoding modes and rates, and a wide range of encoding / decoding complexity, each terminal maintains its own encoding / decoding configuration table. This table describes the various combinations of encoding and decoding modes / rates that the terminal can support. When two or more terminals establish a call and negotiate this codec, they select their respective suitable encoding configurations based on their encoding / decoding configuration tables.
[0264] Each terminal has a factory-default codec configuration table. When establishing a call and negotiating the codec, the terminal can update its own codec configuration table based on other factors and negotiate the codec according to the updated codec configuration table. The updated codec configuration table does not overwrite the default factory codec configuration table.
[0265] In the codec configuration table, supported decoding modes / bitrates and / or encoding modes / bitrates are sorted by experience priority.
[0266] During codec negotiation, one terminal sends its complete or partial codec configuration table to the other party for negotiation. The other terminal then makes a negotiation decision based on all or part of the codec configuration tables from at least two parties. Here, "partial encoding mode and rate" refers to a subset of the encoding modes and rates.
[0267] The decision to negotiate can be based on the principle of maximizing the user experience, or it can be based on other principles, such as the decision-making principles preset in the terminal that makes the negotiation decision, such as prioritizing the other party's terminal for coding with a higher user experience.
[0268] The following example illustrates the audio negotiation process for terminal battery level.
[0269] [Revised from Rule 26 to rule 22.10.2025] When the terminal battery is at 100%, the terminal's configuration table is the initial configuration table, as shown in Table 3 below:
[0270] [Revised from Rule 26 to 22.10.2025] When the terminal's battery level drops, in order to ensure the operation of other terminal functions, the configuration table needs to be adjusted and the configuration reduced. For example, when the battery level is only 10%, the configuration table is adjusted as shown in Table 4 below:
[0271] When the terminal uses only configuration 1 (default hardware configuration / factory configuration), the terminal's configuration table is the initial configuration table, as shown in Table 3 above.
[0272] When the terminal uses configuration 2 (external FOA microphone), the configuration table modifies some combinations of Table 3.
[0273] When the terminal uses configuration 3 (external HOA microphone), the configuration table is to replace combination 3 in Table 3 with HOA and the corresponding bit rate.
[0274] When the terminal uses configuration 4 (external microphone), the terminal-side device can convert the microphone signal into a FOA / HOA / MASA signal supported by the codec. The configuration table includes combination 3 as FOA and the corresponding bit rate, combination 4 as HOA and the corresponding bit rate, and combination 5 as MASA and the corresponding bit rate.
[0275] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0276] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.
[0277] Please refer to Figure 9. An embodiment of this application provides a first terminal device 900, which may include: an acquisition module 901, a determination module 902, a negotiation module 903, and a sending module 904.
[0278] The acquisition module is used to acquire the audio encoding / decoding configuration information of the second terminal device, which is used to indicate the audio encoding / decoding capability of the second terminal device.
[0279] The determining module is used to determine the audio encoding / decoding configuration information of the first terminal device based on the hardware parameters of the first terminal device, wherein the hardware parameters of the first terminal device include at least the battery parameters of the first terminal device.
[0280] The negotiation module is used to determine the encoding / decoding negotiation result of the first terminal device based on the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device. The encoding / decoding negotiation result includes: the first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capability of the second terminal device.
[0281] The sending module is used to send the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0282] As illustrated by the foregoing embodiments, the process involves the following steps: First, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device. This configuration information indicates the audio encoding / decoding capabilities of the second terminal device. Then, the first terminal device determines its own audio encoding / decoding configuration information based on its hardware parameters, which include at least its battery parameters. Next, the first terminal device determines its encoding / decoding negotiation result based on both the first and second terminal device's configuration information. This negotiation result includes first encoding / decoding parameters adapted to the audio signal to be processed, and these parameters are compatible with the second terminal device's audio encoding / decoding capabilities. Finally, the first terminal device sends its own encoding / decoding negotiation result to the second terminal device. In this embodiment, both terminals involved in audio negotiation possess their own audio encoding / decoding configuration information, which describes the encoding / decoding capabilities that the terminal can support under hardware parameter constraints. By transmitting audio encoding / decoding configuration information during the negotiation process, the two parties in the call can negotiate the call configuration more flexibly. In this embodiment of the application, the best user experience can be provided to both parties as much as possible while satisfying the hardware parameter constraints of each party.
[0283] Please refer to Figure 10. An embodiment of this application provides a second terminal device 1000, which may include: an acquisition module 1001, a determination module 1002, a sending module 1003, and a receiving module 1004.
[0284] The acquisition module is used to acquire the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device include at least the battery parameters of the second terminal device.
[0285] The determining module is used to determine the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0286] The sending module is used to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device;
[0287] The receiving module is configured to receive an encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0288] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0289] This application also provides a computer storage medium storing a program that performs some or all of the steps described in the above method embodiments.
[0290] Next, we will introduce another first terminal device provided in the embodiments of this application. Please refer to FIG11. The first terminal device 1100 includes:
[0291] The device comprises a receiver 1101, a transmitter 1102, a processor 1103, and a memory 1104 (the number of processors 1103 in the first terminal device 1100 can be one or more; Figure 11 shows an example of one processor). In some embodiments of this application, the receiver 1101, transmitter 1102, processor 1103, and memory 1104 can be connected via a bus or other means; Figure 11 shows an example of connection via a bus.
[0292] Memory 1104 may include read-only memory and random access memory, and provides instructions and data to processor 1103. A portion of memory 1104 may also include non-volatile random access memory (NVRAM). Memory 1104 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0293] Processor 1103 controls the operation of the first terminal device. Processor 1103 can also be called a central processing unit (CPU). In specific applications, the various components of the first terminal device are coupled together through a bus system. In addition to the data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as the bus system in the figure.
[0294] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuits in the hardware of the processor 1103 or by instructions in software form. The processor 1103 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1104. Processor 1103 reads the information in memory 1104 and, in conjunction with its hardware, completes the steps of the above method.
[0295] The receiver 1101 can be used to receive input digital or character information and generate signal inputs related to the settings and function control of the first terminal device. The transmitter 1102 may include a display device such as a display screen and can be used to output digital or character information through an external interface.
[0296] In this embodiment, the processor 1103 is used to execute the method shown in FIG4 of the aforementioned embodiment, which is executed by the first terminal device.
[0297] Next, we will introduce another second terminal device provided in the embodiments of this application. Please refer to FIG12. The second terminal device 1200 includes:
[0298] The device comprises a receiver 1201, a transmitter 1202, a processor 1203, and a memory 1204 (the number of processors 1203 in the second terminal device 1200 can be one or more; Figure 12 shows an example of one processor). In some embodiments of this application, the receiver 1201, transmitter 1202, processor 1203, and memory 1204 can be connected via a bus or other means; Figure 12 shows an example of connection via a bus.
[0299] Memory 1204 may include read-only memory and random access memory, and provides instructions and data to processor 1203. A portion of memory 1204 may also include NVRAM. Memory 1204 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0300] Processor 1203 controls the operation of the second terminal device; processor 1203 can also be referred to as a CPU. In specific applications, the various components of the second terminal device are coupled together through a bus system, which includes not only a data bus but also a power bus, control bus, and status signal bus, etc. However, for clarity, all buses in the diagram are referred to as the bus system.
[0301] The methods disclosed in the embodiments of this application can be applied to processor 1203, or implemented by processor 1203. Processor 1203 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1203 or by instructions in the form of software. The processor 1203 can be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1204, and processor 1203 reads the information in memory 1204 and completes the steps of the above method in combination with its hardware.
[0302] In this embodiment, processor 1203 is used to execute the method shown in FIG7 of the aforementioned embodiment, which is executed by the second terminal device.
[0303] In another possible design, when the first or second terminal device is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer-executable instructions stored in a storage unit to cause the chip within the terminal to perform any of the methods described in the first aspect above. Optionally, the storage unit may be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit may be a storage unit located outside the chip within the terminal, such as read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0304] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of programs in the first or second aspect of the above methods.
[0305] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0306] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0307] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0308] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. An audio encoding / decoding negotiation method, characterized in that, The method is applied to a second terminal device, and the method includes: Obtain the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device include at least the battery parameters of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device. Send the audio encoding / decoding configuration information of the second terminal device to the first terminal device; The first terminal device receives an encoding / decoding negotiation result, which includes: first encoding / decoding parameters for encoding and decoding the audio signal to be processed by the first terminal device, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.
2. The method according to claim 1, characterized in that, The step of determining the audio encoding / decoding configuration information of the second terminal device based on its hardware parameters includes: The audio encoding / decoding configuration information of the second terminal device is determined based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device.
3. The method according to claim 1 or 2, characterized in that, The hardware parameters also include at least one of the following: parameters of the audio codec chip of the second terminal device, parameters of the audio acquisition module of the second terminal device, parameters of the audio playback module of the first terminal device, and parameters of the off-chip memory of the second terminal device other than the audio codec chip.
4. The method according to claim 2, characterized in that, The step of determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes: The battery life information of the second terminal device is obtained based on the battery parameters of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is updated based on the battery life information and the audio encoding / decoding capability of the second terminal device to obtain the updated audio encoding / decoding configuration information of the second terminal device.
5. The method according to claim 2 or 3, characterized in that, The step of determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes: The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio acquisition module of the second terminal device. The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
6. The method according to claim 2 or 3, characterized in that, The step of determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes: The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the audio codec chip of the second terminal device. The audio encoding / decoding capability of the second terminal device is determined based on the audio signal types and acquisition rates supported by the audio codec chip of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
7. The method according to claim 2 or 3, characterized in that, The step of determining the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters and audio encoding / decoding capabilities of the second terminal device includes: The audio signal types and acquisition rates supported by the audio acquisition module of the second terminal device are determined based on the parameters of the off-chip memory (excluding the audio codec chip) in the second terminal device. The audio encoding / decoding capability of the second terminal device is determined based on the audio signal type and acquisition rate supported by the off-chip memory in the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined based on its audio encoding / decoding capabilities.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtain the audio service standard types supported by the audio codec chip of the first terminal device; When the audio codec chip of the first terminal device and the audio codec chip of the second terminal device each support the preset first audio service standard type, the aforementioned step of obtaining the hardware parameters of the second terminal device is triggered.
9. The method according to any one of claims 1 to 8, characterized in that, The audio encoding / decoding configuration information of the second terminal device includes: the encoding / decoding modes supported by the second terminal device, and the set of rates corresponding to the encoding / decoding modes supported by the second terminal device.
10. The method according to any one of claims 1 to 10, characterized in that, The first encoding / decoding parameters include at least two of the following: mono MONO, stereo STEREO, metadata-assisted spatial audio MASA, high-order stereo reverb HOA, or first-order stereo reverb FOA; The set of rates corresponding to the encoding / decoding modes includes at least two of the following: 9.6kbps, 13.2kbps, 32kbps, and 48kbps.
11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Receive an audio negotiation request from the first terminal device.
12. The method according to claim 11, characterized in that, The audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device; The method further includes: filtering the audio encoding / decoding configuration information of the second terminal device based on the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device; Sending the audio encoding / decoding configuration information of the second terminal device to the first terminal device includes: sending the candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
13. A terminal device, characterized in that, The terminal device is specifically a second terminal device, and the method includes: The acquisition module is used to acquire the hardware parameters of the second terminal device, wherein the hardware parameters of the second terminal device include at least the battery parameters of the second terminal device. The determining module is used to determine the audio encoding / decoding configuration information of the second terminal device based on the hardware parameters of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device. The sending module is used to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device; The receiving module is configured to receive an encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result including: first encoding / decoding parameters of the first terminal device for encoding and decoding the audio signal to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
14. A terminal device, characterized in that, The terminal device includes at least one processor, which is coupled to a memory to read and execute instructions in the memory to implement the method as described in any one of claims 1 to 10.
15. The terminal device according to claim 14, characterized in that, The terminal device also includes the memory.
16. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 12.