Negotiation method for audio encoding / decoding and terminal device

By obtaining the mode control parameters and user description information of the terminal device, flexible negotiation of audio codec configuration is carried out, which solves the problems of poor codec effect and capacity waste caused by improper selection of codec level in the existing technology, and achieves better user experience and resource utilization.

WO2025200769A1PCT designated stage Publication Date: 2025-10-02HUAWEI TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/075103
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-01-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing audio codec negotiation methods are unable to select the optimal codec level when supporting more complex immersive voice codecs, resulting in a stable call connection but poor codec performance and a waste of codec capacity.

Method used

By obtaining the mode control parameters of the terminal device, including user description parameters, the audio codec configuration information is determined and transmitted during the negotiation process, so that both terminals can flexibly negotiate based on their respective audio codec capabilities and select the optimal codec configuration.

Benefits of technology

This achieves the goal of providing both parties with the best user experience while satisfying the constraints of their respective mode control parameters, avoiding the waste of encoding and decoding capabilities and improving the audio negotiation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075103_02102025_PF_FP_ABST
    Figure CN2025075103_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a negotiation method for audio encoding / decoding, the method being applied to a second terminal device, and the method comprising: acquiring a mode control parameter of the second terminal device, the mode control parameter of the second terminal device at least comprising: a user description parameter generated by the second terminal device; on the basis of the mode control parameter of the second terminal device, determining audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used for indicating an audio encoding / decoding capability of the second terminal device; sending to a first terminal device the audio encoding / decoding configuration information of the second terminal device; and receiving an encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result comprising: a first encoding / decoding parameter for the first terminal device encoding / decoding an audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

Audio encoding / decoding negotiation method and terminal equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 26, 2024, with application number 202410361587.2 and invention name “A negotiation method and terminal device for audio encoding / decoding”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of audio coding and decoding, and in particular to an audio coding / decoding negotiation method and terminal device. Background Art

[0003] Compared to monophonic speech codecs, immersive speech codecs support more audio signal types, a wider bitrate range, and, in addition to coding and decoding features, also include rendering features. Taking the audio codec standards currently under development by the 3rd Generation Partnership Project (3GPP) as an example, different audio codec standards support a variety of signal types, and different types of signals support different numbers of channels and bitrates. The rendering features of different signal types support rendering of various signal types to binaural and standard speaker arrays. Because immersive speech codecs support more signal channels, a wider bitrate range, and additional rendering features, the maximum computational complexity and storage complexity of immersive speech codecs have increased significantly compared to traditional speech codecs.

[0004] To support more complex voice codecs, an existing audio codec negotiation method categorizes voice codecs by complexity. However, when two terminals supporting different complexity levels establish a call, the negotiation results in the selection of the lower complexity level to ensure a stable connection. While this method meets the voice requirements for a call connection, it suffers from poor audio codec negotiation performance. Summary of the Invention

[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides an audio codec negotiation method, characterized in that the method is applied to a second terminal device, and the method includes:

[0007] Acquire a mode control parameter of the second terminal device, where the mode control parameter of the second terminal device at least includes: a user description parameter generated by the second terminal device;

[0008] Determining audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device;

[0009] Sending audio encoding / decoding configuration information of the second terminal device to the first terminal device;

[0010] Receive a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result including: first coding / decoding parameters for the first terminal device to encode and decode the audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.

[0011] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control of the second terminal device includes:

[0012] The audio encoding / decoding configuration information of the second terminal device is determined according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.

[0013] In some embodiments of the present application, the mode control parameters of the second terminal device include at least one of the following: the battery life mode of the second terminal device, the audio call user experience of the second terminal device, and the user's power allocation mode for the second terminal device.

[0014] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:

[0015] Acquiring battery life information of the second terminal device according to the battery parameters of the second terminal device;

[0016] The audio encoding / decoding configuration information of the second terminal device is updated according to the battery life information and the audio encoding / decoding capability of the second terminal device to obtain updated audio encoding / decoding configuration information of the second terminal device.

[0017] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0018] When the mode control parameter of the second terminal device meets the preset first software configuration condition, determining the first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;

[0019] When the mode control parameter of the second terminal device does not meet the preset first software configuration condition, determining the second audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;

[0020] The audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.

[0021] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0022] The mode control parameters of the second terminal device include the first battery power collected by the second terminal device with the permission of the user. The audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power collected by the second terminal device with the permission of the user to obtain the updated audio encoding / decoding configuration information of the second terminal device.

[0023] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes:

[0024] Determine a first power range to which the power of the second battery belongs;

[0025] Determining a first audio codec mode corresponding to the first power interval according to a correspondence between the power interval and the audio codec mode of the second terminal device;

[0026] Adjust the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.

[0027] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes:

[0028] When the power level of the first battery is lower than a preset first power threshold and the second terminal device is in a battery charging mode, the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device is maintained or increased.

[0029] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0030] The mode control parameters of the second terminal device include user habit parameters collected by the second terminal device with user permission, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameters collected by the second terminal device with user permission.

[0031] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device based on the user habit parameters collected by the second terminal device with the user's permission includes:

[0032] determining a first estimated standby time of the second terminal device according to the user schedule information of the second terminal device;

[0033] Determining a second audio codec mode corresponding to the first estimated standby time according to a correspondence between the standby time and the audio codec mode of the second terminal device;

[0034] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.

[0035] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device based on the user habit parameters collected by the second terminal device with the user's permission includes:

[0036] determining terminal capability allocation information according to user operation habit information of the second terminal device;

[0037] Determining a third audio codec mode according to a correspondence between the terminal capability allocation and the audio codec mode of the second terminal device;

[0038] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.

[0039] In some embodiments of the present application, the method further includes:

[0040] Obtaining a network packet loss rate of the audio call performed by the second terminal device;

[0041] When the network packet loss rate is higher than a first network signal threshold, the audio encoding / decoding configuration information of the second terminal device is updated.

[0042] In some embodiments of the present application, the method further comprises:

[0043] Receive an audio negotiation request from the first terminal device.

[0044] In some embodiments of the present application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device;

[0045] The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device;

[0046] The sending the audio codec configuration information of the second terminal device to the first terminal device includes: sending candidate audio codec configuration information of the second terminal device to the first terminal device.

[0047] In a second aspect, an embodiment of the present application provides a terminal device, wherein the terminal device is specifically a second terminal device, and the method includes:

[0048] An acquisition module is used for mode control parameters of the second terminal device, where the mode control parameters of the second terminal device at least include: user description parameters generated by the first terminal device;

[0049] a determining module, configured to determine audio encoding / decoding configuration information of the second terminal device according to a mode control parameter of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate an audio encoding / decoding capability of the second terminal device;

[0050] A sending module, configured to send audio encoding / decoding configuration information of the second terminal device to the first terminal device;

[0051] A receiving module is used to receive the encoding / decoding negotiation result from the first terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameters for the first terminal device to encode and decode the audio signal to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.

[0052] In the second aspect of the present application, the constituent modules of the first terminal device may also execute the steps described in the aforementioned first aspect and various possible implementations. For details, please refer to the aforementioned description of the first aspect and various possible implementations.

[0053] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first aspect above.

[0054] In a fourth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect above.

[0055] In the fifth aspect, an embodiment of the present application provides a communication device, which may include entities such as a terminal device or a chip, and the communication device includes: a processor, a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the communication device performs a method as described in any one of the above-mentioned first aspects.

[0056] In a sixth aspect, the present application provides a chip system, which includes a processor for supporting a terminal device to implement the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the terminal device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0057] In the seventh aspect, an embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors; the interface circuit is used to receive signals from a memory of an electronic device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method in the first aspect or any possible implementation of the first aspect.

[0058] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0059] First, the first terminal device obtains the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameters of the first terminal device, and the mode control parameters of the first terminal device at least include: the user description parameters generated by the first terminal device; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameter for encoding and decoding the processed audio signal by the first terminal device, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiment of the present application, both terminals performing audio negotiation each have audio encoding / decoding configuration information, and the audio encoding / decoding configuration information describes the encoding and decoding capability that the terminal can support under the constraints of the mode control parameters. By transmitting audio encoding / decoding configuration information during the negotiation process, the two parties of the call can negotiate the call configuration more flexibly. In the embodiment of the present application, the best user experience can be provided to both parties as much as possible while satisfying the respective mode control parameter constraints of the two parties of the call. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] FIG1 is a schematic diagram of the structure of an audio processing system provided in an embodiment of the present application;

[0061] FIG2a is a schematic diagram of an audio encoder and an audio decoder provided in an embodiment of the present application applied to a terminal device;

[0062] FIG2 b is a schematic diagram of an audio encoder provided by an embodiment of the present application applied to a wireless device or a core network device;

[0063] FIG2c is a schematic diagram of an audio decoder provided in an embodiment of the present application applied to a wireless device or a core network device;

[0064] FIG3a is a schematic diagram of a multi-channel encoder and a multi-channel decoder provided in an embodiment of the present application applied to a terminal device;

[0065] FIG3 b is a schematic diagram of a multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device;

[0066] FIG3c is a schematic diagram of a multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device;

[0067] FIG4 is a schematic diagram of an audio encoding / decoding negotiation method provided in an embodiment of the present application;

[0068] FIG5 is a schematic diagram of a negotiation process in which two terminals perform audio negotiation via SDP according to an embodiment of the present application;

[0069] FIG6 is a schematic diagram of a negotiation process between two terminals on an audio codec mode and a bit rate according to an embodiment of the present application;

[0070] FIG7 is a schematic diagram of an encoding / decoding negotiation result generated after negotiation between two terminals according to an embodiment of the present application;

[0071] FIG8 is a schematic diagram of another negotiation process between two terminals on audio codec mode and bit rate according to an embodiment of the present application;

[0072] FIG9 is a schematic diagram of the composition structure of a first terminal device provided in an embodiment of the present application;

[0073] FIG10 is a schematic diagram of the structure of a second terminal device provided in an embodiment of the present application;

[0074] FIG11 is a schematic diagram of the composition structure of another first terminal device provided in an embodiment of the present application;

[0075] FIG12 is a schematic diagram of the composition structure of another second terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0076] The embodiments of the present application are described below with reference to the accompanying drawings.

[0077] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0078] Traditional monophonic speech codecs (codecs) have bit rates ranging from a few kbps to tens or 128 kbps. Immersive speech codecs (codecs) support more signal types and a wider bit rate range than monophonic speech codecs. In addition to encoding and decoding features, they also include rendering features. For example, the 3GPP's ongoing Immersive Voice and Audio Services (IVAS) speech / audio codec standard supports mono, stereo, multi-channel, multi-object, higher-order ambisonics (HOA) or first-order ambisonics (FOA) signals, and MASA. Multi-channel signals support formats up to 7.1.4, or 12 channels, while HOA signals support up to 3rd-order HOA, or 16 channels. Bit rates range from a low of 5.9kbps for a single channel to 768kbps for 3rd-order HOA. Rendering supports a variety of signal types, including binaural and standard speaker arrays. The increased number of signal channels, wider bitrate range, and additional rendering features significantly increase the maximum computational and storage complexity of immersive voice codecs compared to traditional voice codecs, significantly increasing the difficulty of implementing chips for these codecs.

[0079] To address the high complexity of voice codecs, voice codecs are divided into different levels according to their complexity. However, when two terminals supporting different levels establish a call, the current negotiation result is that only the one with the lower complexity level can be selected to ensure a stable call connection. Although this method meets the voice requirements of the call connection, it suffers from the problem of poor audio codec negotiation. In addition, although this method meets the voice requirements of the call connection, the lower-level terminal cannot decode all types of bitstreams, which imposes codec restrictions on the audio signal of the higher-level terminal, resulting in poor audio signal codec quality. In addition, when the levels of the two terminals are unequal, the higher-level terminal also suffers from the problem of wasted codec capacity.

[0080] Based on the above description, an embodiment of the present application provides an audio codec negotiation method, which can be used to perform precise Codec negotiation when establishing a call, and select the most appropriate codec combination for each communication party based on the actual capabilities of the platforms of the two parties, thereby ensuring the optimal call experience and not wasting power consumption.

[0081] An embodiment of the present application provides an audio coding technology, in particular, a three-dimensional audio coding technology for three-dimensional audio signals, and specifically provides a coding technology that uses fewer channels to represent three-dimensional audio signals to improve traditional audio coding systems. Audio coding (or commonly referred to as coding) includes two parts: audio coding and audio decoding. Audio coding is performed on the source side, including processing (for example, compressing) the original audio to reduce the amount of data required to represent the audio, thereby more efficiently storing and / or transmitting. Audio decoding is performed on the destination side, including performing inverse processing relative to the encoder to reconstruct the original audio. The coding part and the decoding part are also collectively referred to as coding. The implementation of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0082] The technical solutions of the embodiments of the present application can be applied to various audio processing systems. As shown in Figure 1, a schematic diagram of the structure of the audio processing system provided by the embodiments of the present application is shown. The audio processing system 100 may include: an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 may be used to generate a code stream, which may then be transmitted to the audio decoding device 102 via an audio transmission channel. The audio decoding device 102 may receive the code stream and then perform the audio decoding function of the audio decoding device 102 to obtain a reconstructed signal.

[0083] In an embodiment of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio encoding device can be an audio encoder of the above-mentioned terminal device or wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio decoding device can be an audio decoder of the above-mentioned terminal device or wireless device or core network device. For example, the audio encoder can include a wireless access network, a media gateway of the core network, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The audio encoder can also be an audio encoder used in a virtual reality (VR) streaming service.

[0084] In the application embodiment, taking the audio encoding module (audio encoding and audio decoding) suitable for virtual reality streaming (VR streaming) services as an example, the end-to-end audio signal processing flow includes: audio signal A passes through the acquisition module (acquisition) and then undergoes preprocessing (audio PReprocessing). The preprocessing operation includes filtering out the low-frequency part of the signal, which can be divided by 20Hz or 50Hz, extracting the azimuth information in the signal, and then performing encoding processing (audio encoding) and packaging (file / segment encapsulation) before sending (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), and performing binaural rendering (audio rendering) on ​​the decoded signal. The rendered signal is mapped to the listener's headphones (headphones), which can be independent headphones or headphones on a glasses device.

[0085] As shown in Figure 2a, a schematic diagram of the audio encoder and audio decoder provided in an embodiment of the present application applied to a terminal device. Each terminal device may include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used to channel encode the audio signal, and the channel decoder is used to channel decode the audio signal. For example, the first terminal device 20 may include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 may include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22. The first network communication device 22 and the wireless or wired second network communication device 23 are connected via a digital channel. The second terminal device 21 is connected to the wireless or wired second network communication device 23. The above-mentioned wireless or wired network communication device can generally refer to signal transmission equipment, such as a communication base station, data exchange equipment, etc.

[0086] In audio communication, the transmitting terminal device first captures the audio signal, encodes it, and then performs channel coding before transmitting it over a digital channel via a wireless network or core network. The receiving terminal device then performs channel decoding on the received signal to obtain a bitstream. Audio decoding then restores the audio signal, which is then played back by the receiving terminal device.

[0087] As shown in Figure 2b, it is a schematic diagram of the audio encoder provided in an embodiment of the present application applied to a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, another audio decoder 252, an audio encoder 253 provided in an embodiment of the present application, and a channel encoder 254. The other audio decoder 252 refers to an audio decoder other than the audio decoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, then audio decoding is performed using the other audio decoder 252, then audio encoding is performed using the audio encoder 253 provided in an embodiment of the present application, and finally, channel encoding is performed using the channel encoder 254. After channel encoding is completed, the audio signal is transmitted. The other audio decoder 252 performs audio decoding on the bitstream decoded by the channel decoder 251.

[0088] Figure 2c shows a schematic diagram of an audio decoder provided in an embodiment of the present application being applied to a wireless device or core network device. The wireless device or core network device 25 includes a channel decoder 251, an audio decoder 255 provided in an embodiment of the present application, another audio encoder 256, and a channel encoder 254. The other audio encoder 256 refers to an audio encoder other than the audio encoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251. The received audio code stream is then decoded using the audio decoder 255. The other audio encoder 256 then performs audio encoding. Finally, the audio signal is channel-encoded using the channel encoder 254. After channel encoding is complete, the audio signal is transmitted. If transcoding is required within the wireless device or core network device, corresponding audio encoding processing is required. A wireless device refers to a radio frequency-related device in communications, and a core network device refers to a core network-related device in communications.

[0089] In some embodiments of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices with transcoding needs, and core network devices. For example, the audio encoding device can be a multi-channel encoder of the above-mentioned terminal devices, wireless devices, or core network devices. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices with transcoding needs, and core network devices. For example, the audio decoding device can be a multi-channel decoder of the above-mentioned terminal devices, wireless devices, or core network devices.

[0090] Figure 3a shows a schematic diagram of a multi-channel encoder and multi-channel decoder provided in an embodiment of the present application applied to a terminal device. Each terminal device may include a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can implement the audio encoding method provided in an embodiment of the present application, and the multi-channel decoder can implement the audio decoding method provided in an embodiment of the present application. Specifically, the channel encoder is used to perform channel encoding on a multi-channel signal, and the channel decoder is used to perform channel decoding on a multi-channel signal. For example, a first terminal device 30 may include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. A second terminal device 31 may include a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a first wireless or wired network communication device 32. The first network communication device 32 and a second wireless or wired network communication device 33 are connected via a digital channel. The second terminal device 31 is connected to the second wireless or wired network communication device 33. The wireless or wired network communication equipment mentioned above generally refers to signal transmission equipment, such as communication base stations and data exchange equipment. In audio communication, the transmitting terminal device performs multi-channel encoding on the collected multi-channel signal, then performs channel coding, and then transmits the signal over a digital channel via a wireless network or core network. The receiving terminal device then performs channel decoding on the received signal to obtain a multi-channel signal encoded stream. Multi-channel decoding then recovers the multi-channel signal, which is then played back by the receiving terminal device.

[0091] As shown in Figure 3b, it is a schematic diagram of the multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, which is similar to the aforementioned Figure 2b and will not be repeated here.

[0092] As shown in Figure 3c, it is a schematic diagram of the multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, which is similar to the aforementioned Figure 2c and will not be repeated here.

[0093] Among them, the audio encoding process can be a part of the multi-channel encoder, and the audio decoding process can be a part of the multi-channel decoder. For example, multi-channel encoding of the collected multi-channel signal can be to obtain an audio signal after processing the collected multi-channel signal, and then encode the obtained audio signal according to the method provided in the embodiment of the present application; the decoding end encodes the code stream according to the multi-channel signal, decodes the audio signal, and restores the multi-channel signal after upmixing. Therefore, the embodiment of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding is required, corresponding multi-channel encoding processing is required.

[0094] The embodiments of the present application involve two or more terminals that need to conduct a call, for example, they may be terminals of both parties conducting the call, or they may be terminals of multiple parties. The subsequent embodiments will be described using a two-party call as an example, such as a calling terminal and a called terminal, wherein when the calling terminal may be the aforementioned encoding terminal, the called terminal may be the decoding terminal, and when the called terminal may be the aforementioned encoding terminal, the calling terminal may be the decoding terminal. When two parties establish a call in a mobile network, since the voice codec types supported by the two terminals may be different and the bit rates and sampling rates supported for the same voice codec may also be different, it is necessary to negotiate the voice codec before the call is established to determine which voice codec and which codec configuration the two parties will use to conduct the call, wherein the multiple codecs may include EVS, AMR-WB, and IVAS. The transmission rate can also be called the codec rate.

[0095] The negotiation is carried out using the Session Description Protocol (SDP). The negotiation process includes the negotiation of four important parameters: load type, sampling frequency, rate, and packetization duration. Specifically:

[0096] The payload type is a value defined by the standard protocol.

[0097] The higher the sampling frequency, the more sampling points there are, and the closer the voice quality is to the real voice signal.

[0098] The rate is related to the bandwidth utilization rate. The higher the rate, the higher the bandwidth utilization rate.

[0099] The packetization duration refers to the duration of speech contained in each voice packet. A larger packetization duration increases packetization latency, but improves jitter resistance and bandwidth utilization.

[0100] The parameters involved in the negotiation include the codec type and rate. The sampling frequency and packetization duration are fixed for each codec type. During SDP negotiation, the codec type is negotiated first, followed by the rate. For AMR-WB and EVS, which support adaptive rate, the supported adaptive range is selected during negotiation.

[0101] The following is an example of an SDP:

[0102] a=rtpmap:100AMR / 8000 / / 100 is the payload type, AMR indicates the codec type, and 8000 indicates the sampling frequency

[0103] a=fmtp:100mode-set=0,2,4,7;mode-change-neighbor=1;mode-change-period=2 / / mode-set indicates the rate set, rate adjustment period and adjustment method, etc.

[0104] a=ptime:20 / / packaging duration.

[0105] First, an audio encoding / decoding negotiation method provided in an embodiment of the present application is introduced. The method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in the embodiment of the present application can specifically refer to audio encoding, or audio decoding, or audio encoding and decoding. Subsequently, "audio encoding / decoding" will refer to the various encoding and decoding processes mentioned above.

[0106] As shown in Figure 4, the audio codec negotiation method mainly includes the following:

[0107] 401. A second terminal device obtains a mode control parameter of the second terminal device, where the mode control parameter of the second terminal device at least includes a user description parameter generated by the first terminal device.

[0108] In the embodiments of the present application, the second terminal device determines the mode control parameters of the second terminal device. The mode control parameters of the second terminal device may include user description parameters generated by the terminal device. The terminal device generates the user description parameters based on historical information about the user's terminal operations, with the user's permission. In the embodiments of the present application, the types and configuration parameters of the user-set parameters involved are not limited.

[0109] 402. The second terminal device determines audio encoding / decoding configuration information of the second terminal device according to a mode control parameter of the second terminal device. The audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.

[0110] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device based on the mode control parameters of the second terminal device, so that the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device. For example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.

[0111] The audio encoding / decoding capability of the second terminal device refers to the capability of the second terminal device itself. For example, the audio encoding / decoding capability may include computing power, memory, etc. of the second terminal device.

[0112] For example, the audio encoding / decoding configuration information may specifically be a codec mode configuration table.

[0113] 403. The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.

[0114] The second terminal device may send the audio codec configuration information of the second terminal device to the first terminal device, so that the first terminal device may perform audio codec negotiation.

[0115] It is not limited to that, in the embodiment of the present application, the first terminal device may also send the audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device may negotiate the audio encoding / decoding.

[0116] 411. The first terminal device obtains audio encoding / decoding configuration information of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.

[0117] 412. The first terminal device determines audio encoding / decoding configuration information of the first terminal device according to a mode control parameter of the first terminal device. The mode control parameter of the first terminal device includes at least a user description parameter generated by the first terminal device.

[0118] In the embodiment of the present application, the first terminal device determines the mode control parameters of the first terminal device, and the mode control parameters of the first terminal device may include hardware information of various terminal devices. In the embodiment of the present application, the types of hardware involved and the hardware configuration parameters are not limited.

[0119] 413. The first terminal device determines the encoding / decoding negotiation result of the first terminal device based on the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device. The encoding / decoding negotiation result includes: first encoding / decoding parameters for encoding and decoding the audio signal to be processed by the first terminal device, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.

[0120] The first terminal device can respectively obtain the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the first terminal device can perform encoding / decoding negotiation to obtain the encoding / decoding negotiation result of the first terminal device.

[0121] The encoding / decoding negotiation result includes: first encoding / decoding parameters for encoding and decoding the audio signal to be processed by the first terminal device, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capability of the second terminal device.

[0122] 414. The first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device.

[0123] 404. The first terminal device receives a coding / decoding negotiation result from the first terminal device, where the coding / decoding negotiation result includes: first coding / decoding parameters for coding / decoding the audio signal to be processed by the first terminal device, where the first coding / decoding parameters are adapted to the audio coding / decoding capability of the second terminal device.

[0124] In the embodiment of the present application, the first terminal device and the second terminal device can perform static audio negotiation based on their respective mode control parameters, for example, based on the hardware capabilities, processors, and bottom-end chips of the terminal devices, or based on the microphones and memories of the terminal devices.

[0125] In some embodiments of the present application, the first encoding / decoding parameter includes: a first encoding / decoding mode executed by the first terminal device when encoding and decoding the audio signal to be processed, and a first encoding / decoding rate corresponding to the first encoding / decoding mode.

[0126] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device includes:

[0127] When the mode control parameter of the first terminal device meets the preset first hardware configuration condition, determining the first audio encoding / decoding configuration information according to the mode control parameter of the first terminal device;

[0128] When the mode control parameter of the first terminal device does not meet the first hardware configuration condition, determining the second audio encoding / decoding configuration information according to the mode control parameter of the first terminal device;

[0129] The audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.

[0130] In the embodiment of the present application, different mode control parameters of the first terminal device indicate different audio encoding / decoding capabilities, thereby completing audio codec negotiation based on the mode control parameters.

[0131] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device includes:

[0132] The audio encoding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device and the audio encoding / decoding capability of the first terminal device.

[0133] In an embodiment of the present application, when the first terminal device determines the audio encoding / decoding configuration information, in addition to using the mode control parameters of the first terminal device, the audio encoding / decoding capabilities of the first terminal device can also be used, thereby being able to more accurately determine the audio encoding / decoding configuration information of the first terminal device.

[0134] In some embodiments of the present application, the mode control parameters also include at least one of the following: the battery life mode of the first terminal device, the audio call user experience of the first terminal device, and the user's power allocation mode for the first terminal device.

[0135] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device includes:

[0136] The mode control parameters of the first terminal device include the first battery power collected by the first terminal device with the permission of the user. The audio encoding / decoding configuration information of the first terminal device is updated according to the first battery power collected by the first terminal device with the permission of the user to obtain the updated audio encoding / decoding configuration information of the first terminal device.

[0137] In the embodiment of the present application, different battery parameters of the first terminal device indicate different battery life information. For example, when the battery life information indicates more power, a stronger audio encoding / decoding capability can be used. When the power indicated by the battery life information decreases, the audio encoding / decoding capability can be reduced, thereby completing the audio codec negotiation based on the mode control parameters.

[0138] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device with the permission of the user includes:

[0139] determining a first power range to which the power of the first battery belongs;

[0140] Determining a first audio codec mode corresponding to the first power interval according to a correspondence between the power interval and the audio codec mode of the first terminal device;

[0141] The audio encoding / decoding configuration information of the first terminal device is adjusted according to the audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.

[0142] Among them, in the embodiment of the present application, the first terminal device indicates different audio encoding / decoding capabilities according to the correspondence between the power interval and the audio codec mode of the first terminal device, thereby completing the audio codec negotiation based on the mode control parameters.

[0143] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device with the permission of the user includes:

[0144] When the power level of the first battery is lower than a preset first power threshold and the first terminal device is in a battery charging mode, the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device is maintained or increased.

[0145] Among them, in the embodiment of the present application, the power of the first battery is lower than the preset first power threshold, and the first terminal device is in battery charging mode, thereby completing the audio codec negotiation based on the mode control parameters.

[0146] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device includes:

[0147] The mode control parameters of the first terminal device include user habit parameters collected by the first terminal device with user permission, and the audio encoding / decoding configuration information of the first terminal device is determined based on the user habit parameters collected by the first terminal device with user permission.

[0148] Among them, the user habit parameters set by the first terminal device of the first terminal device in the embodiment of the present application indicate different audio encoding / decoding capabilities, thereby completing the audio encoding and decoding negotiation based on the mode control parameters.

[0149] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device based on the user habit parameters collected by the first terminal device with the permission of the user includes:

[0150] determining a first estimated standby time of the first terminal device according to the user schedule information of the first terminal device;

[0151] Determining a second audio codec mode corresponding to the first estimated standby time according to a correspondence between the standby time and the audio codec mode of the first terminal device;

[0152] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.

[0153] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the first terminal device based on the user habit parameters collected by the first terminal device with the permission of the user includes:

[0154] determining terminal capability allocation information according to user operation habit information of the first terminal device;

[0155] Determining a third audio codec mode according to a correspondence between terminal capability allocation and the audio codec mode of the first terminal device;

[0156] The audio encoding / decoding configuration information of the first terminal device is determined according to the audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.

[0157] In some embodiments of the present application, the method further includes:

[0158] Obtaining a network packet loss rate of the audio call performed by the first terminal device;

[0159] When the network packet loss rate is higher than the first network signal threshold, the audio encoding / decoding configuration information of the first terminal device is updated

[0160] In some embodiments of the present application, the method further comprises:

[0161] Obtaining the audio service standard type supported by the audio codec chip of the second terminal device;

[0162] When the audio service standard types supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device are both the preset first audio service standard types, the aforementioned step is triggered to execute: obtaining the audio codec configuration information of the second terminal device.

[0163] The audio service standard type may include at least one of the following: EVS / IVAS / AMR, etc. Before the audio negotiation, the audio service standard type is negotiated first, which can achieve more accurate audio negotiation.

[0164] In some embodiments of the present application, the audio codec configuration information of the second terminal device includes: a codec mode supported by the second terminal device, and a rate set corresponding to the codec mode supported by the second terminal device;

[0165] The audio encoding / decoding configuration information of the first terminal device includes: the encoding / decoding mode supported by the first terminal device, and the rate set corresponding to the encoding / decoding mode supported by the first terminal device.

[0166] For example, the codec mode may include MONO / STEREO / FOA / MASA, and the rate values ​​may vary, depending on the corresponding configuration information.

[0167] In some embodiments of the present application, determining the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device includes:

[0168] Determining an audio codec negotiation order according to the audio codec negotiation priority policy of the first terminal device, the audio codec negotiation order including: performing audio coding negotiation followed by audio decoding negotiation, or performing audio decoding negotiation followed by audio coding negotiation;

[0169] The performing of the audio coding negotiation includes: determining a coding negotiation result of the first terminal device according to the audio coding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device;

[0170] Performing audio decoding negotiation includes: determining a decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device.

[0171] In the embodiment of the present application, there is no limitation on the order of audio encoding negotiation and audio decoding negotiation, which can achieve a more flexible audio negotiation method.

[0172] In some embodiments of the present application, determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device includes:

[0173] According to the highest encoding mode supported by the first terminal device being the first audio mode, and the decoding modes supported by the second terminal device including the first audio mode, determining that the encoding negotiation result of the first terminal device includes the first audio mode;

[0174] Determine a first bit rate intersection between a first audio mode supported by the first terminal device and a first audio mode supported by the second terminal device;

[0175] Determine that the coding negotiation result of the first terminal device includes the maximum code rate in the first code rate intersection.

[0176] In the embodiment of the present application, the coding negotiation result includes the maximum bit rate in the first bit rate intersection. According to the principle of maximizing experience, the coding bit rate with a higher experience effect of the other terminal is preferentially selected to improve the call quality.

[0177] In some embodiments of the present application, determining a decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device includes:

[0178] According to the highest decoding mode supported by the first terminal device being the second audio mode, and the encoding modes supported by the second terminal device including the second audio mode, determining that the decoding negotiation result of the first terminal device includes the second audio mode;

[0179] Determining a second bit rate intersection between a second audio mode supported by the first terminal device and a second audio mode supported by the second terminal device;

[0180] Determine that the decoding negotiation result of the first terminal device includes the maximum code rate in the second code rate intersection.

[0181] Among them, the decoding negotiation result in the embodiment of the present application includes the maximum bit rate in the intersection of the second bit rates. According to the principle of maximizing experience, the encoding bit rate with a higher experience effect of the other terminal is preferentially selected to improve the call quality.

[0182] In some embodiments of the present application, the sum of the bit rate corresponding to the highest encoding mode supported by the first terminal device and the bit rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum bit rate supported by the audio encoding / decoding capability of the first terminal device.

[0183] In the embodiment of the present application, the sum of the bit rate corresponding to the highest encoding mode supported by the first terminal device and the bit rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum bit rate supported by the audio encoding / decoding capability of the first terminal device, thereby making full use of the maximum bit rate of the first terminal device, maximizing the use of the audio encoding / decoding capability of the first terminal device, and improving call quality.

[0184] In some embodiments of the present application, the first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata-assisted spatial audio (MASA), higher-order ambisonics (HOA), or first-order ambisonics (FOA);

[0185] The rate set corresponding to the codec mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.

[0186] In some embodiments of the present application, the method further comprises:

[0187] The first terminal device sends an audio negotiation request to the second terminal device.

[0188] The first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains the mode control parameter of the second terminal device according to the triggering of the audio negotiation request.

[0189] In some embodiments of the present application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device.

[0190] Among them, the capability negotiation request sent by the first terminal device to the second terminal device carries the audio codec configuration information of the first terminal device. The second terminal device can make a preliminary selection based on the audio codec configuration information received from the first terminal device. The audio codec configuration information of the second terminal device obtained can be one or more audio codec configuration information after preliminary screening, and then the first terminal device performs audio negotiation, thereby further improving the efficiency of audio negotiation.

[0191] It further includes: a capability negotiation request sent by the first terminal device to the second terminal device, and may not include the audio encoding / decoding configuration information of the first terminal device, which is not limited here.

[0192] The above embodiment illustrates the method executed by the first terminal device, and the method executed by the second terminal device is described next. The configuration method of the mode control parameters of the second terminal device by the second terminal device is similar to the configuration method of the mode control parameters of the aforementioned first terminal device. The specific process of determining the audio encoding / decoding configuration information of the second terminal device based on the mode control parameters of the second terminal device will not be described in detail later.

[0193] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control of the second terminal device includes:

[0194] The audio encoding / decoding configuration information of the second terminal device is determined according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.

[0195] In some embodiments of the present application, the mode control parameters of the second terminal device include at least one of the following: the battery life mode of the second terminal device, the audio call user experience of the second terminal device, and the user's power allocation mode for the second terminal device.

[0196] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:

[0197] Acquiring battery life information of the second terminal device according to the battery parameters of the second terminal device;

[0198] The audio encoding / decoding configuration information of the second terminal device is updated according to the battery life information and the audio encoding / decoding capability of the second terminal device to obtain updated audio encoding / decoding configuration information of the second terminal device.

[0199] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0200] When the mode control parameter of the second terminal device meets the preset first software configuration condition, determining the first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;

[0201] When the mode control parameter of the second terminal device does not meet the preset first software configuration condition, determining the second audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;

[0202] The audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.

[0203] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0204] The mode control parameters of the second terminal device include the first battery power collected by the second terminal device with the permission of the user. The audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power collected by the second terminal device with the permission of the user to obtain the updated audio encoding / decoding configuration information of the second terminal device.

[0205] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes:

[0206] Determine a first power range to which the power of the second battery belongs;

[0207] Determining a first audio codec mode corresponding to the first power interval according to a correspondence between the power interval and the audio codec mode of the second terminal device;

[0208] Adjust the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.

[0209] In some embodiments of the present application, updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes:

[0210] When the power level of the first battery is lower than a preset first power threshold and the second terminal device is in a battery charging mode, the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device is maintained or increased.

[0211] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:

[0212] The mode control parameters of the second terminal device include user habit parameters collected by the second terminal device with user permission, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameters collected by the second terminal device with user permission.

[0213] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device based on the user habit parameters collected by the second terminal device with the user's permission includes:

[0214] determining a first estimated standby time of the second terminal device according to the user schedule information of the second terminal device;

[0215] Determining a second audio codec mode corresponding to the first estimated standby time according to a correspondence between the standby time and the audio codec mode of the second terminal device;

[0216] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.

[0217] In some embodiments of the present application, determining the audio encoding / decoding configuration information of the second terminal device based on the user habit parameters collected by the second terminal device with the user's permission includes:

[0218] determining terminal capability allocation information according to user operation habit information of the second terminal device;

[0219] Determining a third audio codec mode according to a correspondence between the terminal capability allocation and the audio codec mode of the second terminal device;

[0220] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.

[0221] In some embodiments of the present application, the method further includes:

[0222] Obtaining a network packet loss rate of the audio call performed by the second terminal device;

[0223] When the network packet loss rate is higher than a first network signal threshold, the audio encoding / decoding configuration information of the second terminal device is updated.

[0224] In some embodiments of the present application, the method further comprises:

[0225] Receive an audio negotiation request from the first terminal device.

[0226] In some embodiments of the present application, the audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device;

[0227] The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device;

[0228] The sending the audio codec configuration information of the second terminal device to the first terminal device includes: sending candidate audio codec configuration information of the second terminal device to the first terminal device.

[0229] Next, we will take an actual application scenario as an example to illustrate.

[0230] The method provided in the embodiment of the present application is applicable to IVAS encoding and decoding. Different signal types and corresponding rates can be used, which requires further negotiation. The embodiment of the present application provides a negotiation method for IVAS encoding and decoding modes and rates.

[0231] Each IVAS-supported terminal has its own IVAS codec configuration table, which describes all supported IVAS codec combinations within the terminal's hardware constraints. By passing this codec configuration table during the negotiation process, both parties can more flexibly negotiate call configurations, such as the type of encoded signal and bit rate. Furthermore, it supports calls using different signal types and bit rates. This present invention provides the best possible user experience for both parties while meeting their respective hardware constraints.

[0232] Each mobile phone that supports IVAS will maintain two tables based on its own capabilities (computing power, memory, etc.). One table is the IVAS codec mode configuration table, as shown in Table 1 below:

[0233] The IVAS codec mode configuration table can represent all possible combinations of IVAS encoding and decoding modes that the mobile phone can support, including the encoding or decoding signal type and bit rate. As shown in the example in Table 1 above, the mobile phone supports three types of decoding signals: MONO, STEREO, and FOA. The maximum bit rate supported by MONO decoding is 9.6kbps, the maximum bit rate supported by STEREO decoding is 32kbps, and the maximum bit rate supported by FOA decoding is 48kbps. When the mobile phone decodes MONO signals, it can support three types of encoding signals: MONO, STEREO, and FOA. The maximum bit rates supported when encoding these three signals are 9.6kbps, 32kbps, and 48kbps, respectively. When the phone decodes STEREO signals, it can still support three types of coded signals: MONO, STEREO, and FOA. However, because the overhead of decoding STEREO is greater than that of decoding MONO, the remaining computing power left for encoding will decrease. At this time, the FOA encoding bitrate cannot support 48kbps, and can only support a maximum of 24kbps encoding (the higher the bitrate, the greater the overhead). When the phone decodes FOA signals, due to the greater decoding overhead of FOA, the remaining computing power can no longer support FOA encoding. At this time, the only coded signal types that the phone can support are MONO and STEREO, with maximum supported bitrates of 9.6kbps and 32kbps, respectively.

[0234] From the above examples, we can see that the combination of codec modes is limited by the capabilities of the terminal / chip. If the encoding overhead is large, the decoding overhead must be reduced, and vice versa.

[0235] The other table is the codec rate set table, as shown in Table 2 below:

[0236] The codec rate set table can represent all the encoding rates supported by IVAS for each signal. Combining the maximum rate supported by each encoding / decoding signal in the codec mode configuration table with the codec rate set table, the rate range supported by each encoding / decoding signal in the codec mode configuration table can be obtained.

[0237] Every mobile phone has a default IVAS codec configuration table when it leaves the factory. Depending on factors such as the chip computing power of the mobile phone, the IVAS codec configuration table of different mobile phone platforms may be different.

[0238] The types of IVAS signals that a mobile phone can encode are also related to some external factors. For example, FOA / HOA encoding can only be supported when the mobile phone has FOA or HOA acquisition equipment. Multi-channel or object encoding can only be supported when the voice service contains multi-channel or object audio / sound effect signals. For example, only mobile phones that support the complete MASA solution can support MASA signal encoding. When the mobile phone battery is low, the encoding and decoding of certain high-complexity signals may be restricted. This requires the mobile phone to dynamically update its own IVAS codec configuration table to reflect the current actual capabilities of the mobile phone.

[0239] As shown in FIG5 , the audio negotiation process of terminals UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) is taken as an example for explanation:

[0240] Before A and B can start a voice call, media negotiation is required if the IVAS voice codec is used. This negotiation includes:

[0241] 1. Audio codec negotiation (EVS / IVAS / AMR).

[0242] 2. After the audio codec is negotiated to IVAS, the codec mode (MONO / STEREO / FOA / MASA) and bit rate need to be further negotiated (each codec mode has a corresponding bit rate range). The encoding mode and decoding mode can be negotiated to be different.

[0243] Negotiation process:

[0244] 1. When UE-A initiates a call INVITE, if it supports IVAS voice codec, it sends the IVAS codec configuration table it supports to B in the IVAS SDP information, including a list of encoding and decoding mode combinations supported by A, the maximum rate supported by each mode, and the rate set supported by each encoding mode.

[0245] 2. UE-B performs the following operations based on its capabilities:

[0246] 1) Perform codec negotiation first. If IVAS is supported, the initial negotiation is IVAS.

[0247] 2) Continue to negotiate the codec mode and rate. Based on the local policy, decide whether to negotiate based on your encoding capability or decoding capability first. Take the negotiation based on encoding capability as an example:

[0248] ① Determine the coding mode of UE-B based on the highest coding mode supported by UE-B and whether UE-A supports the decoding mode;

[0249] ② Determine the encoding rate after UE-B negotiation based on the intersection of the code rates supported by the calling and called parties for the encoding mode. If there is no intersection, UE-B's encoding mode needs to be renegotiated;

[0250] ③ Based on the decoding mode corresponding to the coding mode list supported by UE-B and whether UE-A supports the coding mode, determine the decoding mode negotiated by UE-B;

[0251] ④ Determine the decoding bitrate after negotiation by UE-B based on the intersection of the bitrates of the decoding mode used by the calling and called parties.

[0252] If any of the IVAS-negotiated codec modes fails to be negotiated, IVAS will not be adopted.

[0253] As shown in Figures 6 and 7, the SDP sample adds the following parameters to the offer a line of IVAS:

[0254] Bitrate set:

[0255] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.

[0256] [mono,5.9 / 7.2 / 9.6] indicates that the bitrate set supported by mono is 5.9, 7.2, and 9.6.

[0257] Decoding mode set:

[0258] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.

[0259] This list indicates that three sets are supported. Each set corresponds to the following encoding mode set. The decoding mode corresponding to set 1 is [mono / 9.6], which means that mono is supported and the maximum bit rate is 9.6.

[0260] Encoding mode set:

[0261] enc-mode-list={[mono / 9.6,stereo / 32,foa / 48],[mono / 9.6,stereo / 32,foa / 24],[mono / 9.6,stereo / 32]}.

[0262] This list indicates that three sets are supported, which correspond one-to-one to the above decoding mode sets. The encoding mode corresponding to set 1 is = {[mono / 9.6,stereo / 32,foa / 48], which means that the decoding modes from low to high are mono / stereo / foa, and the highest bit rates are 9.6 / 32 / 48 respectively.

[0263] Add the following parameters to the answer a line of IVAS:

[0264] Negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]};

[0265] Negotiated coding mode and rate set: enc-mode = {[foa,24 / 32]}.

[0266] Example of negotiation process:

[0267] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A and the maximum rate combination supported by each mode, as well as the rate set supported by each coding mode.

[0268] 2. Mobile phone B decides based on local policy whether to negotiate based on its encoding capabilities or its decoding capabilities. Taking encoding capability negotiation as an example:

[0269] ① Based on the fact that the highest coding mode supported by UE-B is FOA and that the decoding mode of UE-A includes FOA, the coding mode negotiated by UE-B is determined to be FOA;

[0270] ② Determine the intersection of the bit rates supported by the calling and called FOA modes, and determine that the encoding bit rates negotiated by UE-B are 24 and 32;

[0271] ③ According to the decoding mode corresponding to the coding mode column supported by UE-B, which is MONO, and the coding mode of UE-A includes MONO, it is determined that the decoding mode negotiated by UE-B is MONO;

[0272] ④ Determine the intersection of the bit rates of the MONO mode supported by the calling and called parties, and determine that the decoding bit rates after UE-B negotiation are 5.9, 7.2, and 9.6.

[0273] Figure 8 shows another example of the negotiation process:

[0274] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the codec modes supported by A and the maximum rate combination supported by each mode, as well as the rate set supported by each coding mode.

[0275] 2. Mobile phone B decides based on local policy whether to negotiate based on its encoding capabilities or its decoding capabilities. Taking encoding capability negotiation as an example:

[0276] ① Based on the fact that the highest decoding mode supported by UE-B is FOA and that the decoding mode of UE-A includes FOA, it is determined that the decoding mode negotiated by UE-B is FOA;

[0277] ② Determine the intersection of the code rates of the FOA modes supported by the calling and called parties, and determine that the encoding code rates after UE-B negotiation are 24, 32, and 48;

[0278] ③ According to the coding mode column supported by UE-B, the highest coding mode is STEREO, and UE-A's decoding mode includes STEREO. Therefore, the coding mode negotiated by UE-B is determined to be STEREO.

[0279] ④ Determine the intersection of the STEREO bit rates supported by the calling and called parties, and ensure that the decoding bit rates after negotiation by UE-B are 13.2 and 24.

[0280] As can be seen from the preceding examples, for real-time audio codecs with numerous encoding modes and rates, and a wide range of codec complexity, each terminal maintains a specific codec configuration table. This table describes the various encoding and decoding mode / rate combinations supported by the terminal. When two or more terminals establish a call and negotiate the codec, they select the appropriate encoding configuration for each terminal based on their respective codec configuration tables.

[0281] Each terminal has a factory-default codec configuration table. During codec negotiation during a call, the terminal can update its own codec configuration table based on other factors and perform codec negotiation based on the updated codec configuration table. The updated codec configuration table does not overwrite the factory-default codec configuration table.

[0282] In the codec configuration table, the supported decoding modes / bitrates and / or encoding modes / bitrates are sorted by experience priority.

[0283] During codec negotiation, one terminal sends its complete or partial codec configuration table to the other terminal for negotiation. The terminal makes a negotiation decision based on the complete or partial codec configuration table of at least two parties. This refers to partial coding modes and rates, which is a subset.

[0284] The negotiation decision may be made based on the principle of maximizing experience, or based on other principles, such as a decision principle preset in the terminal making the negotiation decision, such as giving priority to the other terminal for encoding with a higher experience effect.

[0285] The following example illustrates the audio negotiation process for terminal power.

[0286] When the terminal power is 100%, the terminal configuration table is the initial configuration table, such as shown in Table 3 below:

[0287] When the terminal battery power is low, in order to ensure the operation of other terminal functions, it is necessary to adjust the configuration table and reduce the configuration. For example, when the battery power is only 10%, the configuration table is adjusted as shown in Table 4 below:

[0288] When the terminal uses only configuration 1 (default hardware configuration / factory configuration), the terminal's configuration table is the initial configuration table, as shown in Table 3 above.

[0289] When the terminal uses configuration 2 (external FOA microphone), the configuration table modifies some combinations of Table 3.

[0290] When the terminal uses configuration 3 (external HOA microphone), the configuration table replaces combination 3 in Table 3 with HOA and the corresponding bit rate.

[0291] When the terminal uses configuration 4 (external microphone), the terminal-side device can convert the microphone signal into FOA / HOA / MASA signal supported by the Codec. The configuration table includes combination 3 for FOA and the corresponding bit rate, combination 4 for HOA and the corresponding bit rate, and combination 5 for MASA and the corresponding bit rate.

[0292] The following examples illustrate the application scenarios and corresponding solutions involved in the embodiments of this application. The computing power and memory of the chip (Complexity and RAM / ROM Memory in the IVAS document) are fixed in terms of the application modes that can be supported under a given hardware configuration.

[0293] Assuming the IVAS standard supports a total of 150 operating modes, and that low-power modes generally correspond to lower call quality or a lower sense of immersion, the lowest complexity operating mode is 1, the highest complexity mode is 150, and the others are arranged in ascending order of complexity.

[0294] IVAS's maximum computing power and memory requirements are approximately ten times that of EVS. However, for a mobile phone product, it may only have six times the hardware capabilities of EVS, so it can only support the first 80 working modes, for example. This is a limitation of the maximum complexity working mode caused by the hardware configuration.

[0295] When the battery is fully charged, the phone can support the 1-80 working mode. However, when the battery drops to 50%, it may selectively support the 1-50 working mode and give up support for the high-complexity working mode. When the battery drops to 20%, the supported working mode can be further reduced to 1-30, thereby maximizing the phone's usage time.

[0296] Alternatively, the user terminal can determine the optimal power consumption mode based on the user's schedule. For example, if the user's schedule today is not busy, with only one short work schedule, the user's historical usage data indicates that the battery level is sufficient. At this time, the phone's application configuration can maintain a high-quality call mode (1-80). However, tomorrow, the user's schedule is very tight, such as requiring a conference call. Therefore, the phone's configuration is more conservative from the beginning of the day (for example, only using 1-40). This ensures that the user can maintain the phone's battery level throughout the day without charging, avoiding battery warnings on the phone.

[0297] Another situation is that the user is currently on a call and the battery is seriously low. The mobile phone's automatic configuration mechanism determines that it is necessary to use low-power mode (for example, 1-30), but the user himself determines that there will be an opportunity to charge within a short time, so the user can set the phone to continue to remain in high-quality call mode (1-80).

[0298] In another implementation scenario, due to a busy mobile network, the user's received signal may be extremely unstable, resulting in frequent call interruptions. In this case, the best solution is to switch to a low-rate mode. This is a decision that the network negotiation mechanism should participate in, but the terminal cannot perceive the characteristics of the user's received signal. In this case, it is desirable for the user's phone to intelligently participate in audio negotiation decisions.

[0299] In the embodiment of the present application, IVAS can achieve intelligent adjustment due to its multi-functional and multi-mode capabilities, thereby further improving the user experience.

[0300] For another example, if a user allows the collection of historical information about their terminal operations, artificial intelligence (AI) or statistics can be used to determine, for example, that users of the first type use their terminals to play games and consume more power, users of the second type use their terminals to make video calls, and users of the third type use more video calls. After collecting user usage habits, the phone can determine the user's characteristics and determine if their usage habits are likely to remain unchanged. This allows the phone to determine the optimal operating mode for the user.

[0301] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0302] In order to better implement the above-mentioned solutions of the embodiments of the present application, relevant devices for implementing the above-mentioned solutions are also provided below.

[0303] As shown in FIG9 , a first terminal device 900 provided in an embodiment of the present application may include: an acquisition module 901 , a determination module 902 , a negotiation module 903 and a sending module 904 , wherein:

[0304] an acquisition module, configured to acquire audio encoding / decoding configuration information of a second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device;

[0305] a determining module, configured to determine audio encoding / decoding configuration information of the first terminal device according to a mode control parameter of the first terminal device, where the mode control parameter of the first terminal device at least includes: a user description parameter generated by the first terminal device;

[0306] a negotiation module, configured to determine an encoding / decoding negotiation result of the first terminal device based on the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result comprising: first encoding / decoding parameters for encoding and decoding the audio signal to be processed by the first terminal device, the first encoding / decoding parameters being adapted to the audio encoding / decoding capabilities of the second terminal device;

[0307] A sending module is used to send the encoding / decoding negotiation result of the first terminal device to the second terminal device.

[0308] Through the examples of the above embodiments, it can be seen that first, the first terminal device obtains the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameters of the first terminal device, and the mode control parameters of the first terminal device at least include: the user description parameters generated by the first terminal device; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameter for encoding and decoding the processed audio signal by the first terminal device, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiment of the present application, both terminals performing audio negotiation each have audio encoding / decoding configuration information, and the audio encoding / decoding configuration information describes the encoding and decoding capabilities that the terminal can support under the constraints of the mode control parameters. By transmitting audio encoding / decoding configuration information during the negotiation process, the two parties of the call can negotiate the call configuration more flexibly. In the embodiment of the present application, the best user experience can be provided to both parties as much as possible while satisfying the respective mode control parameter constraints of the two parties of the call.

[0309] Referring to FIG. 10 , a second terminal device 1000 provided in an embodiment of the present application may include: an acquisition module 1001 , a determination module 1002 , a sending module 1003 , and a receiving module 1004 , wherein:

[0310] an acquisition module, configured to acquire a mode control parameter of the second terminal device, where the mode control parameter of the second terminal device at least includes: a user description parameter generated by the first terminal device;

[0311] a determining module, configured to determine audio encoding / decoding configuration information of the second terminal device according to a mode control parameter of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate an audio encoding / decoding capability of the second terminal device;

[0312] A sending module, configured to send audio encoding / decoding configuration information of the second terminal device to the first terminal device;

[0313] A receiving module is used to receive the encoding / decoding negotiation result from the first terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameters for the first terminal device to encode and decode the audio signal to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.

[0314] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.

[0315] An embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a program, and the program executes some or all of the steps recorded in the above method embodiment.

[0316] Next, another first terminal device provided in an embodiment of the present application is introduced. Referring to FIG11 , the first terminal device 1100 includes:

[0317] Receiver 1101, transmitter 1102, processor 1103, and memory 1104 (wherein the number of processors 1103 in the first terminal device 1100 may be one or more, and FIG11 takes one processor as an example). In some embodiments of the present application, the receiver 1101, transmitter 1102, processor 1103, and memory 1104 may be connected via a bus or other means, wherein FIG11 takes connection via a bus as an example.

[0318] Memory 1104 may include read-only memory and random access memory, and provides instructions and data to processor 1103. A portion of memory 1104 may also include non-volatile random access memory (NVRAM). Memory 1104 stores an operating system and operating instructions, executable modules, or data structures, or subsets or extensions thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0319] Processor 1103 controls the operation of the first terminal device and may also be referred to as a central processing unit (CPU). In specific applications, the various components of the first terminal device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all of these buses are collectively referred to as a bus system in the figure.

[0320] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1103. Processor 1103 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1103. The above processor 1103 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in storage media well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1104 , and the processor 1103 reads the information in the memory 1104 and completes the steps of the above method in combination with its hardware.

[0321] The receiver 1101 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the first terminal device. The transmitter 1102 may include a display device such as a display screen. The transmitter 1102 can be used to output digital or character information through an external interface.

[0322] In the embodiment of the present application, the processor 1103 is used to execute the method performed by the first terminal device shown in Figure 4 of the aforementioned embodiment.

[0323] Next, another second terminal device provided in an embodiment of the present application is introduced. Referring to FIG. 12 , the second terminal device 1200 includes:

[0324] Receiver 1201, transmitter 1202, processor 1203, and memory 1204 (wherein the number of processors 1203 in the second terminal device 1200 may be one or more, and FIG12 takes one processor as an example). In some embodiments of the present application, the receiver 1201, transmitter 1202, processor 1203, and memory 1204 may be connected via a bus or other means, wherein FIG12 takes connection via a bus as an example.

[0325] Memory 1204 may include read-only memory and random access memory, and provides instructions and data to processor 1203. A portion of memory 1204 may also include NVRAM. Memory 1204 stores an operating system and operating instructions, executable modules, or data structures, or subsets or extensions thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0326] Processor 1203 controls the operation of the second terminal device and may also be referred to as a CPU. In specific applications, the various components of the second terminal device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0327] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1203. Processor 1203 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 1203 or by software instructions. Processor 1203 can be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory 1204. Processor 1203 reads information from memory 1204 and, in conjunction with its hardware, completes the steps of the above method.

[0328] In an embodiment of the present application, the processor 1203 is used to execute the method performed by the second terminal device as shown in Figure 7 of the aforementioned embodiment.

[0329] In another possible design, when the first terminal device or the second terminal device is a chip within a terminal, the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip within the terminal executes any one of the methods of the first aspect described above. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit within the terminal located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0330] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect or second aspect method.

[0331] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0332] Through the description of the above embodiments, it is clear to those skilled in the art that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course it can also be implemented by means of dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0333] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0334] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, a computer, a server or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

Claims

1. An audio codec negotiation method, characterized in that: The method is applied to a second terminal device, and the method includes: Acquire a mode control parameter of the second terminal device, where the mode control parameter of the second terminal device at least includes: a user description parameter generated by the second terminal device; Determining audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device; Sending audio encoding / decoding configuration information of the second terminal device to the first terminal device; Receive a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result including: first coding / decoding parameters for the first terminal device to encode and decode the audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.

2. The method according to claim 1, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the mode control of the second terminal device includes: The audio encoding / decoding configuration information of the second terminal device is determined according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.

3. The method according to claim 1 or 2, characterized in that The mode control parameter of the second terminal device includes at least one of the following: a battery life mode of the second terminal device, an audio call user experience of the second terminal device, and a power allocation mode of the user to the second terminal device.

4. The method according to claim 2, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes: Acquiring battery life information of the second terminal device according to the battery parameters of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is updated according to the battery life information and the audio encoding / decoding capability of the second terminal device to obtain updated audio encoding / decoding configuration information of the second terminal device.

5. The method according to claim 2 or 3, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes: When the mode control parameter of the second terminal device meets the preset first software configuration condition, determining the first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device; When the mode control parameter of the second terminal device does not meet the preset first software configuration condition, determining the second audio encoding / decoding configuration information according to the mode control parameter of the second terminal device; The audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.

6. The method according to claim 2 or 3, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes: The mode control parameters of the second terminal device include the first battery power collected by the second terminal device with the permission of the user. The audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power collected by the second terminal device with the permission of the user to obtain the updated audio encoding / decoding configuration information of the second terminal device.

7. The method according to claim 6, characterized in that The updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes: Determine a first power range to which the power of the second battery belongs; Determining a first audio codec mode corresponding to the first power interval according to a correspondence between the power interval and the audio codec mode of the second terminal device; Adjust the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.

8. The method according to any one of claims 1 to 7, characterized in that The updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with the permission of the user includes: When the power level of the first battery is lower than a preset first power threshold and the second terminal device is in a battery charging mode, the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device is maintained or increased.

9. The method according to any one of claims 1 to 8, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes: The mode control parameters of the second terminal device include user habit parameters collected by the second terminal device with user permission, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameters collected by the second terminal device with user permission.

10. The method according to claim 9, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the user habit parameters collected by the second terminal device with the permission of the user includes: determining a first estimated standby time of the second terminal device according to the user schedule information of the second terminal device; Determining a second audio codec mode corresponding to the first estimated standby time according to a correspondence between the standby time and the audio codec mode of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.

11. The method according to claim 9, characterized in that The determining the audio encoding / decoding configuration information of the second terminal device according to the user habit parameters collected by the second terminal device with the permission of the user includes: determining terminal capability allocation information according to user operation habit information of the second terminal device; Determining a third audio codec mode according to a correspondence between the terminal capability allocation and the audio codec mode of the second terminal device; The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.

12. The method according to any one of claims 1 to 11, characterized in that The method further comprises: Obtaining a network packet loss rate of the audio call performed by the second terminal device; When the network packet loss rate is higher than a first network signal threshold, the audio encoding / decoding configuration information of the second terminal device is updated.

13. The method according to any one of claims 1 to 12, characterized in that The method further comprises: Receive an audio negotiation request from the first terminal device.

14. The method according to claim 13, characterized in that The audio negotiation request includes: audio encoding / decoding configuration information of the first terminal device; The method further includes: screening the audio codec configuration information of the second terminal device according to the audio codec configuration information of the first terminal device to obtain candidate audio codec configuration information of the second terminal device; The sending the audio codec configuration information of the second terminal device to the first terminal device includes: sending candidate audio codec configuration information of the second terminal device to the first terminal device.

15. A terminal device, characterized in that: The terminal device is specifically a second terminal device, and the method includes: An acquisition module is used for mode control parameters of the second terminal device, where the mode control parameters of the second terminal device at least include: user description parameters generated by the first terminal device; a determining module, configured to determine audio encoding / decoding configuration information of the second terminal device according to a mode control parameter of the second terminal device, where the audio encoding / decoding configuration information of the second terminal device is used to indicate an audio encoding / decoding capability of the second terminal device; A sending module, configured to send audio encoding / decoding configuration information of the second terminal device to the first terminal device; A receiving module is used to receive the encoding / decoding negotiation result from the first terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameters for the first terminal device to encode and decode the audio signal to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capabilities of the second terminal device.

16. A terminal device, characterized in that: The terminal device includes at least one processor, and the at least one processor is configured to be coupled to a memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 14.

17. The terminal device according to claim 16, characterized in that The terminal device further includes: the memory.

18. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Processing method of audio data and voice communication terminal

    CN104952454A

  • Method and device for determining set of codec modes for traffic communications

    CN112398854A

  • Voice communication method, electronic equipment and readable medium

    CN113573233A

  • Voice processing method, terminal device and storage medium

    CN114694662A

  • Voice coding method and device, electronic equipment and storage medium

    CN114898760A