Negotiation method for audio encoding / decoding and terminal device
By acquiring the mode control parameters of the terminal device, determining its audio codec configuration information, and negotiating, the problem of matching the codec capabilities of immersive voice codecs between terminals is solved, achieving optimal user experience and resource utilization.
Patent Information
- Application Number
- PCT/CN2025/075112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-01-26
- Publication Date
- 2025-11-27
AI Technical Summary
Existing audio codec negotiation methods cannot accurately match the codec capabilities between terminals when supporting highly complex immersive voice codecs. This results in stable call connections but poor codec performance and wasted resources.
By obtaining the mode control parameters of the terminal device, its audio codec configuration information is determined and transmitted during the negotiation process, so that both terminals can flexibly negotiate according to their respective capabilities and select the optimal codec configuration.
This achieves the goal of providing the best user experience for both parties while satisfying the control parameter constraints of their respective modes, thus avoiding problems such as resource waste and poor encoding and decoding quality.
Smart Images

Figure CN2025075112_27112025_PF_FP_ABST
Abstract
Description
Method for negotiating audio encoding / decoding and terminal device
[0001] The present application claims priority from the Chinese patent application No. 202410357092.2 filed on March 26, 2024, and entitled "Method for negotiating audio encoding / decoding and terminal device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of audio encoding and decoding, and in particular to a method for negotiating audio encoding / decoding and a terminal device. BACKGROUND
[0003] Compared with a speech codec using a single channel, an immersive speech codec supports more types of audio signals, supports a larger range of code rates, and includes rendering characteristics in addition to codec characteristics. For example, in the audio codec standard being developed by the 3rd Generation Partnership Project (3GPP), different audio codec standards support multiple types of signals, different types of signals support different numbers of channels and code rates, and the rendering characteristics of different types of signals support rendering of various types of signals to binaural and standard loudspeaker arrays. Due to the support of more signal channel numbers, a larger range of code rates, and additional rendering characteristics, the immersive speech codec has a very significant increase in maximum computational complexity and storage complexity compared to the traditional speech codec.
[0004] In order to support a speech codec with high complexity, in an existing method for negotiating audio encoding / decoding, the speech codec is classified according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result can only select a level with lower complexity for the call to ensure that the call connection is stable. Although this method meets the speech requirements of the call connection, it has the problem of poor audio encoding / decoding negotiation effect. SUMMARY
[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a method for negotiating audio encoding / decoding, characterized in that the method is applied to a second terminal device, and the method comprises:
[0007] obtaining a mode control parameter of the second terminal device, wherein the mode control parameter of the second terminal device at least includes a setting parameter of the second terminal device set by a user;
[0008] determining audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device;
[0009] sending the audio encoding / decoding configuration information of the second terminal device to the first terminal device;
[0010] receiving a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result including first coding / decoding parameters of the first terminal device for encoding / decoding audio signals, the first coding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0011] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:
[0012] determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0013] In some embodiments of the present application, the mode control parameter of the second terminal device includes at least one of a battery endurance mode of the second terminal device, an audio call user experience of the second terminal device, and a power distribution mode of the second terminal device.
[0014] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device includes:
[0015] obtaining battery endurance information of the second terminal device according to the battery parameter of the second terminal device;
[0016] updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0017] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device includes:
[0018] when the mode control parameter of the second terminal device meets a preset first software configuration condition, determining first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;
[0019] determining second audio coding configuration information according to the mode control parameter of the second terminal device when the mode control parameter of the second terminal device does not satisfy a preset first software configuration condition;
[0020] The audio coding capability corresponding to the first audio coding configuration information is higher than the audio coding capability corresponding to the second audio coding configuration information.
[0021] In some embodiments of the present application, the determining of the audio coding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0022] The mode control parameter of the second terminal device comprises a first battery power set by a user for the second terminal device, and the audio coding configuration information of the second terminal device is updated according to the first battery power set by the user for the second terminal device to obtain updated audio coding configuration information of the second terminal device.
[0023] In some embodiments of the present application, the updating of the audio coding configuration information of the second terminal device according to the first battery power set by the user for the second terminal device comprises:
[0024] determining a first power interval to which the second battery power belongs;
[0025] determining a first audio coding mode corresponding to the first power interval according to a correspondence between power intervals and audio coding modes of the second terminal device;
[0026] adjusting the audio coding configuration information of the second terminal device according to the audio coding capability corresponding to the first audio coding mode.
[0027] In some embodiments of the present application, the updating of the audio coding configuration information of the second terminal device according to the first battery power set by the user for the second terminal device comprises:
[0028] when the first battery power is lower than a preset first power threshold and the second terminal device is in a battery charging mode, maintaining or increasing the audio coding capability indicated by the audio coding configuration information of the second terminal device.
[0029] In some embodiments of the present application, the determining of the audio coding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0030] The mode control parameter of the second terminal device comprises a user habit parameter set by a user to the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter set by the user to the second terminal device.
[0031] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter set by the user to the second terminal device, comprising:
[0032] The first estimated standby time length of the second terminal device is determined according to user schedule information of the second terminal device;
[0033] The second audio encoding / decoding mode corresponding to the first estimated standby time length is determined according to a corresponding relationship between the standby time length and the audio encoding / decoding mode of the second terminal device;
[0034] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.
[0035] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter set by the user to the second terminal device, comprising:
[0036] The terminal capability allocation information is determined according to the user operation habit information of the second terminal device;
[0037] The third audio encoding / decoding mode is determined according to a corresponding relationship between the terminal capability allocation and the audio encoding / decoding mode of the second terminal device;
[0038] The audio encoding / decoding configuration information of the second terminal device is determined according to the audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.
[0039] In some embodiments of the present application, the method further comprises:
[0040] The network packet loss rate of the second terminal device for audio communication is acquired;
[0041] When the network packet loss rate is higher than a first network signal threshold, the audio encoding / decoding configuration information of the second terminal device is updated.
[0042] In some embodiments of the present application, the method further comprises:
[0043] The audio negotiation request from the first terminal device is received.
[0044] In some embodiments of the present application, the audio negotiation request comprises the audio encoding / decoding configuration information of the first terminal device.
[0045] The method further includes: screening the audio coding / decoding configuration information of the second terminal device according to the audio coding / decoding configuration information of the first terminal device to obtain candidate audio coding / decoding configuration information of the second terminal device.
[0046] The sending of the audio coding / decoding configuration information of the second terminal device to the first terminal device includes: sending the candidate audio coding / decoding configuration information of the second terminal device to the first terminal device.
[0047] In a second aspect of the present application, the component module of the second terminal device can also perform the steps described in the foregoing first aspect and various possible implementation manners, for details, refer to the foregoing description of the first aspect and various possible implementation manners.
[0048] In a third aspect, the embodiments of the present application provide a computer readable storage medium, which stores instructions, when the instructions are run on a computer, cause the computer to execute the method in the foregoing first aspect.
[0049] In a fourth aspect, the embodiments of the present application provide a computer program product containing instructions, when the instructions are run on a computer, cause the computer to execute the method in the foregoing first aspect.
[0050] In a fifth aspect, the embodiments of the present application provide a communication apparatus, which can include a terminal device or a chip and the like entity, the communication apparatus includes: a processor, a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the communication apparatus executes the method in any one of the foregoing first aspect.
[0051] In a sixth aspect, the embodiments of the present application provide a chip system, which includes a processor, used to support the terminal device to realize the functions involved in the foregoing aspects, for example, sending or processing the data and / or information involved in the foregoing method. In a possible design, the chip system further includes a memory, the memory is used to save the necessary program instructions and data of the terminal device. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0052] In a seventh aspect, the embodiments of the present application provide a chip, which includes one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of an electronic device, and send a signal to the processor, the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, causes the electronic device to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0053] From the above technical solution, the embodiments of the present application have the following advantages:
[0054] First, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, which is used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, which at least includes the setting parameter of the first terminal device set by the user; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, which includes the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiments of the present application, both the two terminals for audio negotiation have audio encoding / decoding configuration information, which describes the encoding / decoding capability that the terminal can support under the constraint of the mode control parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiments of the present application, the optimal user experience can be provided for the two parties as much as possible under the constraint of the mode control parameter of the two parties of the call. BRIEF DESCRIPTION OF DRAWINGS
[0055] FIG. 1 is a schematic diagram of the composition structure of the audio processing system provided by the embodiments of the present application;
[0056] FIG. 2a is a schematic diagram of the application of the audio encoder and the audio decoder to the terminal device provided by the embodiments of the present application;
[0057] FIG. 2b is a schematic diagram of the application of the audio encoder to the wireless device or the core network device provided by the embodiments of the present application;
[0058] FIG. 2c is a schematic diagram of the application of the audio decoder to the wireless device or the core network device provided by the embodiments of the present application;
[0059] FIG. 3a is a schematic diagram of the application of the multi-channel encoder and the multi-channel decoder to the terminal device provided by the embodiments of the present application;
[0060] FIG. 3b is a schematic diagram of the application of the multi-channel encoder to the wireless device or the core network device provided by the embodiments of the present application;
[0061] FIG. 3c is a schematic diagram of the application of the multi-channel decoder to the wireless device or the core network device provided by the embodiments of the present application;
[0062] Fig. 4 is a schematic diagram of a negotiation method of audio encoding / decoding provided by an embodiment of the present application;
[0063] Fig. 5 is a schematic diagram of a negotiation process of audio negotiation by two terminals through SDP provided by an embodiment of the present application;
[0064] Fig. 6 is a schematic diagram of a negotiation process of audio encoding / decoding mode and code rate by two terminals provided by an embodiment of the present application;
[0065] Fig. 7 is a schematic diagram of a negotiation result of encoding / decoding provided by two terminals after negotiation provided by an embodiment of the present application;
[0066] Fig. 8 is a schematic diagram of another negotiation process of audio encoding / decoding mode and code rate by two terminals provided by an embodiment of the present application;
[0067] Fig. 9 is a schematic diagram of a structure of a first terminal device provided by an embodiment of the present application;
[0068] Fig. 10 is a schematic diagram of a structure of a second terminal device provided by an embodiment of the present application;
[0069] Fig. 11 is a schematic diagram of another structure of a first terminal device provided by an embodiment of the present application;
[0070] Fig. 12 is a schematic diagram of another structure of a second terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0071] The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0072] The terms "first", "second", and the like in the description and the claims of the present application and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the terms thus used can be interchanged under appropriate circumstances, and are merely used to distinguish between similar objects in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, a method, a system, a product, or an apparatus including a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to the process, the method, the product, or the apparatus.
[0073] The code rate of traditional single-channel speech codec is between several kbps and tens or 128 kbps. Compared with the single-channel speech codec, the immersive speech codec supports more signal types, supports a larger range of code rates, and additionally includes rendering characteristics in addition to the codec characteristics. Taking the 3GPP ongoing immersive voice and audio service (IVAS) speech / audio codec standard as an example, the signal types supported by the IVAS speech / audio codec include single-channel, stereo, multi-channel, multi-object, higher order ambisonics (HOA) signal or first order ambisonics (FOA) signal, MASA, etc. Among them, the multi-channel signal supports up to 7.1.4 channel format, i.e., 12 channels, and the HOA signal supports up to 3-order HOA, i.e., 16 channels, and the supported code rate can be as high as 768 kbps from the lowest single-channel 5.9 kbps. The rendering supports rendering of various signal types to binaural and standard loudspeaker arrays. More signal channels, a larger range of code rates, and additional rendering characteristics significantly increase the maximum computational complexity and storage complexity of the immersive speech codec compared with the traditional speech codec, which will result in a greater difficulty in landing the chip of this type of codec.
[0074] In order to solve the problem of high complexity of the speech codec, the speech codec is divided into levels according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result is to select the lower level for the call to ensure stable call connection. This method meets the speech demand of the call connection, but has the problem of poor audio codec negotiation effect. In addition, this method meets the speech demand of the call connection, but the terminal of the lower level cannot decode all types of code streams, which causes the coding and decoding limitation of the audio signal of the terminal of the higher level, and has the problem of poor quality of the audio signal coding and decoding. In addition, when the levels of the two terminals are not equal, the terminal of the higher level also has the problem of waste of coding and decoding capability.
[0075] Based on the above description, the embodiment of the present application provides an audio coding / decoding negotiation method, which can perform accurate codec negotiation when establishing a call, and select the most appropriate codec combination according to the actual capability of the communication platform of each party, so as to ensure the optimal call experience and not waste power consumption.
[0076] The embodiments of the present application provide an audio coding technology, in particular, provide a three-dimensional audio coding technology for three-dimensional audio signals, and specifically provide an encoding technology for representing three-dimensional audio signals by using fewer channels to improve a conventional audio coding system. Audio coding (or commonly referred to as encoding) includes two parts of audio encoding and audio decoding. Audio encoding is performed at a source side, including processing (for example, compression) of original audio to reduce the amount of data required to represent the audio, thereby more efficiently storing and / or transmitting. Audio decoding is performed at a destination side, including inverse processing relative to the encoder to reconstruct the original audio. The encoding part and the decoding part are also collectively referred to as encoding. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0077] The technical solutions of the embodiments of the present application can be applied to various audio processing systems. As shown in FIG. 1, an audio processing system provided by the embodiments of the present application includes an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 can be used to generate a bitstream, and then the audio encoding bitstream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the bitstream and then perform the audio decoding function of the audio decoding device 102, and finally obtain the reconstructed signal.
[0078] In the embodiments of the present application, the audio encoding device can be applied to various terminal devices that need audio communication, wireless devices that need transcoding, and core network devices. For example, the audio encoding device can be an audio encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding device can be applied to various terminal devices that need audio communication, wireless devices that need transcoding, and core network devices. For example, the audio decoding device can be an audio decoder of the terminal device, the wireless device, or the core network device. For example, the audio encoder can include a radio access network, a media gateway of a core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc. The audio encoder can also be an audio encoder applied to a virtual reality (VR) streaming service.
[0079] In the application embodiment, taking the audio encoding module and audio decoding module suitable for virtual reality streaming service as an example, the processing flow of the audio signal in an end-to-end manner includes: the audio signal A is preprocessed (audio PReprocessing) after being acquired by the acquisition module, the preprocessing operation includes filtering out the low-frequency part in the signal, which can be divided by 20 Hz or 50 Hz, extracting the azimuth information in the signal, and then performing encoding processing (audio encoding) and packaging (file / segment encapsulation) and then delivering (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), and then decodes (audio decoding), performs binaural rendering (audio rendering) processing on the decoded signal, and maps the processed signal to the listener's headphones (headphones), which can be independent headphones or headphones on glasses devices.
[0080] As shown in FIG. 2a, the audio encoder and the audio decoder provided by the application embodiment are applied to a schematic diagram of a terminal device. Each terminal device can include an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is configured to perform channel encoding on the audio signal, and the channel decoder is configured to perform channel decoding on the audio signal. For example, the first terminal device 20 can include a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 can include a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a first network communication device 22 via a wireless or wired connection, the first network communication device 22 and a second network communication device 23 are connected via a digital channel, and the second terminal device 21 is connected to the second network communication device 23 via a wireless or wired connection. The wireless or wired network communication device described above can be a signal transmission device, such as a communication base station or a data exchange device.
[0081] In audio communication, the terminal device as a sending end first acquires audio, encodes the acquired audio signal, and then performs channel encoding and transmits the signal in a digital channel through a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a code stream, and then restores the audio signal through audio decoding, and the terminal device as the receiving end plays back the audio signal.
[0082] As shown in FIG. 2b, a schematic diagram of the audio encoder provided by the embodiments of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 comprises a channel decoder 251, another audio decoder 252, the audio encoder provided by the embodiments of the present application 253, and a channel encoder 254, wherein the another audio decoder 252 refers to an audio decoder other than the audio decoder. In the wireless device or the core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, then audio-decoded by the another audio decoder 252, then audio-encoded by the audio encoder provided by the embodiments of the present application 253, and finally channel-encoded by the channel encoder 254, and the channel-encoded audio signal is then transmitted out. The another audio decoder 252 is used to audio-decode the code stream decoded by the channel decoder 251.
[0083] As shown in FIG. 2c, a schematic diagram of the audio decoder provided by the embodiments of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 comprises a channel decoder 251, the audio decoder provided by the embodiments of the present application 255, another audio encoder 256, and a channel encoder 254, wherein the another audio encoder 256 refers to an audio encoder other than the audio encoder. In the wireless device or the core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, then decoded by the audio decoder 255, then audio-encoded by the another audio encoder 256, and finally channel-encoded by the channel encoder 254, and the channel-encoded audio signal is then transmitted out. In the wireless device or the core network device, if transcoding is needed, corresponding audio encoding is needed. The wireless device refers to a radio frequency related device in communication, and the core network device refers to a core network related device in communication.
[0084] In some embodiments of the present application, the audio encoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices having transcoding needs, for example, the audio encoding apparatus can be a multi-channel encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices having transcoding needs, for example, the audio decoding apparatus can be a multi-channel decoder of the terminal device, the wireless device, or the core network device.
[0085] As shown in FIG. 3a, a schematic diagram of the multi-channel encoder and the multi-channel decoder provided by the embodiment of the present application applied to terminal devices is shown, and each terminal device can include a multi-channel encoder, a channel encoder, a multi-channel decoder and a channel decoder. The multi-channel encoder can perform the audio encoding method provided by the embodiment of the present application, and the multi-channel decoder can perform the audio decoding method provided by the embodiment of the present application. Specifically, the channel encoder is configured to perform channel encoding on the multi-channel signal, and the channel decoder is configured to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 can include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303 and a first channel decoder 304. The second terminal device 31 can include a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313 and a second channel encoder 314. The first terminal device 30 is connected to a first network communication device 32 via a wireless or wired connection, the first network communication device 32 and a second network communication device 33 are connected via a digital channel, and the second terminal device 31 is connected to the second network communication device 33 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device. In the audio communication, the terminal device as a sending end performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and transmits the signal in the digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a multi-channel signal encoding stream, and then performs multi-channel decoding to restore the multi-channel signal, and the terminal device as the receiving end plays back the multi-channel signal.
[0086] As shown in FIG. 3b, a schematic diagram of the multi-channel encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, another audio decoder 352, a multi-channel encoder 353 and a channel encoder 354. The above-mentioned FIG. 2b is similar, and will not be described herein.
[0087] As shown in FIG. 3c, a schematic diagram of the multi-channel decoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, a multi-channel decoder 355, another audio encoder 356 and a channel encoder 354. The above-mentioned FIG. 2c is similar, and will not be described herein.
[0088] The audio encoding process can be part of a multi-channel encoder, and the audio decoding process can be part of a multi-channel decoder. For example, multi-channel encoding of a captured multi-channel signal can include processing the captured multi-channel signal to obtain an audio signal, and encoding the obtained audio signal according to the method provided in the embodiments of the present application. The decoding end decodes the multi-channel signal encoding stream to obtain an audio signal, and recovers the multi-channel signal after upmix processing. Therefore, the embodiments of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In a wireless device or a core network device, if transcoding is required, corresponding multi-channel encoding processing needs to be performed.
[0089] The embodiments of the present application involve two or more terminals that need to communicate, for example, two terminals that communicate, or multiple terminals that communicate. In the following embodiments, two terminals that communicate are taken as an example for description, for example, a calling terminal and a called terminal. The calling terminal can be the encoding end as described above, and the called terminal can be the decoding end. The called terminal can be the encoding end as described above, and the calling terminal can be the decoding end. When two terminals in a mobile network establish a communication, because the voice Codec types supported by the two terminals can be different, and the code rate, sampling rate, etc. supported by the same voice Codec can also be different, the voice Codec needs to be negotiated before the communication is established, to determine which voice Codec and which encoding and decoding configuration are used for the communication. The multiple codecs can include EVS, AMR-WB, and IVAS. The transmission rate can also be referred to as the encoding and decoding rate.
[0090] The negotiation is performed using the Session Description Protocol (SDP). The negotiation process includes negotiation of four important parameters: load type, sampling frequency, rate, and packet length. Specifically:
[0091] The load type is a value defined in a standard protocol.
[0092] The higher the sampling frequency, the more sampling points, and the closer the voice quality to the real voice signal.
[0093] The rate is related to the bandwidth occupancy rate. The higher the rate, the higher the bandwidth occupancy rate.
[0094] The packet length refers to the voice duration contained in each voice packet. The larger the packet length, the larger the packet delay, but the stronger the anti-jitter capability and the higher the bandwidth utilization rate.
[0095] The parameters involved in the negotiation are: codec type and rate, sampling frequency and packet length, which are fixed for each codec type. In the SDP negotiation, the codec type is negotiated first, and then the rate. For AMR-WB and EVS, adaptive rate is supported, and the adaptive range supported is selected during the negotiation.
[0096] An SDP example is as follows:
[0097] a=rtpmap:100 AMR / 8000 / / 100 is the payload type, AMR represents the codec type, and 8000 represents the sampling frequency
[0098] a=fmtp:100 mode-set=0,2,4,7; mode-change-neighbor=1; mode-change-period=2 / / mode-set represents the rate set, rate adjustment period and adjustment mode, etc.
[0099] a=ptime:20 / / packet length.
[0100] First, a method for negotiating audio encoding / decoding provided by an embodiment of the present application is introduced. The method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in the embodiment of the present application can refer to audio encoding, or audio decoding, or audio encoding and decoding. Hereinafter, the audio encoding / decoding refers to the above-mentioned encoding and decoding processes.
[0101] As shown in FIG. 4, the method for negotiating audio encoding / decoding mainly includes the following steps:
[0102] 401. The second terminal device acquires mode control parameters of the second terminal device, wherein the mode control parameters of the second terminal device at least include user-set parameters of the second terminal device.
[0103] In the embodiment of the present application, the second terminal device determines the mode control parameters of the second terminal device, and the mode control parameters of the second terminal device can include user configuration information of multiple terminal devices. The types and configuration parameters of the user-set parameters involved in the embodiment of the present application are not limited.
[0104] 402. The second terminal device determines audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device, wherein the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0105] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, so that the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device, for example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.
[0106] The audio encoding / decoding capability of the second terminal device refers to the capability of the second terminal device itself, for example, the audio encoding / decoding capability can include the computing power, memory, etc. of the second terminal device.
[0107] For example, the audio encoding / decoding configuration information can be a codec mode configuration table.
[0108] 403. The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0109] The second terminal device can send the audio encoding / decoding configuration information of the second terminal device to the first terminal device, so that the first terminal device can perform audio encoding / decoding negotiation.
[0110] It is not limited that in the embodiment of the application, the first terminal device can also send the audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device can perform audio encoding / decoding negotiation.
[0111] 411. The first terminal device obtains the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0112] 412. The first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, and the mode control parameter of the first terminal device at least includes: the setting parameter of the first terminal device by the user.
[0113] In the embodiment of the application, the first terminal device determines the mode control parameter of the first terminal device, and the mode control parameter of the first terminal device can include hardware information of multiple terminal devices. The types and hardware configuration parameters of the hardware involved in the embodiment of the application are not limited.
[0114] 413. The first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the encoding / decoding negotiation result includes: the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0115] The first terminal device can obtain the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device respectively, and the first terminal device can perform the encoding / decoding negotiation to obtain the encoding / decoding negotiation result of the first terminal device.
[0116] The encoding / decoding negotiation result includes the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0117] 414. The first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0118] 404. The first terminal device receives the encoding / decoding negotiation result of the first terminal device, and the encoding / decoding negotiation result includes the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0119] In the embodiments of the present application, the first terminal device and the second terminal device can perform static audio negotiation according to the mode control parameters of each terminal device. For example, the negotiation is based on the hardware capability, processor and bottom chip of the terminal device, or based on the microphone and memory of the terminal device.
[0120] In some embodiments of the present application, the first encoding / decoding parameter includes the first encoding / decoding mode performed by the first terminal device when encoding / decoding the audio signal to be processed and the first encoding / decoding rate corresponding to the first encoding / decoding mode.
[0121] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device, including:
[0122] When the mode control parameter of the first terminal device meets the preset first hardware configuration condition, the first audio encoding / decoding configuration information is determined according to the mode control parameter of the first terminal device;
[0123] When the mode control parameter of the first terminal device does not meet the first hardware configuration condition, the second audio encoding / decoding configuration information is determined according to the mode control parameter of the first terminal device;
[0124] The audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0125] In the embodiments of the present application, different mode control parameters of the first terminal device indicate different audio encoding / decoding capabilities, so as to complete the audio encoding / decoding negotiation based on the mode control parameter.
[0126] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device, including:
[0127] The audio encoding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device and the audio encoding / decoding capability of the first terminal device.
[0128] When the first terminal device determines the audio encoding / decoding configuration information in the embodiments of the present application, in addition to using the mode control parameter of the first terminal device, the audio encoding / decoding capability of the first terminal device can also be used, so that the audio encoding / decoding configuration information of the first terminal device can be more accurately determined.
[0129] In some embodiments of the present application, the mode control parameter further includes at least one of the following: the battery endurance mode of the first terminal device, the audio call user experience of the first terminal device, and the power distribution mode of the first terminal device set by the user.
[0130] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device, including:
[0131] The mode control parameter of the first terminal device includes the first battery power set by the user for the first terminal device, and the audio encoding / decoding configuration information of the first terminal device is updated according to the first battery power set by the user for the first terminal device to obtain the updated audio encoding / decoding configuration information of the first terminal device.
[0132] In the embodiments of the present application, different battery parameters of the first terminal device indicate different battery endurance information, for example, when the battery endurance information indicates more power, stronger audio encoding / decoding capability can be used, and when the battery endurance information indicates that the power decreases, the audio encoding / decoding capability can be reduced, so as to complete the audio encoding / decoding negotiation based on the mode control parameter.
[0133] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is updated according to the first battery power set by the user for the first terminal device, including:
[0134] A first power interval to which the first battery power belongs is determined;
[0135] A first audio encoding / decoding mode corresponding to the first power interval is determined according to the correspondence between the power interval and the audio encoding / decoding mode of the first terminal device;
[0136] adjust the audio coding / decoding configuration information of the first terminal device according to the audio coding / decoding capability corresponding to the first audio coding / decoding mode.
[0137] In the embodiments of the present application, the first terminal device indicates different audio coding / decoding capabilities according to the correspondence between the power interval and the audio coding / decoding mode of the first terminal device, thereby completing the audio coding / decoding negotiation based on the mode control parameter.
[0138] In some embodiments of the present application, the updating of the audio coding / decoding configuration information of the first terminal device according to the first battery power set by the user for the first terminal device comprises:
[0139] When the first battery power is lower than the preset first power threshold and the first terminal device is in a battery charging mode, the audio coding / decoding capability indicated by the audio coding / decoding configuration information of the first terminal device is maintained or increased.
[0140] In the embodiments of the present application, when the first battery power is lower than the preset first power threshold and the first terminal device is in a battery charging mode, thereby completing the audio coding / decoding negotiation based on the mode control parameter.
[0141] In some embodiments of the present application, the determining of the audio coding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0142] The mode control parameter of the first terminal device comprises a user habit parameter set by the user for the first terminal device, and the audio coding / decoding configuration information of the first terminal device is determined according to the user habit parameter set by the user for the first terminal device.
[0143] In the embodiments of the present application, the user habit parameter set by the first terminal device for the first terminal device indicates different audio coding / decoding capabilities, thereby completing the audio coding / decoding negotiation based on the mode control parameter.
[0144] In some embodiments of the present application, the determining of the audio coding / decoding configuration information of the first terminal device according to the user habit parameter set by the user for the first terminal device comprises:
[0145] determining a first estimated standby time length of the first terminal device according to user schedule information of the first terminal device;
[0146] determining a second audio coding / decoding mode corresponding to the first estimated standby time length according to the correspondence between the standby time length and the audio coding / decoding mode of the first terminal device;
[0147] The audio coding / decoding configuration information of the first terminal device is determined according to an audio coding / decoding capability corresponding to the second audio coding / decoding mode.
[0148] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the user habit parameter set by the user for the first terminal device, and the method further comprises:
[0149] The terminal capability distribution information is determined according to the user operation habit information of the first terminal device.
[0150] The third audio coding / decoding mode is determined according to a correspondence between the terminal capability distribution and the audio coding / decoding mode of the first terminal device.
[0151] The audio coding / decoding configuration information of the first terminal device is determined according to an audio coding / decoding capability corresponding to the third audio coding / decoding mode.
[0152] In some embodiments of the present application, the method further comprises:
[0153] The network packet loss rate of the first terminal device for audio communication is acquired.
[0154] When the network packet loss rate is higher than a first network signal threshold, the audio coding / decoding configuration information of the first terminal device is updated.
[0155] In some embodiments of the present application, the method further comprises:
[0156] The audio service standard type supported by the audio coding / decoding chip of the second terminal device is acquired.
[0157] When the audio service standard type supported by the audio coding / decoding chip of the second terminal device and the audio service standard type supported by the audio coding / decoding chip of the first terminal device are both a preset first audio service standard type, the step of acquiring the audio coding / decoding configuration information of the second terminal device is triggered.
[0158] The audio service standard type can include at least one of EVS, IVAS, AMR, etc., and the audio service standard type is negotiated before the audio negotiation, so that more accurate audio negotiation can be achieved.
[0159] In some embodiments of the present application, the audio coding / decoding configuration information of the second terminal device comprises a coding / decoding mode supported by the second terminal device and a rate set corresponding to the coding / decoding mode supported by the second terminal device.
[0160] The audio coding / decoding configuration information of the first terminal device comprises a coding / decoding mode supported by the first terminal device and a rate set corresponding to the coding / decoding mode supported by the first terminal device.
[0161] For example, the coding mode can include MONO / STEREO / FOA / MASA, and the rate can have different values, and the specific value depends on the corresponding configuration information.
[0162] In some embodiments of the present application, the coding / decoding negotiation result of the first terminal device is determined according to the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device, including:
[0163] The audio coding / decoding negotiation order is determined according to the audio coding / decoding negotiation priority policy of the first terminal device, and the audio coding / decoding negotiation order includes: performing audio coding negotiation first and then performing audio decoding negotiation, or performing audio decoding negotiation first and then performing audio coding negotiation.
[0164] The audio coding negotiation includes: determining the coding negotiation result of the first terminal device according to the audio coding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device.
[0165] The audio decoding negotiation includes: determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio coding capability of the second terminal device.
[0166] In the embodiments of the present application, the order of audio coding negotiation and audio decoding negotiation is not limited, and a more flexible audio negotiation mode can be realized.
[0167] In some embodiments of the present application, the coding negotiation result of the first terminal device is determined according to the audio coding capability indicated by the audio coding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device, including:
[0168] According to the first audio mode supported by the highest coding mode of the first terminal device, and the decoding mode supported by the second terminal device includes the first audio mode, the coding negotiation result of the first terminal device includes the first audio mode.
[0169] The first bit rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device is determined.
[0170] The coding negotiation result of the first terminal device includes the maximum bit rate in the first bit rate intersection.
[0171] In the embodiments of the present application, the coding negotiation result includes the maximum bit rate in the first bit rate intersection, according to the principle of maximizing experience, the coding bit rate with higher experience effect of the opposite terminal is preferentially selected, and the call quality is improved.
[0172] In some embodiments of the present application, the decoding negotiation result of the first terminal device is determined according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device, including:
[0173] According to the highest decoding mode supported by the first terminal device being the second audio mode and the encoding mode supported by the second terminal device including the second audio mode, the decoding negotiation result of the first terminal device includes the second audio mode;
[0174] Determining the second code rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device;
[0175] Determining that the decoding negotiation result of the first terminal device includes the maximum code rate in the second code rate intersection.
[0176] In the embodiments of the present application, the decoding negotiation result includes the maximum code rate in the second code rate intersection, according to the principle of maximizing experience, the encoding code rate with higher experience effect is preferentially selected for the opposite terminal, and the call quality is improved.
[0177] In some embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0178] In the embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device, so that the maximum code rate of the first terminal device can be fully utilized, the maximum use of the audio encoding / decoding capability of the first terminal device is improved, and the call quality is improved.
[0179] In some embodiments of the present application, the first encoding / decoding parameter includes at least two of the following: mono (MONO), stereo (STEREO), metadata assisted spatial audio (MASA), high order ambisonic (HOA), or first order ambisonic (FOA).
[0180] The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0181] In some embodiments of the present application, the method further includes:
[0182] The first terminal device sends an audio negotiation request to the second terminal device.
[0183] The first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains mode control parameters of the second terminal device according to triggering of the audio negotiation request.
[0184] In some embodiments of the present application, the audio negotiation request comprises audio encoding / decoding configuration information of the first terminal device.
[0185] The capability negotiation request sent by the first terminal device to the second terminal device carries the audio encoding / decoding configuration information of the first terminal device, and the second terminal device can preliminarily select the audio encoding / decoding configuration information of the second terminal device according to the received audio encoding / decoding configuration information of the first terminal device. The audio encoding / decoding configuration information of the second terminal device obtained after the preliminary selection can be one or more audio encoding / decoding configuration information after the preliminary selection, and then the first terminal device performs audio negotiation, thereby further improving the efficiency of audio negotiation.
[0186] Further comprising: the capability negotiation request sent by the first terminal device to the second terminal device can also not comprise the audio encoding / decoding configuration information of the first terminal device, which is not limited here.
[0187] The foregoing embodiments illustrate the method performed by the first terminal device, and the method performed by the second terminal device is described as follows. The configuration mode of the mode control parameters of the second terminal device is similar to the configuration mode of the mode control parameters of the first terminal device, and the specific process of determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device is not described in detail.
[0188] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control of the second terminal device comprises:
[0189] The audio encoding / decoding configuration information of the second terminal device is determined according to the mode control parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0190] In some embodiments of the present application, the mode control parameters of the second terminal device comprise at least one of the following: a battery endurance mode of the second terminal device, an audio call user experience of the second terminal device, and a power distribution mode of the second terminal device.
[0191] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0192] obtaining battery endurance information of the second terminal device according to a battery parameter of the second terminal device;
[0193] updating audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and an audio encoding / decoding capability of the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0194] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0195] when the mode control parameter of the second terminal device meets a preset first software configuration condition, determining first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;
[0196] when the mode control parameter of the second terminal device does not meet the preset first software configuration condition, determining second audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;
[0197] wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0198] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0199] the mode control parameter of the second terminal device comprises a first battery power set by a user for the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power set by the user for the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
[0200] In some embodiments of the present application, the updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power set by the user for the second terminal device comprises:
[0201] determining a first power interval to which the second battery power belongs;
[0202] determining a first audio encoding / decoding mode corresponding to the first power interval according to a correspondence between power intervals and audio encoding / decoding modes of the second terminal device;
[0203] adjusting the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.
[0204] In some embodiments of the present application, the updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery level set by the user for the second terminal device comprises:
[0205] When the first battery level is lower than the preset first battery level threshold and the second terminal device is in a battery charging mode, the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device is maintained or increased.
[0206] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0207] The mode control parameter of the second terminal device comprises a user habit parameter set by the user for the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter set by the user for the second terminal device.
[0208] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter set by the user for the second terminal device comprises:
[0209] The first estimated standby time length of the second terminal device is determined according to user schedule information of the second terminal device;
[0210] The second audio encoding / decoding mode corresponding to the first estimated standby time length is determined according to a correspondence between standby time lengths and audio encoding / decoding modes of the second terminal device;
[0211] The audio encoding / decoding configuration information of the second terminal device is determined according to an audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.
[0212] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter set by the user for the second terminal device comprises:
[0213] The terminal capability allocation information is determined according to user operation habit information of the second terminal device;
[0214] The third audio encoding / decoding mode is determined according to a correspondence between terminal capability allocation and audio encoding / decoding modes of the second terminal device;
[0215] The audio encoding / decoding configuration information of the second terminal device is determined according to an audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.
[0216] In some embodiments of the present application, the method further comprises:
[0217] obtaining a network packet loss rate at which the second terminal device conducts an audio call;
[0218] updating audio encoding / decoding configuration information of the second terminal device when the network packet loss rate is higher than a first network signal threshold.
[0219] In some embodiments of the present application, the method further comprises:
[0220] receiving an audio negotiation request from the first terminal device.
[0221] In some embodiments of the present application, the audio negotiation request comprises audio encoding / decoding configuration information of the first terminal device.
[0222] The method further comprises: screening audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device.
[0223] The sending of the audio encoding / decoding configuration information of the second terminal device to the first terminal device comprises sending the candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0224] Next, an actual application scenario is taken as an example for description.
[0225] The method provided by the embodiments of the present application is applicable to IVAS encoding and decoding, and different signal types and corresponding rates can be used, and further negotiation is required. The embodiments of the present application provide a negotiation method for IVAS encoding and decoding modes and rates.
[0226] Each terminal supporting IVAS has an IVAS encoding and decoding configuration table owned by itself, which describes all IVAS encoding-decoding combinations that can be supported by the terminal under the hardware constraints. By transmitting the encoding and decoding configuration table in the negotiation process, the call parties can more flexibly negotiate the call configuration, such as the type of encoded signal, the code rate, etc. In addition, different signal types and different code rates can be supported for the call between the call parties. The present application can provide the best user experience for both parties as much as possible under the condition of meeting the hardware constraints of both parties.
[0227] Each IVAS-supporting mobile phone will maintain two tables according to the capabilities (computing power, memory, etc.) of the mobile phone itself. One table is an IVAS encoding and decoding mode configuration table, as shown in Table 1 below:
[0228] The IVAS codec mode configuration table can represent all possible combinations of IVAS encoding modes and decoding modes that the mobile phone can support, including signal types and bit rates for encoding or decoding. As an example in Table 1 above, the mobile phone supports three types of decoded signals, namely MONO, STEREO, and FOA, with the maximum bit rate of 9.6 kbps for MONO decoding, 32 kbps for STEREO decoding, and 48 kbps for FOA decoding. When the mobile phone decodes a MONO signal, it can support three types of encoded signals, namely MONO, STEREO, and FOA, with the maximum bit rates of 9.6 kbps, 32 kbps, and 48 kbps, respectively. When the mobile phone decodes a STEREO signal, it can still support three types of encoded signals, namely MONO, STEREO, and FOA, but since the decoding overhead for STEREO is larger than that for MONO, the remaining computing power for encoding will decrease, and the encoding bit rate for FOA cannot support 48 kbps, but only up to 24 kbps (the higher the bit rate, the larger the overhead). When the mobile phone decodes a FOA signal, since the decoding overhead for FOA is even larger, the remaining computing power cannot support FOA encoding, and the mobile phone can only support two types of encoded signals, namely MONO and STEREO, with the maximum bit rates of 9.6 kbps and 32 kbps, respectively.
[0229] As can be seen from the above example, the combination of encoding and decoding modes is limited by the capabilities of the terminal / chip. If the encoding overhead is large, the decoding overhead will decrease, and vice versa.
[0230] Another table is the encoding and decoding bit rate set table, as shown in Table 2 below:
[0231] The encoding and decoding bit rate set table can represent all encoding bit rates supported by IVAS for each type of signal. By combining the maximum bit rate supported by each encoding / decoding signal in the encoding and decoding mode configuration table and the encoding and decoding bit rate set table, the bit rate range supported by each encoding / decoding signal in the encoding and decoding mode configuration table can be obtained.
[0232] Each mobile phone will have a default IVAS encoding and decoding configuration table when it is shipped. Due to factors such as chip computing power, the IVAS encoding and decoding configuration table of different mobile phone platforms may be different.
[0233] Since the IVAS signal type that the mobile phone can encode is also related to some external factors, for example, only when the mobile phone has FOA or HOA acquisition equipment can it support the encoding of FOA / HOA, for example, only when the voice service contains multi-channel or object audio / sound effect signals can it support multi-channel or object encoding, for example, only when the mobile phone supports the complete MASA solution can it support the encoding of MASA signals, for example, when the mobile phone has low power, it may limit the encoding and decoding of some high complexity signals, and so on, so the mobile phone needs to dynamically update its IVAS encoding and decoding configuration table to reflect the current actual capability of the mobile phone.
[0234] As shown in FIG. 5, taking the audio negotiation process of terminal UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) as an example, an example is described:
[0235] Before A and B voice calls, if IVAS voice encoding and decoding is used for communication, media negotiation is needed, which includes:
[0236] 1. Audio encoding and decoding negotiation (EVS / IVAS / AMR).
[0237] 2. When the audio encoding and decoding negotiation is IVAS, further negotiation of encoding mode (MONO / STEREO / FOA / MASA) and code rate (each encoding mode has a corresponding code rate range) is needed, and the encoding mode and the decoding mode can be negotiated to be different.
[0238] Negotiation process:
[0239] 1. When UE-A initiates a call INVITE, if it supports IVAS voice encoding and decoding, it sends its supported IVAS encoding and decoding configuration table to B in the SDP information of IVAS, including the list of supported encoding mode and decoding mode combinations of A, the maximum rate supported by each mode, and the rate set supported by each encoding mode.
[0240] 2. UE-B processes as follows according to its own capability:
[0241] 1) First, encoding and decoding negotiation is performed, such as IVAS, which is preliminarily negotiated as IVAS;
[0242] 2) Continue to negotiate the encoding mode and the rate. According to the local strategy, decide whether to negotiate according to the encoding capability or the decoding capability, taking the priority of the encoding capability negotiation as an example:
[0243] ① Determine the encoding mode of UE-B according to the highest encoding mode supported by UE-B and whether UE-A supports the decoding mode;
[0244] ② According to the intersection of the supported bit rates of the calling and called parties for the coding mode, determine the UE-B negotiated coding bit rate. If there is no intersection, re-negotiate the coding mode of UE-B;
[0245] ③ According to the decoding mode corresponding to the UE-B supported coding mode list, and whether the UE-A supports the coding mode, determine the UE-B negotiated decoding mode;
[0246] ④ According to the intersection of the bit rates of the calling and called parties for the decoding mode, determine the UE-B negotiated decoding bit rate.
[0247] If any of the IVAS negotiated codec modes fails to negotiate, IVAS will not be used finally.
[0248] As shown in FIGS. 6 and 7, the SDP example adds the following parameters in the offer a line of IVAS:
[0249] Bit rate set:
[0250] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.
[0251] [mono,5.9 / 7.2 / 9.6] indicates that the mono supported bit rate set is 5.9, 7.2 and 9.6.
[0252] Decoding mode set:
[0253] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.
[0254] The list indicates that 3 sets are supported, each set corresponds to the following coding mode set one by one, set 1 corresponds to the decoding mode [mono / 9.6], indicating that mono is supported, and the highest bit rate is 9.6.
[0255] Coding mode set:
[0256] enc-mode-list={[mono / 9.6,stereo / 32,foa / 48],[mono / 9.6,stereo / 32,foa / 24],[mono / 9.6,stereo / 32]}.
[0257] The list represents support for 3 sets, one-to-one corresponding to the above decoding mode set, set 1 corresponds to the encoding mode = {[mono / 9.6, stereo / 32, foa / 48]} indicates that the decoding mode from low to high is mono / stereo / foa, and the highest code rate is 9.6 / 32 / 48.
[0258] Add the following parameters in the answer a line of IVAS:
[0259] The negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]};
[0260] The negotiated encoding mode and rate set: enc-mode = {[foa, 24 / 32]}.
[0261] Negotiation process example:
[0262] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the encoding and decoding modes supported by A and the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0263] 2. Mobile phone B decides according to its local strategy whether to negotiate according to its encoding capability or according to its decoding capability, taking the encoding capability negotiation as an example:
[0264] ① According to the highest encoding mode supported by UE-B, which is FOA, and the decoding mode of UE-A containing FOA, it is determined that the encoding mode of UE-B after negotiation is FOA;
[0265] ② Determine the intersection of the code rates supported by the FOA mode of the main and called parties, and determine that the encoding code rate of UE-B after negotiation is 24 and 32;
[0266] ③ According to the decoding mode corresponding to the encoding mode column supported by UE-B, which is MONO, and the encoding mode of UE-A containing MONO, it is determined that the decoding mode of UE-B after negotiation is MONO;
[0267] ④ Determine the intersection of the code rates supported by the MONO mode of the main and called parties, and determine that the decoding code rate of UE-B after negotiation is 5.9, 7.2 and 9.6.
[0268] As shown in FIG. 8, another negotiation process example:
[0269] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the encoding and decoding modes supported by A and the maximum rate combination supported by each mode, and the rate set supported by each encoding mode.
[0270] 2. Mobile phone B decides whether to negotiate based on its own encoding capabilities or its own decoding capabilities according to its local policy. Taking encoding capability negotiation as an example:
[0271] ① Based on the fact that the highest decoding mode supported by UE-B is FOA, and the decoding mode of UE-A includes FOA, it is determined that the decoding mode negotiated by UE-B is FOA;
[0272] ② Determine the intersection of the code rates of the FOA modes supported by the calling and called parties, and determine that the UE-B negotiated code rates are 24, 32, and 48;
[0273] ③ Based on the fact that the highest encoding mode corresponding to the decoding mode column supported by UE-B is STEREO, and the decoding mode of UE-A includes STEREO, it is determined that the encoding mode negotiated by UE-B is STEREO.
[0274] ④ Determine the intersection of the code rates of the STEREO modes supported by the calling and called parties, and determine that the decoding code rates after UE-B negotiation are 13.2 and 24.
[0275] As illustrated by the foregoing examples, for real-time communication audio codecs with numerous encoding modes and rates, and a wide range of encoding / decoding complexity, each terminal maintains its own encoding / decoding configuration table. This table describes the various combinations of encoding and decoding modes / rates that the terminal can support. When two or more terminals establish a call and negotiate this codec, they select their respective suitable encoding configurations based on their encoding / decoding configuration tables.
[0276] Each terminal has a factory-default codec configuration table. When establishing a call and negotiating the codec, the terminal can update its own codec configuration table based on other factors and negotiate the codec according to the updated codec configuration table. The updated codec configuration table does not overwrite the default factory codec configuration table.
[0277] In the codec configuration table, supported decoding modes / bitrates and / or encoding modes / bitrates are sorted by experience priority.
[0278] During codec negotiation, one terminal sends its complete or partial codec configuration table to the other party for negotiation. The other terminal then makes a negotiation decision based on all or part of the codec configuration tables from at least two parties. Here, "partial encoding mode and rate" refers to a subset of the encoding modes and rates.
[0279] The decision to negotiate can be based on the principle of maximizing the user experience, or it can be based on other principles, such as the decision-making principles preset in the terminal that makes the negotiation decision, such as prioritizing the other party's terminal for coding with a higher user experience.
[0280] Next, an audio negotiation process for terminal power is illustrated by way of example.
[0281] When the terminal power is 100%, the configuration table of the terminal is the initial configuration table, for example, as shown in Table 3 below:
[0282] When the terminal power decreases, in order to ensure the operation of other functions of the terminal, the configuration table needs to be adjusted to reduce the configuration. For example, when the power is only 10%, the configuration table is adjusted as shown in Table 4 below:
[0283] When the terminal only uses configuration 1 (default hardware configuration / factory configuration), the configuration table of the terminal is the initial configuration table, as shown in Table 3 above.
[0284] When the terminal uses configuration 2 (external FOA microphone), the configuration table is modified for some combinations of Table 3.
[0285] When the terminal uses configuration 3 (external HOA microphone), the configuration table is to replace combination 3 in Table 3 with HOA and the corresponding code rate.
[0286] When the terminal uses configuration 4 (external microphone), the terminal side device can convert the microphone signal to the Codec supported FOA / HOA / MASA signal, and the configuration table includes combination 3 as FOA and the corresponding code rate, combination 4 as HOA and the corresponding code rate, and combination 5 as MASA and the corresponding code rate.
[0287] Next, the application scenarios and corresponding solutions involved in the embodiments of the present application are illustrated by way of example. The computing power and memory of the chip (Complexity and RAM / ROM Memory in the IVAS document) can support a fixed application mode under a given hardware configuration.
[0288] Assuming that the IVAS standard can support a total of 150 working modes, and the low-power mode generally corresponds to low call quality or low immersion effect. Assuming that the lowest complexity working mode is 1, and the maximum complexity mode is 150, and the others are arranged in ascending order of complexity.
[0289] The maximum computing power and memory requirement of IVAS is about ten times that of EVS, but assuming that a mobile phone product may have only six times the hardware capability of EVS, it can only support, for example, the first 80 working modes, which is a limitation of the maximum complexity working mode caused by the hardware configuration.
[0290] In the case of full battery, the phone can support 1-80 working modes, but when the battery is reduced to 50%, it may selectively support 1-50 working modes, and give up support for high complexity working modes. When the battery is reduced to, for example, 20%, the supported working mode can be further reduced to 1-30, thereby maximizing the use time of the phone.
[0291] Another condition is that the user terminal can determine the most ideal power consumption mode according to the user's schedule. For example, the user's schedule today is not busy, and only has one work arrangement and a short time, so according to the user's historical use data, it is inferred that the battery power today is sufficient, at this time, the phone application configuration can maintain the high-quality call mode (1-80), and the user's schedule tomorrow is very tight, for example, needs to occupy the phone terminal for a telephone conference, so the phone configuration is more conservative (such as only 1-40) from the beginning of the day, to ensure that the user can maintain the power of the phone terminal without charging within a day, avoiding the situation of power alarm of the phone terminal.
[0292] Another case is that the user is currently talking, and the power is seriously insufficient. The judgment conclusion of the automatic configuration mechanism of the phone is to use a low-power mode (for example, 1-30), but the user himself determines that there will be a charging opportunity in a short time, so the user can set the phone to continue to remain in the high-quality call mode (1-80).
[0293] In another implementation scenario, due to the busy network of the phone, the signal received by the user can be extremely unstable, and the call is frequently interrupted. At this time, the best solution is to switch to a low-rate working mode. This is a decision that the network negotiation mechanism should participate in, but the terminal cannot perceive the characteristics of the signal received by the user. At this time, it is hoped that the user's phone can intelligently participate in the decision of audio negotiation.
[0294] In the embodiment of the application, IVAS can realize intelligent adjustment due to its multi-functional and multi-mode capability, thereby further improving the user experience.
[0295] For example, when the user allows the collection of historical information of the user operating the terminal, the user's use habits can be determined by artificial intelligence (AI) or statistical methods, for example, the first type of user operating the terminal for games occupies more power, the second type of user operating the terminal for video calls, and the third user operating the terminal for video calls. After the phone collects the user's use habits, the user's characteristics can be determined, and the use habits are likely to remain unchanged. The best working mode is determined for the user.
[0296] It should be noted that, for the foregoing method embodiments, for the purpose of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0297] In order to better implement the above scheme of the embodiments of the present application, the related device for implementing the above scheme is also provided below.
[0298] Please refer to Fig. 9, the first terminal device 900 provided by the embodiments of the present application can include: an acquisition module 901, a determination module 902, a negotiation module 903 and a sending module 904, wherein,
[0299] The acquisition module is configured to acquire audio encoding / decoding configuration information of a second terminal device, wherein the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0300] The determination module is configured to determine the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, wherein the mode control parameter of the first terminal device at least includes the setting parameter of the first terminal device set by the user.
[0301] The negotiation module is configured to determine the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, wherein the encoding / decoding negotiation result includes the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0302] The sending module is configured to send the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0303] As can be known from the foregoing examples, first, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, the mode control parameter of the first terminal device at least including: the setting parameter of the first terminal device set by the user; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including: the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiment of the application, both the two terminal devices for audio negotiation have the audio encoding / decoding configuration information, which describes the encoding / decoding capability that can be supported by the terminal under the constraint of the mode control parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiment of the application, the optimal user experience can be provided for the two parties as much as possible under the condition of meeting the constraint of the mode control parameter of the two parties of the call.
[0304] As shown in FIG. 10, the second terminal device 1000 provided by the embodiment of the application can include: an acquisition module 1001, a determination module 1002, a sending module 1003 and a receiving module 1004, wherein,
[0305] The acquisition module is used to acquire the mode control parameter of the second terminal device, the mode control parameter of the second terminal device at least including: the setting parameter of the second terminal device set by the user.
[0306] The determination module is used to determine the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device.
[0307] The sending module is used to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0308] The receiving module is used to receive the encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result including: the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device.
[0309] It should be noted that the information interaction, execution process and the like between the modules / units of the apparatus are based on the same concept as the method embodiments of the present application, and the technical effects brought by the same are the same as the method embodiments of the present application. For details, refer to the description in the foregoing method embodiments of the present application, which will not be repeated here.
[0310] The embodiments of the present application further provide a computer storage medium, wherein the computer storage medium stores a program, and the program performs part or all of the steps recorded in the foregoing method embodiments.
[0311] Next, another first terminal device provided by the embodiments of the present application is introduced. Referring to FIG. 11, the first terminal device 1100 includes:
[0312] The receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 (wherein the number of the processor 1103 in the first terminal device 1100 can be one or more, and one processor is taken as an example in FIG. 11). In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 11.
[0313] The memory 1104 can include read-only memory and random access memory, and provide the processor 1103 with instructions and data. A part of the memory 1104 can also include non-volatile random access memory (NVRAM). The memory 1104 stores an operating system and operation instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0314] The processor 1103 controls the operation of the first terminal device, and the processor 1103 can also be referred to as a central processing unit (CPU). In specific applications, various components of the first terminal device are coupled together through a bus system, wherein the bus system can include a data bus, a power bus, a control bus and a status signal bus, etc. in addition to the data bus. However, in order to clearly illustrate, all kinds of buses are referred to as a bus system in the figure.
[0315] The method disclosed in the embodiments of the present application can be applied to the processor 1103 or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the method can be completed by hardware integrated logic circuits in the processor 1103 or by instructions in the form of software. The processor 1103 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the storage 1104, and the processor 1103 reads the information in the storage 1104 and combines the hardware to complete the steps of the above method.
[0316] The receiver 1101 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the first terminal device. The transmitter 1102 can include a display device such as a display screen, and the transmitter 1102 can be used to output digital or character information through an external interface.
[0317] In the embodiments of the present application, the processor 1103 is configured to execute the method performed by the first terminal device shown in Fig. 4.
[0318] Next, another second terminal device provided by the embodiments of the present application is introduced. Please refer to Fig. 12, the second terminal device 1200 includes:
[0319] The receiver 1201, the transmitter 1202, the processor 1203 and the storage 1204 (wherein the number of the processor 1203 in the second terminal device 1200 can be one or more, and one processor is taken as an example in Fig. 12). In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the storage 1204 can be connected by a bus or other means, wherein the connection by the bus is taken as an example in Fig. 12.
[0320] The memory 1204 can include read-only memory and random access memory, and provide instructions and data to the processor 1203. A portion of the memory 1204 can also include a NVRAM. The memory 1204 stores operating systems and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0321] The processor 1203 controls the operation of the second terminal device, and the processor 1203 can also be referred to as a CPU. In a specific application, various components of the second terminal device are coupled together through a bus system, which can include a data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, various buses are referred to as a bus system in the figure.
[0322] The method disclosed in the embodiments of the present application can be applied in the processor 1203 or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with a signal processing capability. In the implementation process, the steps of the above method can be completed by hardware integrated logic circuits in the processor 1203 or by instructions in the form of software. The processor 1203 described above can be a general purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and combines the hardware to complete the steps of the above method.
[0323] In the embodiments of the present application, the processor 1203 is configured to execute the method performed by the second terminal device shown in Fig. 7 of the preceding embodiments.
[0324] In another possible design, when the first terminal device or the second terminal device is a chip in the terminal, the chip includes a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute computer-executed instructions stored in a storage unit, so that the chip in the terminal performs the method of any one of the first aspect. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the terminal, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), etc.
[0325] The processor mentioned in any one of the above can be a general central processor, a microprocessor, an ASIC, or one or more integrated circuits for controlling execution of the programs of the method of the first aspect or the second aspect.
[0326] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, in the apparatus embodiment provided in the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0327] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits, or special circuits, etc. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the embodiments of the present application.
[0328] In the above-described embodiments, all or some of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the embodiments can be implemented in the form of a computer program product.
[0329] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.
Claims
1. An audio encoding / decoding negotiation method, characterized by, The method is applied to a second terminal device, and the method comprises: obtaining mode control parameters of the second terminal device, wherein the mode control parameters of the second terminal device at least comprise setting parameters of the second terminal device by a user; determining audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device, wherein the audio encoding / decoding configuration information of the second terminal device is used to indicate audio encoding / decoding capability of the second terminal device; sending the audio encoding / decoding configuration information of the second terminal device to a first terminal device; receiving encoding / decoding negotiation results from the first terminal device, wherein the encoding / decoding negotiation results comprise first encoding / decoding parameters of the first terminal device for encoding / decoding audio signals to be processed, and the first encoding / decoding parameters are adapted to the audio encoding / decoding capability of the second terminal device.
2. The method of claim 1, wherein, The method further comprises: determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device and the audio encoding / decoding capability of the second terminal device.
3. The method according to claim 1 or 2, characterized in that, The mode control parameters of the second terminal device comprise at least one of the following: a battery endurance mode of the second terminal device, an audio call user experience of the second terminal device, and a power distribution mode of the second terminal device by the user.
4. The method of claim 2, wherein, The method further comprises: obtaining battery endurance information of the second terminal device according to the battery parameters of the second terminal device; updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
5. The method according to claim 2 or 3, characterized in that, The method further comprises: when the mode control parameters of the second terminal device satisfy a preset first software configuration condition, determining first audio encoding / decoding configuration information according to the mode control parameters of the second terminal device; when the mode control parameters of the second terminal device do not satisfy the preset first software configuration condition, determining second audio encoding / decoding configuration information according to the mode control parameters of the second terminal device; wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
6. The method according to claim 2 or 3, characterized in that, The method further comprises: The mode control parameter of the second terminal device comprises a first battery power set by a user for the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power set by the user for the second terminal device, to obtain updated audio encoding / decoding configuration information of the second terminal device.
7. The method of claim 6, wherein, The updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery power set by the user for the second terminal device comprises: determining a first power interval to which the second battery power belongs; determining a first audio encoding / decoding mode corresponding to the first power interval according to a correspondence between power intervals and audio encoding / decoding modes of the second terminal device; adjusting the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.
8. The method according to any one of claims 1 to 7, characterized in that, The updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery power set by the user for the second terminal device comprises: when the first battery power is lower than a preset first power threshold and the second terminal device is in a battery charging mode, maintaining or increasing an audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device.
9. The method according to any one of claims 1 to 8, characterized in that, The determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises: The mode control parameter of the second terminal device comprises a user habit parameter set by a user for the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter set by the user for the second terminal device.
10. The method of claim 9, wherein, The determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter set by the user for the second terminal device comprises: determining a first estimated standby time length of the second terminal device according to user schedule information of the second terminal device; determining a second audio encoding / decoding mode corresponding to the first estimated standby time length according to a correspondence between standby time lengths and audio encoding / decoding modes of the second terminal device; determining the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the second audio encoding / decoding mode.
11. The method of claim 9, wherein, The determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter set by the user for the second terminal device comprises: determining terminal capability allocation information according to user operation habit information of the second terminal device; determining a third audio encoding / decoding mode according to a correspondence between terminal capability allocation and audio encoding / decoding modes of the second terminal device; determining the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the third audio encoding / decoding mode.
12. The method according to any one of claims 1 to 11, characterized in that, The method further comprises: obtaining a network packet loss rate of the second terminal device in an audio call; when the network packet loss rate is higher than a first network signal threshold, updating the audio encoding / decoding configuration information of the second terminal device.
13. The method according to any one of claims 1 to 12, characterized in that, The method further comprises: receiving an audio negotiation request from the first terminal device.
14. The method of claim 13, wherein, The audio negotiation request comprises audio encoding / decoding configuration information of the first terminal device. The method further comprises: screening the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device. The sending of the audio encoding / decoding configuration information of the second terminal device to the first terminal device comprises: sending the candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
15. A terminal device, comprising: The terminal device is specifically a second terminal device, and the method comprises: An obtaining module, configured to obtain mode control parameters of the second terminal device, the mode control parameters of the second terminal device at least comprising: setting parameters of the second terminal device set by a user; A determining module, configured to determine audio encoding / decoding configuration information of the second terminal device according to the mode control parameters of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate audio encoding / decoding capability of the second terminal device; A sending module, configured to send the audio encoding / decoding configuration information of the second terminal device to a first terminal device; A receiving module, configured to receive a coding / decoding negotiation result from the first terminal device, the coding / decoding negotiation result comprising: first coding / decoding parameters of the first terminal device for encoding / decoding of an audio signal to be processed, the first coding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
16. A terminal device, comprising: The terminal device comprises at least one processor, which is coupled with a memory, reads and executes instructions in the memory to implement the method in any one of claims 1 to 14.
17. The terminal device of claim 16, wherein, The terminal device further comprises the memory.
18. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to carry out the method of any one of claims 1 to 14.