Audio encoding / decoding negotiation method and terminal device
By obtaining the audio codec configuration information of both terminals and conducting precise codec negotiation based on the mode control parameters, the problems of poor audio codec negotiation effect and wasted capacity in the existing technology are solved, and better audio quality and resource utilization are achieved.
Patent Information
- Application Number
- PCT/CN2025/075119
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-01-26
- Publication Date
- 2025-11-27
AI Technical Summary
Existing audio codec negotiation methods cannot achieve optimal codec performance when selecting highly complex speech codecs, resulting in stable call connections but poor audio quality, and also waste of codec capabilities.
By obtaining the audio codec configuration information of both terminals, the codec configuration is determined according to their respective mode control parameters, precise codec negotiation is carried out, and suitable codec parameters are selected to optimize the call experience and save power.
It achieves the optimal user experience while satisfying the control parameter constraints of each mode, avoids wasting encoding and decoding capabilities, and improves the encoding and decoding quality of audio signals.
Smart Images

Figure CN2025075119_27112025_PF_FP_ABST
Abstract
Description
Method for negotiating audio codec and terminal device
[0001] The present application claims priority from the Chinese patent application No. 202410354992.1 filed on March 26, 2024, and entitled "Method for negotiating audio codec and terminal device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of audio codec, and in particular, to a method for negotiating audio codec and a terminal device. BACKGROUND
[0003] Compared with a speech codec using a single channel, an immersive speech codec supports more types of audio signals, supports a larger range of code rates, and includes rendering characteristics in addition to codec characteristics. For example, in the audio codec standard being developed by the 3rd Generation Partnership Project (3GPP), different audio codec standards support multiple types of signals, different types of signals support different numbers of channels and code rates, and the rendering characteristics of different types of signals support rendering of various types of signals to binaural and standard loudspeaker arrays. Due to supporting more signal channel numbers, a larger range of code rates, and additional rendering characteristics, the immersive speech codec has a very significant increase in maximum computational complexity and storage complexity compared to the traditional speech codec.
[0004] In order to support a speech codec with high complexity, in an existing method for negotiating audio codec, the speech codec is classified according to complexity. However, when two terminals supporting different levels establish a call, the current negotiation result can only select a level with lower complexity for the call to ensure that the call connection is stable. Although this method meets the speech requirements of the call connection, it has the problem of poor audio codec negotiation effect. SUMMARY
[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a method for negotiating audio codec, the method being applied to a first terminal device, and the method comprising:
[0007] obtaining audio codec configuration information of a second terminal device, the audio codec configuration information of the second terminal device being used to indicate an audio codec capability of the second terminal device;
[0008] determining audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, wherein the mode control parameter of the first terminal device at least comprises user description parameter generated by the first terminal device;
[0009] determining encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, wherein the encoding / decoding negotiation result comprises first encoding / decoding parameter of the first terminal device for encoding / decoding audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0010] In some embodiments of the present application, the mode control parameter of the first terminal device is used to indicate configuration information of terminal device running mode used by the first terminal device when performing audio encoding / decoding.
[0011] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0012] determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device and the audio encoding / decoding capability of the first terminal device.
[0013] In some embodiments of the present application, the mode control parameter of the first terminal device comprises at least one of the following: battery endurance mode of the first terminal device, audio call user experience of the first terminal device, and user's power distribution mode for the first terminal device.
[0014] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0015] when the mode control parameter of the first terminal device satisfies preset first software configuration condition, determining first audio encoding / decoding configuration information according to the mode control parameter of the first terminal device;
[0016] when the mode control parameter of the first terminal device does not satisfy preset first software configuration condition, determining second audio encoding / decoding configuration information according to the mode control parameter of the first terminal device;
[0017] wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0018] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0019] The mode control parameter of the first terminal device comprises a first battery power collected by the first terminal device in a user-allowed case, and the audio encoding / decoding configuration information of the first terminal device is updated according to the first battery power collected by the first terminal device in the user-allowed case to obtain updated audio encoding / decoding configuration information of the first terminal device.
[0020] In some embodiments of the present application, the updating the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device in the user-allowed case comprises:
[0021] determining a first power interval to which the first battery power belongs;
[0022] determining a first audio encoding / decoding mode corresponding to the first power interval according to a correspondence between power intervals and audio encoding / decoding modes of the first terminal device;
[0023] adjusting the audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the first audio encoding / decoding mode.
[0024] In some embodiments of the present application, the updating the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device in the user-allowed case comprises:
[0025] when the first battery power is lower than a preset first power threshold and the first terminal device is in a battery charging mode, maintaining or increasing an audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device.
[0026] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0027] The mode control parameter of the first terminal device comprises a user habit parameter collected by the first terminal device in a user-allowed case, and the audio encoding / decoding configuration information of the first terminal device is determined according to the user habit parameter collected by the first terminal device in the user-allowed case.
[0028] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the first terminal device according to the user habit parameter collected by the first terminal device in the user-allowed case comprises:
[0029] determining a first estimated standby time length of the first terminal device according to user schedule information of the first terminal device;
[0030] determining a second audio codec mode corresponding to the first estimated standby time length according to a correspondence between standby time length and audio codec mode of the first terminal device;
[0031] determining audio encoding / decoding configuration information of the first terminal device according to audio encoding / decoding capability corresponding to the second audio codec mode.
[0032] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the first terminal device according to the user habit parameters collected by the first terminal device with user permission comprises:
[0033] determining terminal capability distribution information according to user operation habit information of the first terminal device;
[0034] determining a third audio codec mode according to a correspondence between terminal capability distribution and audio codec mode of the first terminal device;
[0035] determining audio encoding / decoding configuration information of the first terminal device according to audio encoding / decoding capability corresponding to the third audio codec mode.
[0036] In some embodiments of the present application, the method further comprises:
[0037] obtaining a network packet loss rate of the first terminal device for audio communication;
[0038] updating the audio encoding / decoding configuration information of the first terminal device when the network packet loss rate is higher than a first network signal threshold.
[0039] In some embodiments of the present application, the determining of the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device comprises:
[0040] determining an audio encoding / decoding negotiation order according to an audio encoding / decoding negotiation priority strategy of the first terminal device, the audio encoding / decoding negotiation order comprising: performing audio encoding negotiation before audio decoding negotiation, or performing audio decoding negotiation before audio encoding negotiation;
[0041] wherein the performing of the audio encoding negotiation comprises: determining an encoding negotiation result of the first terminal device according to audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and audio decoding capability of the second terminal device;
[0042] The performing the audio decoding negotiation comprises determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device.
[0043] In some embodiments of the present application, the determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device comprises:
[0044] determining the encoding negotiation result of the first terminal device to include a first audio mode according to that the highest encoding mode supported by the first terminal device is the first audio mode and that the decoding mode supported by the second terminal device includes the first audio mode;
[0045] determining a first code rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device;
[0046] determining the encoding negotiation result of the first terminal device to include a maximum code rate in the first code rate intersection.
[0047] In some embodiments of the present application, the determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device comprises:
[0048] determining the decoding negotiation result of the first terminal device to include a second audio mode according to that the highest decoding mode supported by the first terminal device is the second audio mode and that the encoding mode supported by the second terminal device includes the second audio mode;
[0049] determining a second code rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device;
[0050] determining the decoding negotiation result of the first terminal device to include a maximum code rate in the second code rate intersection.
[0051] In some embodiments of the present application, a sum of a code rate corresponding to the highest encoding mode supported by the first terminal device and a code rate corresponding to an audio decoding mode performed by the first terminal device is less than a maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0052] In some embodiments of the present application, the method further comprises:
[0053] sending an audio negotiation request to the second terminal device.
[0054] In some embodiments of the present application, the audio negotiation request comprises audio encoding / decoding configuration information of the first terminal device.
[0055] In a second aspect, embodiments of the present application provide a terminal device, specifically a first terminal device, and the method comprises:
[0056] The obtaining module is configured to obtain audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate audio encoding / decoding capability of the second terminal device.
[0057] The determining module is configured to determine audio encoding / decoding configuration information of the first terminal device according to mode control parameters of the first terminal device, the mode control parameters of the first terminal device at least comprising user description parameters generated by the first terminal device.
[0058] The negotiating module is configured to determine encoding / decoding negotiation results of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation results comprising first encoding / decoding parameters of the first terminal device for encoding / decoding audio signals to be processed, the first encoding / decoding parameters being adapted to the audio encoding / decoding capability of the second terminal device.
[0059] The sending module is configured to send the encoding / decoding negotiation results of the first terminal device to the second terminal device.
[0060] In the second aspect of the present application, the composing module of the first terminal device can also perform the steps described in the foregoing first aspect and various possible implementation manners, for details, refer to the foregoing description of the first aspect and various possible implementation manners.
[0061] In a third aspect, embodiments of the present application provide a computer readable storage medium, wherein instructions are stored in the computer readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the method in the foregoing first aspect.
[0062] In a fourth aspect, embodiments of the present application provide a computer program product comprising instructions, and when the instructions are run on a computer, the computer is caused to execute the method in the foregoing first aspect.
[0063] In a fifth aspect, embodiments of the present application provide a communication apparatus, which can comprise a terminal device or a chip and the like entity, and the communication apparatus comprises a processor and a memory, the memory is configured to store instructions, and the processor is configured to execute the instructions in the memory, so that the communication apparatus executes the method in any one of the foregoing first aspect.
[0064] In a sixth aspect, the present application provides a chip system, which comprises a processor for supporting a terminal device to implement functions involved in the above aspects, such as transmitting or processing data and / or information involved in the above methods. In a possible design, the chip system further comprises a memory, and the memory is configured to store necessary program instructions and data of the terminal device. The chip system can be composed of a chip, or can comprise a chip and other discrete devices.
[0065] In a seventh aspect, an embodiment of the present application provides a chip, comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of an electronic device and send a signal to the processor, and the signal comprises computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the method in the first aspect or any possible implementation manner of the first aspect.
[0066] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0067] First, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, and the mode control parameter of the first terminal device at least comprises a user description parameter generated by the first terminal device; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, and the encoding / decoding negotiation result comprises a first encoding / decoding parameter for the first terminal device to encode / decode the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiments of the present application, both the two terminal devices for audio negotiation have audio encoding / decoding configuration information, and the audio encoding / decoding configuration information describes the encoding / decoding capability that the terminal can support under the constraint of the mode control parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiments of the present application, the two parties can be provided with the optimal user experience as much as possible under the constraint of the mode control parameter of the two parties. BRIEF DESCRIPTION OF DRAWINGS
[0068] FIG. 1 is a schematic diagram of the composition structure of an audio processing system provided by an embodiment of the present application;
[0069] Fig. 2a is a schematic diagram of an audio encoder and an audio decoder applied to a terminal device according to an embodiment of the present application;
[0070] Fig. 2b is a schematic diagram of an audio encoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0071] Fig. 2c is a schematic diagram of an audio decoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0072] Fig. 3a is a schematic diagram of a multi-channel encoder and a multi-channel decoder applied to a terminal device according to an embodiment of the present application;
[0073] Fig. 3b is a schematic diagram of a multi-channel encoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0074] Fig. 3c is a schematic diagram of a multi-channel decoder applied to a wireless device or a core network device according to an embodiment of the present application;
[0075] Fig. 4 is a schematic diagram of a negotiation method of audio encoding / decoding according to an embodiment of the present application;
[0076] Fig. 5 is a schematic diagram of a negotiation process of audio negotiation by two terminals through SDP according to an embodiment of the present application;
[0077] Fig. 6 is a schematic diagram of a negotiation process of audio encoding / decoding mode and code rate by two terminals according to an embodiment of the present application;
[0078] Fig. 7 is a schematic diagram of a negotiation result of encoding / decoding after negotiation by two terminals according to an embodiment of the present application;
[0079] Fig. 8 is a schematic diagram of another negotiation process of audio encoding / decoding mode and code rate by two terminals according to an embodiment of the present application;
[0080] Fig. 9 is a schematic diagram of a structure of a first terminal device according to an embodiment of the present application;
[0081] Fig. 10 is a schematic diagram of a structure of a second terminal device according to an embodiment of the present application;
[0082] Fig. 11 is a schematic diagram of another structure of a first terminal device according to an embodiment of the present application;
[0083] Fig. 12 is a schematic diagram of another structure of a second terminal device according to an embodiment of the present application. DETAILED DESCRIPTION
[0084] Embodiments of the present application will be described below with reference to the accompanying drawings.
[0085] The terms "first", "second", and the like, in the description and in the claims of the present application and above drawings, are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the present application are capable of functioning in other sequences, except where it is inherent from the disclosure. Furthermore, the terms "comprise", "comprising", "include", "including", and the like, are specifically intended to be open-ended. These terms mean that a process, method, article, or apparatus that "comprises", "comprising", "include", or "including" something lacks one or more of the features that are specifically named, but that the process, method, article, or apparatus otherwise includes the specifically named features.
[0086] The bit rate of traditional monophonic speech codec is between several kbps to tens or 128 kbps. Immersive speech codec supports more signal types, supports a larger range of bit rates, and contains rendering features in addition to codec features, compared with monophonic speech codec. Taking the 3GPP ongoing Immersive Voice and Audio Services (IVAS) speech / audio codec standard as an example, the signal types supported by IVAS speech / audio codec include monophonic, stereo, multi-channel, multi-object, higher order ambisonics (HOA) signal or first order ambisonics (FOA) signal, MASA, etc. Among them, the multi-channel signal supports up to 7.1.4 channel format, i.e. 12 channels, and the HOA signal supports up to 3rd order HOA, i.e. 16 channels, and the supported bit rate ranges from the lowest single channel 5.9 kbps to 768 kbps for 3rd order HOA. Rendering supports rendering of various signal types to binaural and standard loudspeaker arrays. More signal channels, larger bit rate range, and additional rendering features make the maximum computational complexity and storage complexity of immersive speech codec significantly higher than that of traditional speech codec, which will lead to a significant increase in the difficulty of landing chips for such codecs.
[0087] In order to solve the problem of high complexity of the speech codec, the speech codec is classified according to complexity, but when two terminals supporting different levels establish a call, the current negotiation result is to select a lower complexity level for the call to ensure stable call connection. This method meets the speech requirements of the call connection, but has the problem of poor audio codec negotiation effect. In addition, this method meets the speech requirements of the call connection, but the lower level terminal cannot decode all types of code streams, which causes the coding and decoding limitation of the audio signal of the higher level terminal, and has the problem of poor audio signal coding and decoding quality. In addition, when the levels of the two terminals are not equal, the higher level terminal also has the problem of waste of coding and decoding capability.
[0088] Based on the above description, the embodiment of the present application provides an audio coding / decoding negotiation method, which can perform accurate Codec negotiation when establishing a call, and select the most appropriate coding and decoding combination according to the actual capability of the communication platform of each party, so as to ensure optimal call experience and not waste power consumption.
[0089] The embodiment of the present application provides an audio coding technology, in particular, provides a three-dimensional audio coding technology for three-dimensional audio signals, and specifically provides an encoding technology for representing three-dimensional audio signals by using fewer channels, so as to improve the traditional audio coding system. Audio coding (or commonly referred to as coding) includes two parts of audio encoding and audio decoding. Audio encoding is performed at the source side, including processing (for example, compressing) the original audio to reduce the amount of data required to represent the audio, so as to more efficiently store and / or transmit. Audio decoding is performed at the destination side, including inverse processing relative to the encoder, to reconstruct the original audio. The encoding part and the decoding part are also collectively referred to as coding. The implementation of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0090] The technical solution of the embodiment of the present application can be applied to various audio processing systems. As shown in FIG. 1, it is a schematic diagram of the composition structure of the audio processing system provided by the embodiment of the present application. The audio processing system 100 can include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 can be used to generate a code stream, and then the audio encoding code stream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the code stream, then perform the audio decoding function of the audio decoding device 102, and finally obtain the reconstructed signal.
[0091] In the embodiments of the present application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices and core network devices requiring transcoding, for example, the audio encoding device can be an audio encoder of the terminal device, wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices and core network devices requiring transcoding, for example, the audio decoding device can be an audio decoder of the terminal device, wireless device or core network device. For example, the audio encoder can include a wireless access network, a media gateway of a core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc., and the audio encoder can also be an audio encoder applied to a virtual reality (VR) streaming service.
[0092] In the embodiments of the present application, taking the audio encoding module (audio encoding and audio decoding) applied to a virtual reality streaming (VR streaming) service as an example, the processing flow of the end-to-end audio signal includes: the audio signal A is preprocessed (audio PReprocessing) after being acquired by the acquisition module, the preprocessing operation includes filtering out the low frequency part of the signal, which can be divided by 20Hz or 50Hz, and extracting the orientation information of the signal, then the encoding processing (audio encoding) is performed, the file / segment encapsulation is performed, and then the delivery is performed to the decoding end. The decoding end first performs file / segment decapsulation, then performs decoding (audio decoding), performs binaural rendering (audio rendering) processing on the decoded signal, and maps the rendered signal to the listener's headphones (headphones), which can be independent headphones or headphones on glasses devices.
[0093] As shown in FIG. 2a, a schematic diagram of the audio encoder and the audio decoder provided by the embodiment of the present application applied to a terminal device is shown. For each terminal device, an audio encoder, a channel encoder, an audio decoder and a channel decoder can be included. Specifically, the channel encoder is used to perform channel encoding on the audio signal, and the channel decoder is used to perform channel decoding on the audio signal. For example, the first terminal device 20 can include a first audio encoder 201, a first channel encoder 202, a first audio decoder 203 and a first channel decoder 204. The second terminal device 21 can include a second audio decoder 211, a second channel decoder 212, a second audio encoder 213 and a second channel encoder 214. The first terminal device 20 is connected to a first network communication device 22 via a wireless or wired connection, the first network communication device 22 and a second network communication device 23 are connected via a digital channel, and the second terminal device 21 is connected to the second network communication device 23 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device.
[0094] In the audio communication, the terminal device as a sending end first performs audio acquisition, encodes the acquired audio signal, and then performs channel encoding before transmitting the signal in a digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding according to the received signal to obtain a code stream, and then performs audio decoding to restore the audio signal, which is played back by the terminal device as the receiving end.
[0095] As shown in FIG. 2b, a schematic diagram of the audio encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown. The wireless device or the core network device 25 includes a channel decoder 251, another audio decoder 252, the audio encoder 253 provided by the embodiment of the present application, and a channel encoder 254. The another audio decoder 252 refers to an audio decoder other than the audio decoder. In the wireless device or the core network device 25, the signal entering the device is first decoded by the channel decoder 251, and then decoded by the another audio decoder 252, and then encoded by the audio encoder 253 provided by the embodiment of the present application, and finally encoded by the channel encoder 254 before being transmitted out. The another audio decoder 252 is used to decode the code stream decoded by the channel decoder 251.
[0096] As shown in FIG. 2c, the audio decoder provided by the embodiments of the present application is applied to a wireless device or a core network device. The wireless device or the core network device 25 comprises a channel decoder 251, the audio decoder 255 provided by the embodiments of the present application, another audio encoder 256, and a channel encoder 254, wherein the another audio encoder 256 refers to an audio encoder other than the audio encoder. In the wireless device or the core network device 25, the signal entering the device is firstly channel-decoded by the channel decoder 251, then the received audio encoded code stream is decoded by the audio decoder 255, then the audio is encoded by the another audio encoder 256, and finally the audio signal is channel-encoded by the channel encoder 254, and the channel-encoded audio signal is transmitted out. In the wireless device or the core network device, if transcoding is needed, corresponding audio encoding processing is needed. The wireless device refers to a radio frequency related device in communication, and the core network device refers to a core network related device in communication.
[0097] In some embodiments of the present application, the audio encoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices, for example, the audio encoding apparatus can be a multi-channel encoder of the terminal device, the wireless device, or the core network device. Similarly, the audio decoding apparatus can be applied to various terminal devices having audio communication needs, wireless devices having transcoding needs, and core network devices, for example, the audio decoding apparatus can be a multi-channel decoder of the terminal device, the wireless device, or the core network device.
[0098] As shown in FIG. 3a, a schematic diagram of the multi-channel encoder and the multi-channel decoder provided by the embodiment of the present application applied to terminal devices is shown, and each terminal device can include a multi-channel encoder, a channel encoder, a multi-channel decoder and a channel decoder. The multi-channel encoder can perform the audio encoding method provided by the embodiment of the present application, and the multi-channel decoder can perform the audio decoding method provided by the embodiment of the present application. Specifically, the channel encoder is configured to perform channel encoding on the multi-channel signal, and the channel decoder is configured to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 can include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303 and a first channel decoder 304. The second terminal device 31 can include a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313 and a second channel encoder 314. The first terminal device 30 is connected to a first network communication device 32 via a wireless or wired connection, the first network communication device 32 and a second network communication device 33 are connected via a digital channel, and the second terminal device 31 is connected to the second network communication device 33 via a wireless or wired connection. The network communication device can be a signal transmission device, such as a communication base station or a data exchange device. In the audio communication, the terminal device as a sending end performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and transmits the signal in the digital channel via a wireless network or a core network. The terminal device as a receiving end performs channel decoding on the received signal to obtain a multi-channel signal encoding stream, and then performs multi-channel decoding to restore the multi-channel signal, and the terminal device as the receiving end plays back the multi-channel signal.
[0099] As shown in FIG. 3b, a schematic diagram of the multi-channel encoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, another audio decoder 352, a multi-channel encoder 353 and a channel encoder 354. The above-mentioned wireless device or the core network device 35 is similar to the wireless device or the core network device 35 shown in FIG. 2b, and thus will not be described herein.
[0100] As shown in FIG. 3c, a schematic diagram of the multi-channel decoder provided by the embodiment of the present application applied to a wireless device or a core network device is shown, wherein the wireless device or the core network device 35 includes a channel decoder 351, a multi-channel decoder 355, another audio encoder 356 and a channel encoder 354. The above-mentioned wireless device or the core network device 35 is similar to the wireless device or the core network device 35 shown in FIG. 2c, and thus will not be described herein.
[0101] The audio encoding process can be part of a multichannel encoder, and the audio decoding process can be part of a multichannel decoder. For example, multichannel encoding of a captured multichannel signal can include processing the captured multichannel signal to obtain an audio signal, and encoding the obtained audio signal according to the method provided in the embodiments of the present application. The decoding end decodes the multichannel signal encoding stream to obtain an audio signal, and recovers the multichannel signal after upmix processing. Therefore, the embodiments of the present application can also be applied to multichannel encoders and multichannel decoders in terminal devices, wireless devices, and core network devices. In a wireless device or a core network device, if transcoding is required, corresponding multichannel encoding processing needs to be performed.
[0102] The embodiments of the present application involve two or more terminals that need to communicate, for example, two terminals that communicate, or multiple terminals that communicate. In the following embodiments, two terminals that communicate are taken as an example for description, for example, a calling terminal and a called terminal. The calling terminal can be the encoding end as described above, and the called terminal can be the decoding end. The called terminal can be the encoding end as described above, and the calling terminal can be the decoding end. When two terminals in a mobile network establish a communication, because the voice Codec types supported by the two terminals can be different, and the code rate, sampling rate, etc. supported by the same voice Codec can also be different, the voice Codec needs to be negotiated before the communication is established, to determine which voice Codec and which codec configuration are used for the communication. The multiple codecs can include EVS, AMR-WB, and IVAS. The transmission rate can also be referred to as the codec rate.
[0103] The negotiation is performed using the Session Description Protocol (SDP). The negotiation process includes negotiation of four important parameters: load type, sampling frequency, rate, and packet length. Specifically:
[0104] The load type is a value defined in a standard protocol.
[0105] The higher the sampling frequency, the more sampling points, and the closer the voice quality to the real voice signal.
[0106] The rate is related to the bandwidth occupancy rate. The higher the rate, the higher the bandwidth occupancy rate.
[0107] The packet length refers to the voice duration contained in each voice packet. The larger the packet length, the larger the packet delay, but the stronger the anti-jitter capability and the higher the bandwidth utilization rate.
[0108] The parameters involved in the negotiation are: codec type and rate, sampling frequency and packet length, which are fixed for each codec type. In the SDP negotiation, the codec type is negotiated first, and then the rate. For AMR-WB and EVS, adaptive rate is supported, and the adaptive range supported is selected during negotiation.
[0109] An SDP example is as follows:
[0110] a=rtpmap:100 AMR / 8000 / / 100 is the payload type, AMR represents the codec type, and 8000 represents the sampling frequency
[0111] a=fmtp:100 mode-set=0,2,4,7;mode-change-neighbor=1;mode-change-period=2 / / mode-set represents the rate set, rate adjustment period and adjustment mode, etc.
[0112] a=ptime:20 / / packet length.
[0113] First, a method for negotiating audio encoding / decoding provided by an embodiment of the present application is introduced. The method can be executed by a first terminal device and a second terminal device. For example, the first terminal device can be a calling terminal device, and the second terminal device can be a called terminal device. Alternatively, the second terminal device can be a calling terminal device, and the first terminal device can be a called terminal device. The audio encoding / decoding involved in the embodiment of the present application can refer to audio encoding, or audio decoding, or audio encoding and decoding. Hereinafter, “audio encoding / decoding” refers to the above-mentioned encoding and decoding processes.
[0114] As shown in FIG. 4, the method for negotiating audio encoding / decoding mainly includes the following steps:
[0115] 401. The second terminal device acquires the mode control parameter of the second terminal device, wherein the mode control parameter of the second terminal device at least includes a user description parameter generated by the first terminal device.
[0116] In the embodiment of the present application, the second terminal device determines the mode control parameter of the second terminal device, which can include a user description parameter generated by the terminal device. The user description parameter generated by the terminal device is generated according to the historical information of the user operating the terminal device under the condition that the user allows. The types and configuration parameters of the user setting parameter involved in the embodiment of the present application are not limited.
[0117] 402. The second terminal device determines the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, wherein the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0118] The second terminal device can determine the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, so that the audio encoding / decoding configuration information of the second terminal device can indicate the audio encoding / decoding capability of the second terminal device, for example, the audio encoding / decoding configuration information includes: audio encoding configuration information, audio decoding configuration information, or audio encoding and decoding configuration information.
[0119] The audio encoding / decoding capability of the second terminal device refers to the capability of the second terminal device itself, for example, the audio encoding / decoding capability can include the computing power, memory, etc. of the second terminal device.
[0120] For example, the audio encoding / decoding configuration information can be a codec mode configuration table.
[0121] 403. The second terminal device sends the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0122] The second terminal device can send the audio encoding / decoding configuration information of the second terminal device to the first terminal device, so that the first terminal device can negotiate the audio encoding / decoding.
[0123] It is not limited that in the embodiment of the application, the first terminal device can also send the audio encoding / decoding configuration information of the first terminal device to the second terminal device, and the second terminal device can negotiate the audio encoding / decoding.
[0124] 411. The first terminal device obtains the audio encoding / decoding configuration information of the second terminal device, and the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0125] 412. The first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, and the mode control parameter of the first terminal device at least includes: the user description parameter generated by the first terminal device.
[0126] In the embodiment of the application, the first terminal device determines the mode control parameter of the first terminal device, and the mode control parameter of the first terminal device can include hardware information of multiple terminal devices. The types and hardware configuration parameters of the hardware involved in the embodiment of the application are not limited.
[0127] 413、The first terminal device determines a coding / decoding negotiation result of the first terminal device according to the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device, the coding / decoding negotiation result comprising: first coding / decoding parameters of the first terminal device for coding / decoding the audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.
[0128] The first terminal device can obtain the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device respectively, and can perform coding / decoding negotiation to obtain the coding / decoding negotiation result of the first terminal device.
[0129] The coding / decoding negotiation result comprises: first coding / decoding parameters of the first terminal device for coding / decoding the audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.
[0130] 414、The first terminal device sends the coding / decoding negotiation result of the first terminal device to the second terminal device.
[0131] 404、The first terminal device receives the coding / decoding negotiation result of the first terminal device from the second terminal device, the coding / decoding negotiation result comprising: first coding / decoding parameters of the first terminal device for coding / decoding the audio signal to be processed, the first coding / decoding parameters being adapted to the audio coding / decoding capability of the second terminal device.
[0132] In the embodiments of the present application, the first terminal device and the second terminal device can perform static audio negotiation according to their respective mode control parameters. For example, negotiation is performed based on the hardware capability, processor and bottom chip of the terminal device, or based on the microphone and memory of the terminal device.
[0133] In some embodiments of the present application, the first coding / decoding parameters comprise: a first coding / decoding mode performed by the first terminal device when coding / decoding the audio signal to be processed, and a first coding / decoding rate corresponding to the first coding / decoding mode.
[0134] In some embodiments of the present application, determining the audio coding / decoding configuration information of the first terminal device according to the mode control parameters of the first terminal device comprises:
[0135] When the mode control parameters of the first terminal device satisfy a preset first hardware configuration condition, determining the first audio coding / decoding configuration information according to the mode control parameters of the first terminal device;
[0136] When the mode control parameters of the first terminal device do not satisfy the first hardware configuration condition, determining the second audio coding / decoding configuration information according to the mode control parameters of the first terminal device;
[0137] The audio coding / decoding capability corresponding to the first audio coding / decoding configuration information is higher than the audio coding / decoding capability corresponding to the second audio coding / decoding configuration information.
[0138] The different mode control parameters of the first terminal device in the embodiments of the present application indicate different audio coding / decoding capabilities, so that the audio coding / decoding negotiation based on the mode control parameters is completed.
[0139] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device, comprising:
[0140] The audio coding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device and the audio coding / decoding capability of the first terminal device.
[0141] In the embodiments of the present application, when the first terminal device determines the audio coding / decoding configuration information, in addition to using the mode control parameter of the first terminal device, the audio coding / decoding capability of the first terminal device can also be used, so that the audio coding / decoding configuration information of the first terminal device can be more accurately determined.
[0142] In some embodiments of the present application, the mode control parameter further comprises at least one of the following: the battery endurance mode of the first terminal device, the audio call user experience of the first terminal device, and the user's power distribution mode of the first terminal device.
[0143] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is determined according to the mode control parameter of the first terminal device, comprising:
[0144] The mode control parameter of the first terminal device comprises the first battery power collected by the first terminal device with the user's permission, and the audio coding / decoding configuration information of the first terminal device is updated according to the first battery power collected by the first terminal device with the user's permission, to obtain the updated audio coding / decoding configuration information of the first terminal device.
[0145] In the embodiments of the present application, different battery parameters of the first terminal device indicate different battery endurance information, for example, when the battery endurance information indicates more power, stronger audio coding / decoding capability can be used, and when the battery endurance information indicates that the power decreases, the audio coding / decoding capability can be reduced, so that the audio coding / decoding negotiation based on the mode control parameter is completed.
[0146] In some embodiments of the present application, the audio coding / decoding configuration information of the first terminal device is updated according to the first battery power collected by the first terminal device with the user's permission, comprising:
[0147] determining a first power interval to which the first battery power belongs;
[0148] determining a first audio codec mode corresponding to the first power interval according to a correspondence between power intervals and audio codec modes of the first terminal device;
[0149] adjusting audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the first audio codec mode.
[0150] In the embodiments of the present application, the first terminal device indicates different audio encoding / decoding capabilities according to the correspondence between power intervals and audio codec modes of the first terminal device, thereby completing the audio codec negotiation based on the mode control parameter.
[0151] In some embodiments of the present application, the updating of the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device with user permission comprises:
[0152] when the first battery power is lower than a preset first power threshold and the first terminal device is in a battery charging mode, maintaining or increasing the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device.
[0153] In the embodiments of the present application, the first battery power is lower than a preset first power threshold and the first terminal device is in a battery charging mode, thereby completing the audio codec negotiation based on the mode control parameter.
[0154] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises:
[0155] The mode control parameter of the first terminal device comprises a user habit parameter collected by the first terminal device with user permission, and the audio encoding / decoding configuration information of the first terminal device is determined according to the user habit parameter collected by the first terminal device with user permission.
[0156] In the embodiments of the present application, the user habit parameter set by the first terminal device of the first terminal device indicates different audio encoding / decoding capabilities, thereby completing the audio codec negotiation based on the mode control parameter.
[0157] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the first terminal device according to the user habit parameter collected by the first terminal device with user permission comprises:
[0158] determining a first estimated standby time length of the first terminal device according to user schedule information of the first terminal device;
[0159] determining a second audio codec mode corresponding to the first estimated standby time length according to a correspondence between standby time length and audio codec mode of the first terminal device;
[0160] determining audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the second audio codec mode.
[0161] In some embodiments of the present application, the audio encoding / decoding configuration information of the first terminal device is determined according to the user habit parameters collected by the first terminal device with user permission, comprising:
[0162] determining terminal capability allocation information according to user operation habit information of the first terminal device;
[0163] determining a third audio codec mode according to a correspondence between terminal capability allocation and audio codec mode of the first terminal device;
[0164] determining audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the third audio codec mode.
[0165] In some embodiments of the present application, the method further comprises:
[0166] obtaining a network packet loss rate of the first terminal device for audio communication;
[0167] updating the audio encoding / decoding configuration information of the first terminal device when the network packet loss rate is higher than a first network signal threshold.
[0168] In some embodiments of the present application, the method further comprises:
[0169] obtaining an audio service standard type supported by an audio codec chip of the second terminal device;
[0170] triggering the aforementioned step of obtaining audio encoding / decoding configuration information of the second terminal device when the audio service standard type supported by the audio codec chip of the second terminal device and the audio codec chip of the first terminal device is a preset first audio service standard type.
[0171] The audio service standard type can include at least one of EVS / IVAS / AMR, etc. The audio service standard type is negotiated before audio negotiation, which can achieve more accurate audio negotiation.
[0172] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device comprises: a codec mode supported by the second terminal device, and a rate set corresponding to the codec mode supported by the second terminal device.
[0173] The audio encoding / decoding configuration information of the first terminal device comprises: a codec mode supported by the first terminal device, and a rate set corresponding to the codec mode supported by the first terminal device.
[0174] For example, the codec mode can comprise MONO / STEREO / FOA / MASA, and the rate has different values, and the specific value depends on the corresponding configuration information.
[0175] In some embodiments of the present application, the encoding / decoding negotiation result of the first terminal device is determined according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, comprising:
[0176] The audio encoding / decoding negotiation order is determined according to the audio encoding / decoding negotiation priority policy of the first terminal device, and the audio encoding / decoding negotiation order comprises: performing audio encoding negotiation first or performing audio decoding negotiation first.
[0177] The audio encoding negotiation comprises: determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device.
[0178] The audio decoding negotiation comprises: determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device.
[0179] The order of the audio encoding negotiation and the audio decoding negotiation in the embodiments of the present application is not limited, and a more flexible audio negotiation mode can be realized.
[0180] In some embodiments of the present application, the encoding negotiation result of the first terminal device is determined according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device, comprising:
[0181] According to the first audio mode supported by the first terminal device and the decoding mode supported by the second terminal device, if the first audio mode is included in the decoding mode supported by the second terminal device, the encoding negotiation result of the first terminal device comprises the first audio mode.
[0182] The first rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device is determined.
[0183] The coding negotiation result of the first terminal device includes a maximum code rate in the first code rate intersection.
[0184] The coding negotiation result of the first terminal device includes a maximum code rate in the first code rate intersection.
[0185] In some embodiments of the present application, the decoding negotiation result of the first terminal device is determined according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device, including:
[0186] According to the second audio mode supported by the highest decoding mode supported by the first terminal device, and the second audio mode included in the encoding mode supported by the second terminal device, the decoding negotiation result of the first terminal device includes the second audio mode.
[0187] A second code rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device is determined.
[0188] The decoding negotiation result of the first terminal device includes a maximum code rate in the second code rate intersection.
[0189] According to the second audio mode supported by the highest decoding mode supported by the first terminal device, and the second audio mode included in the encoding mode supported by the second terminal device, the decoding negotiation result of the first terminal device includes the second audio mode.
[0190] In some embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0191] In some embodiments of the present application, the sum of the code rate corresponding to the highest encoding mode supported by the first terminal device and the code rate corresponding to the audio decoding mode executed by the first terminal device is less than the maximum code rate supported by the audio encoding / decoding capability of the first terminal device.
[0192] In some embodiments of the present application, the first encoding / decoding parameter includes at least two of the following: monaural MONO, stereo dual sound STEREO, metadata auxiliary spatial audio MASA, high-order ambisonic HOA, or first-order ambisonic FOA.
[0193] The rate set corresponding to the encoding / decoding mode includes at least two of the following: 9.6 kbps, 13.2 kbps, 32 kbps, and 48 kbps.
[0194] In some embodiments of the present application, the method further comprises:
[0195] The first terminal device sends an audio negotiation request to the second terminal device.
[0196] The first terminal device sends an audio negotiation request to the second terminal device, and the second terminal device obtains the mode control parameter of the second terminal device according to the trigger of the audio negotiation request.
[0197] In some embodiments of the present application, the audio negotiation request comprises audio encoding / decoding configuration information of the first terminal device.
[0198] The capability negotiation request sent by the first terminal device to the second terminal device carries the audio encoding / decoding configuration information of the first terminal device, and the second terminal device can preliminarily select the audio encoding / decoding configuration information of the second terminal device according to the received audio encoding / decoding configuration information of the first terminal device, which can be one or more audio encoding / decoding configuration information after preliminary screening, and then the first terminal device performs audio negotiation, thereby further improving the efficiency of audio negotiation.
[0199] Further comprising: the capability negotiation request sent by the first terminal device to the second terminal device can also not include the audio encoding / decoding configuration information of the first terminal device, which is not limited here.
[0200] The foregoing embodiments illustrate the method performed by the first terminal device, and the method performed by the second terminal device is described as follows. The configuration mode of the mode control parameter of the second terminal device is similar to the configuration mode of the mode control parameter of the first terminal device, and the specific process of determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device is not described in detail.
[0201] In some embodiments of the present application, the audio encoding / decoding configuration information of the second terminal device is determined according to the mode control of the second terminal device, comprising:
[0202] The audio encoding / decoding configuration information of the second terminal device is determined according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device.
[0203] In some embodiments of the present application, the mode control parameter of the second terminal device comprises at least one of the following: the battery endurance mode of the second terminal device, the audio call user experience of the second terminal device, and the power distribution mode of the second terminal device.
[0204] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device and the audio encoding / decoding capability of the second terminal device comprises:
[0205] obtaining the battery endurance information of the second terminal device according to the battery parameter of the second terminal device;
[0206] updating the audio encoding / decoding configuration information of the second terminal device according to the battery endurance information and the audio encoding / decoding capability of the second terminal device to obtain the updated audio encoding / decoding configuration information of the second terminal device.
[0207] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0208] when the mode control parameter of the second terminal device meets the preset first software configuration condition, determining the first audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;
[0209] when the mode control parameter of the second terminal device does not meet the preset first software configuration condition, determining the second audio encoding / decoding configuration information according to the mode control parameter of the second terminal device;
[0210] wherein the audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than the audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
[0211] In some embodiments of the present application, the determining the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0212] the mode control parameter of the second terminal device comprises the first battery power collected by the second terminal device in the case of user permission, and the audio encoding / decoding configuration information of the second terminal device is updated according to the first battery power collected by the second terminal device in the case of user permission to obtain the updated audio encoding / decoding configuration information of the second terminal device.
[0213] In some embodiments of the present application, the updating the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device in the case of user permission comprises:
[0214] determining the first power interval to which the second battery power belongs;
[0215] determining a first audio codec mode corresponding to the first battery power interval according to a correspondence relationship between the battery power interval and the audio codec mode of the second terminal device;
[0216] adjusting the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the first audio codec mode.
[0217] In some embodiments of the present application, the updating of the audio encoding / decoding configuration information of the second terminal device according to the first battery power collected by the second terminal device with user permission comprises:
[0218] when the first battery power is lower than a preset first battery power threshold and the second terminal device is in a battery charging mode, maintaining or increasing the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the second terminal device.
[0219] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device comprises:
[0220] The mode control parameter of the second terminal device comprises a user habit parameter collected by the second terminal device with user permission, and the audio encoding / decoding configuration information of the second terminal device is determined according to the user habit parameter collected by the second terminal device with user permission.
[0221] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter collected by the second terminal device with user permission comprises:
[0222] determining a first estimated standby time length of the second terminal device according to user schedule information of the second terminal device;
[0223] determining a second audio codec mode corresponding to the first estimated standby time length according to a correspondence relationship between the standby time length and the audio codec mode of the second terminal device;
[0224] adjusting the audio encoding / decoding configuration information of the second terminal device according to an audio encoding / decoding capability corresponding to the second audio codec mode.
[0225] In some embodiments of the present application, the determining of the audio encoding / decoding configuration information of the second terminal device according to the user habit parameter collected by the second terminal device with user permission comprises:
[0226] determining terminal capability allocation information according to user operation habit information of the second terminal device;
[0227] determining a third audio codec mode according to the correspondence between the terminal capability and the audio codec mode of the second terminal device;
[0228] determining the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding capability corresponding to the third audio codec mode.
[0229] In some embodiments of the present application, the method further comprises:
[0230] obtaining a network packet loss rate of the second terminal device for audio communication;
[0231] updating the audio encoding / decoding configuration information of the second terminal device when the network packet loss rate is higher than a first network signal threshold.
[0232] In some embodiments of the present application, the method further comprises:
[0233] receiving an audio negotiation request from the first terminal device.
[0234] In some embodiments of the present application, the audio negotiation request comprises the audio encoding / decoding configuration information of the first terminal device.
[0235] The method further comprises: screening the audio encoding / decoding configuration information of the second terminal device according to the audio encoding / decoding configuration information of the first terminal device to obtain candidate audio encoding / decoding configuration information of the second terminal device.
[0236] The sending of the audio encoding / decoding configuration information of the second terminal device to the first terminal device comprises sending the candidate audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0237] Next, an actual application scenario example is used for illustration.
[0238] The method provided by the embodiments of the present application is applicable to IVAS encoding and decoding, and different signal types and corresponding rates can be used, further negotiation is needed, and the embodiments of the present application provide a negotiation method for IVAS codec mode and rate.
[0239] Each terminal supporting IVAS has an IVAS codec configuration table owned by itself, and the configuration table describes all IVAS encoding-decoding combinations that can be supported by the terminal under the hardware constraints. By transmitting the codec configuration table in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, such as the type of encoded signal, the code rate, etc. In addition, different signal types and different code rates can be supported for the two parties of the call. The present application can provide the best user experience for the two parties as much as possible under the condition of meeting the hardware constraints of the two parties of the call.
[0240] Each IVAS-enabled phone will maintain two tables according to its own capability (computing power, memory, etc.). One table is the IVAS codec mode configuration table, as shown in Table 1 below:
[0241] The IVAS codec mode configuration table can represent all possible combinations of IVAS encoding and decoding modes that the phone can support, including the signal type and the bit rate. As an example in Table 1 above, the phone supports three types of decoded signals, which are MONO, STEREO, and FOA, with the maximum bit rate of 9.6 kbps for MONO decoding, 32 kbps for STEREO decoding, and 48 kbps for FOA decoding. When the phone decodes a MONO signal, it can support three types of encoded signals, which are MONO, STEREO, and FOA, with the maximum bit rate of 9.6 kbps, 32 kbps, and 48 kbps, respectively. When the phone decodes a STEREO signal, it can still support three types of encoded signals, which are MONO, STEREO, and FOA, but since the decoding overhead for STEREO is larger than that for MONO, the remaining computing power for encoding will decrease, and thus the encoding bit rate for FOA cannot be supported to 48 kbps, but only to 24 kbps (the higher the bit rate, the larger the overhead). When the phone decodes a FOA signal, since the decoding overhead for FOA is even larger, the remaining computing power for encoding is not enough, and thus the phone can only support two types of encoded signals, which are MONO and STEREO, with the maximum bit rate of 9.6 kbps and 32 kbps, respectively.
[0242] As can be seen from the above example, the combination of encoding and decoding modes is limited by the capability of the terminal / chip, and if the encoding overhead is large, the decoding overhead will decrease, and vice versa.
[0243] The other table is the codec bit rate set table, as shown in Table 2 below:
[0244] The codec bit rate set table can represent all encoding bit rates supported by IVAS for each type of signal. By combining the maximum bit rate supported by each encoding / decoding signal in the codec mode configuration table and the codec bit rate set table, the range of bit rates supported by each encoding / decoding signal in the codec mode configuration table can be obtained.
[0245] Each phone will have a default IVAS codec configuration table when it is shipped, and the IVAS codec configuration table for different phone platforms may be different depending on the chip computing power and other factors.
[0246] Since the IVAS signal type that the mobile phone can encode is also related to some external factors, for example, only when the mobile phone has FOA or HOA acquisition equipment can it support the encoding of FOA / HOA, for example, only when the voice service contains multi-channel or object audio / sound effect signals can it support the encoding of multi-channel or objects, for example, only when the mobile phone supports the complete MASA solution can it support the encoding of MASA signals, for example, when the mobile phone has low power, it may limit the encoding and decoding of some high complexity signals, etc., so the mobile phone needs to dynamically update its IVAS encoding and decoding configuration table to reflect the current actual ability of the mobile phone.
[0247] As shown in FIG. 5, taking the audio negotiation process of terminal UE-A (hereinafter referred to as A) and UE-B (hereinafter referred to as B) as an example, an example is described:
[0248] Before A and B voice calls, if IVAS voice encoding and decoding is used for communication, media negotiation is needed, which includes:
[0249] 1. Audio encoding and decoding negotiation (EVS / IVAS / AMR).
[0250] 2. After the audio encoding and decoding negotiation is IVAS, further negotiation of encoding mode (MONO / STEREO / FOA / MASA) and code rate (each encoding mode has a corresponding code rate range) is needed. The encoding mode and the decoding mode can be negotiated to be different.
[0251] Negotiation process:
[0252] 1. When UE-A initiates a call INVITE, if it supports IVAS voice encoding and decoding, it sends its supported IVAS encoding and decoding configuration table to B in the SDP information of IVAS, including the list of supported encoding mode and decoding mode combinations of A, the maximum rate supported by each mode, and the rate set supported by each encoding mode.
[0253] 2. UE-B processes as follows according to its own ability:
[0254] 1) First, perform encoding and decoding negotiation, such as supporting IVAS, then preliminarily negotiate IVAS;
[0255] 2) Continue to negotiate the encoding mode and the rate. According to the local strategy, decide whether to negotiate according to the encoding ability or the decoding ability, taking the priority of the encoding ability negotiation as an example:
[0256] ① Determine the encoding mode of UE-B according to the highest encoding mode supported by UE-B and whether UE-A supports the decoding mode;
[0257] ②According to the intersection of the supported bitrates of the encoding mode of the calling and called parties, determine the UE-B negotiated encoding bitrate. If there is no intersection, re-negotiate the encoding mode of UE-B;
[0258] ③According to the decoding mode corresponding to the encoding mode list supported by UE-B, and whether UE-A supports the encoding mode, determine the UE-B negotiated decoding mode;
[0259] ④According to the intersection of the bitrates of the decoding mode of the calling and called parties, determine the UE-B negotiated decoding bitrate.
[0260] If any of the IVAS codec negotiation fails, IVAS will not be used finally.
[0261] As shown in FIG. 6 and FIG. 7, the SDP example adds the following parameters in the offer a line of IVAS:
[0262] Bitrate set:
[0263] bite-rate-list={[mono,5.9 / 7.2 / 9.6],[stereo,13.2 / 24 / 32],[foa,24 / 32 / 48]}.
[0264] [mono,5.9 / 7.2 / 9.6] indicates that the bitrate set supported by mono is 5.9, 7.2 and 9.6.
[0265] Decoding mode set:
[0266] dec-mode-list={[mono / 9.6],[stereo / 32],[foa / 48]}.
[0267] The list indicates that 3 sets are supported, each set corresponds to the following encoding mode set one by one, set 1 corresponds to the decoding mode [mono / 9.6], which indicates that mono is supported and the highest bitrate is 9.6.
[0268] Encoding mode set:
[0269] enc-mode-list={[mono / 9.6,stereo / 32,foa / 48],[mono / 9.6,stereo / 32,foa / 24],[mono / 9.6,stereo / 32]}.
[0270] The list represents support for 3 sets, one-to-one corresponding to the above decoding mode set, set 1 corresponds to the encoding mode = {[mono / 9.6, stereo / 32, foa / 48]} indicates that the decoding mode from low to high is mono / stereo / foa, and the highest code rate is 9.6 / 32 / 48.
[0271] The following parameters are added in the answer a line of IVAS:
[0272] The negotiated decoding mode and rate set: dec-mode = {[mono, 5.9 / 7.2 / 9.6]};
[0273] The negotiated encoding mode and rate set: enc-mode = {[foa, 24 / 32]}.
[0274] An example of the negotiation process:
[0275] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the encoding and decoding modes supported by A and the maximum rate combination supported by each mode, as well as the rate set supported by each encoding mode.
[0276] 2. Mobile phone B decides according to its local strategy whether to negotiate according to its encoding ability or according to its decoding ability, taking the encoding ability negotiation as an example:
[0277] ① According to the highest encoding mode supported by UE-B, which is FOA, and the decoding mode of UE-A containing FOA, the encoding mode negotiated by UE-B is determined to be FOA;
[0278] ② Determine the intersection of the code rates supported by the FOA mode of the main and called, and determine the encoding code rate negotiated by UE-B to be 24 and 32;
[0279] ③ According to the decoding mode corresponding to the encoding mode column supported by UE-B, which is MONO, and the encoding mode of UE-A containing MONO, the decoding mode negotiated by UE-B is determined to be MONO;
[0280] ④ Determine the intersection of the code rates supported by the MONO mode of the main and called, and determine the decoding code rate negotiated by UE-B to be 5.9, 7.2 and 9.6.
[0281] As shown in FIG. 8, another example of the negotiation process:
[0282] 1. Mobile phone A sends its IVAS codec configuration table to mobile phone B, including the encoding and decoding modes supported by A and the maximum rate combination supported by each mode, as well as the rate set supported by each encoding mode.
[0283] 2. Mobile phone B decides whether to negotiate based on its own encoding capabilities or its own decoding capabilities according to its local policy. Taking encoding capability negotiation as an example:
[0284] ① Based on the fact that the highest decoding mode supported by UE-B is FOA, and the decoding mode of UE-A includes FOA, it is determined that the decoding mode negotiated by UE-B is FOA;
[0285] ② Determine the intersection of the code rates of the FOA modes supported by the calling and called parties, and determine that the UE-B negotiated code rates are 24, 32, and 48;
[0286] ③ Based on the fact that the highest encoding mode corresponding to the decoding mode column supported by UE-B is STEREO, and the decoding mode of UE-A includes STEREO, it is determined that the encoding mode negotiated by UE-B is STEREO.
[0287] ④ Determine the intersection of the code rates of the STEREO modes supported by the calling and called parties, and determine that the decoding code rates after UE-B negotiation are 13.2 and 24.
[0288] As illustrated by the foregoing examples, for real-time communication audio codecs with numerous encoding modes and rates, and a wide range of encoding / decoding complexity, each terminal maintains its own encoding / decoding configuration table. This table describes the various combinations of encoding and decoding modes / rates that the terminal can support. When two or more terminals establish a call and negotiate this codec, they select their respective suitable encoding configurations based on their encoding / decoding configuration tables.
[0289] Each terminal has a factory-default codec configuration table. When establishing a call and negotiating the codec, the terminal can update its own codec configuration table based on other factors and negotiate the codec according to the updated codec configuration table. The updated codec configuration table does not overwrite the default factory codec configuration table.
[0290] In the codec configuration table, supported decoding modes / bitrates and / or encoding modes / bitrates are sorted by experience priority.
[0291] During codec negotiation, one terminal sends its complete or partial codec configuration table to the other party for negotiation. The other terminal then makes a negotiation decision based on all or part of the codec configuration tables from at least two parties. Here, "partial encoding mode and rate" refers to a subset of the encoding modes and rates.
[0292] The decision to negotiate can be based on the principle of maximizing the user experience, or it can be based on other principles, such as the decision-making principles preset in the terminal that makes the negotiation decision, such as prioritizing the other party's terminal for coding with a higher user experience.
[0293] Next, an audio negotiation process for terminal power is illustrated by way of example.
[0294] When the terminal power is 100%, the configuration table of the terminal is the initial configuration table, for example, as shown in Table 3 below:
[0295] When the terminal power decreases, in order to ensure the operation of other functions of the terminal, the configuration table needs to be adjusted to reduce the configuration. For example, when the power is only 10%, the configuration table is adjusted as shown in Table 4 below:
[0296] When the terminal only uses configuration 1 (default hardware configuration / factory configuration), the configuration table of the terminal is the initial configuration table, as shown in Table 3 above.
[0297] When the terminal uses configuration 2 (external FOA microphone), the configuration table is modified for some combinations of Table 3.
[0298] When the terminal uses configuration 3 (external HOA microphone), the configuration table is to replace combination 3 in Table 3 with HOA and the corresponding code rate.
[0299] When the terminal uses configuration 4 (external microphone), the terminal side device can convert the microphone signal to the Codec supported FOA / HOA / MASA signal, and the configuration table includes combination 3 as FOA and the corresponding code rate, combination 4 as HOA and the corresponding code rate, and combination 5 as MASA and the corresponding code rate.
[0300] Next, the application scenarios and corresponding solutions involved in the embodiments of the present application are illustrated by way of example. The computing power and memory of the chip (Complexity and RAM / ROM Memory in the IVAS document) can support a fixed application mode under a given hardware configuration.
[0301] Assuming that the IVAS standard can support a total of 150 working modes, and the low-power mode generally corresponds to low call quality or low immersion effect. Assuming that the lowest complexity working mode is 1, and the maximum complexity mode is 150, and the others are arranged in ascending order of complexity.
[0302] The maximum computing power and memory requirement of IVAS is about ten times that of EVS, but assuming that a mobile phone product has only six times the hardware capability of EVS, it can only support, for example, the first 80 working modes, which is a limitation of the maximum complexity working mode caused by the hardware configuration.
[0303] In the case of full battery, the phone can support 1-80 working modes, but when the battery is reduced to 50%, it may selectively support 1-50 working modes, and give up support for high complexity working modes. When the battery is reduced to, for example, 20%, the supported working mode can be further reduced to 1-30, thereby maximizing the use time of the phone.
[0304] Another condition is that the user terminal can determine the most ideal power consumption mode according to the user's schedule. For example, the user's schedule today is not busy, and only has one work arrangement and a short time, so according to the user's historical use data, it is inferred that the battery power today is sufficient, at this time, the phone application configuration can maintain a high-quality call mode (1-80), and the user's schedule tomorrow is very tight, for example, needs to occupy the phone terminal for a telephone conference, so the phone configuration is more conservative (such as only 1-40) from the beginning of the day, to ensure that the user can maintain the power of the phone terminal without charging within a day, avoiding the situation of power alarm of the phone terminal.
[0305] Another case is that the user is currently talking, and the power is seriously insufficient. The judgment conclusion of the automatic configuration mechanism of the phone is to use a low-power mode (for example, 1-30), but the user himself determines that there will be a charging opportunity in a short time, so the user can set the phone to continue to remain in a high-quality call mode (1-80).
[0306] In another implementation scenario, due to the busy network of the phone, the signal received by the user can be extremely unstable, and the call is frequently interrupted. At this time, the best solution is to switch to a low-rate working mode. This is a decision that the network negotiation mechanism should participate in, but the terminal cannot perceive the characteristics of the signal received by the user. At this time, it is hoped that the user's phone can intelligently participate in the decision of audio negotiation.
[0307] In the embodiment of the application, IVAS can realize intelligent adjustment due to its multi-functional and multi-mode capability, thereby further improving the user experience.
[0308] For example, when the user allows the collection of historical information of the user operating the terminal, the user's use habits can be determined by artificial intelligence (AI) or statistical methods, for example, the first type of user operating the terminal for games consumes more power, the second type of user operating the terminal for video calls, and the third type of user operating the terminal for video calls. After the phone collects the user's use habits, the user's characteristics can be determined, and the use habits are likely to remain unchanged. The best working mode is determined for the user.
[0309] It should be noted that, for the foregoing method embodiments, for the purpose of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0310] To better implement the above scheme of the embodiments of the present application, the related device for implementing the above scheme is also provided.
[0311] Please refer to Fig. 9, the first terminal device 900 provided by the embodiments of the present application can include: an acquisition module 901, a determination module 902, a negotiation module 903 and a sending module 904, wherein,
[0312] The acquisition module is configured to acquire audio encoding / decoding configuration information of a second terminal device, wherein the audio encoding / decoding configuration information of the second terminal device is used to indicate the audio encoding / decoding capability of the second terminal device.
[0313] The determination module is configured to determine the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, wherein the mode control parameter of the first terminal device at least includes a user description parameter generated by the first terminal device.
[0314] The negotiation module is configured to determine the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, wherein the encoding / decoding negotiation result includes a first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, and the first encoding / decoding parameter is adapted to the audio encoding / decoding capability of the second terminal device.
[0315] The sending module is configured to send the encoding / decoding negotiation result of the first terminal device to the second terminal device.
[0316] As can be known from the foregoing examples, first, the first terminal device acquires the audio encoding / decoding configuration information of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device; then the first terminal device determines the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device, the mode control parameter of the first terminal device at least including: the user description parameter generated by the first terminal device; next, the first terminal device determines the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the encoding / decoding negotiation result including: the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device; finally, the first terminal device sends the encoding / decoding negotiation result of the first terminal device to the second terminal device. In the embodiment of the application, both the two terminals for audio negotiation have the audio encoding / decoding configuration information, which describes the encoding / decoding capability that can be supported by the terminal under the constraint of the mode control parameter. By transmitting the audio encoding / decoding configuration information in the negotiation process, the two parties of the call can more flexibly negotiate the call configuration, and in the embodiment of the application, the optimal user experience can be provided for the two parties as much as possible under the condition of meeting the constraint of the mode control parameter of the two parties of the call.
[0317] As shown in FIG. 10, the second terminal device 1000 provided by the embodiment of the application can include: an acquisition module 1001, a determination module 1002, a sending module 1003 and a receiving module 1004, wherein,
[0318] The acquisition module is used to acquire the mode control parameter of the second terminal device, the mode control parameter of the second terminal device at least including: the user description parameter generated by the first terminal device.
[0319] The determination module is used to determine the audio encoding / decoding configuration information of the second terminal device according to the mode control parameter of the second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate the audio encoding / decoding capability of the second terminal device.
[0320] The sending module is used to send the audio encoding / decoding configuration information of the second terminal device to the first terminal device.
[0321] The receiving module is used to receive the encoding / decoding negotiation result from the first terminal device, the encoding / decoding negotiation result including: the first encoding / decoding parameter of the first terminal device for encoding / decoding the audio signal to be processed, the first encoding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device.
[0322] It should be noted that the information interaction, execution process and the like between the modules / units of the apparatus are based on the same concept as the method embodiments of the present application, and the technical effects brought by the same are the same as the method embodiments of the present application. For details, refer to the description in the foregoing method embodiments of the present application, which will not be repeated here.
[0323] The embodiments of the present application further provide a computer storage medium, wherein the computer storage medium stores a program, and the program performs part or all of the steps recorded in the foregoing method embodiments.
[0324] Next, another first terminal device provided by the embodiments of the present application is introduced. Referring to FIG. 11, the first terminal device 1100 includes:
[0325] The receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 (wherein the number of the processor 1103 in the first terminal device 1100 can be one or more, and one processor is taken as an example in FIG. 11). In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 can be connected through a bus or other means, and the connection through the bus is taken as an example in FIG. 11.
[0326] The memory 1104 can include read-only memory and random access memory, and provide the processor 1103 with instructions and data. A part of the memory 1104 can also include non-volatile random access memory (NVRAM). The memory 1104 stores an operating system and operation instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0327] The processor 1103 controls the operation of the first terminal device, and the processor 1103 can also be referred to as a central processing unit (CPU). In specific applications, various components of the first terminal device are coupled together through a bus system, wherein the bus system can include a data bus, a power bus, a control bus and a state signal bus, etc. in addition to the data bus. However, in order to clearly illustrate, all kinds of buses are referred to as a bus system in the figure.
[0328] The method disclosed in the embodiments of the present application can be applied to the processor 1103 or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the method can be completed by hardware integrated logic circuits in the processor 1103 or by instructions in the form of software. The processor 1103 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the storage 1104, and the processor 1103 reads the information in the storage 1104 and combines the hardware to complete the steps of the above method.
[0329] The receiver 1101 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the first terminal device. The transmitter 1102 can include a display device such as a display screen. The transmitter 1102 can be used to output digital or character information through an external interface.
[0330] In the embodiments of the present application, the processor 1103 is configured to execute the method performed by the first terminal device shown in Fig. 4.
[0331] Next, another second terminal device provided by the embodiments of the present application is introduced. Please refer to Fig. 12, the second terminal device 1200 includes:
[0332] The receiver 1201, the transmitter 1202, the processor 1203 and the storage 1204 (wherein the number of the processor 1203 in the second terminal device 1200 can be one or more, and one processor is taken as an example in Fig. 12). In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the storage 1204 can be connected by a bus or other means, wherein the connection by the bus is taken as an example in Fig. 12.
[0333] The memory 1204 can include read-only memory and random access memory, and provide instructions and data to the processor 1203. A portion of the memory 1204 can also include a NVRAM. The memory 1204 stores operating systems and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic services and processing hardware-based tasks.
[0334] The processor 1203 controls the operation of the second terminal device, and the processor 1203 can also be referred to as a CPU. In a specific application, various components of the second terminal device are coupled together through a bus system, which can include a data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, various buses are referred to as a bus system in the figure.
[0335] The method disclosed in the above embodiments of the present application can be applied in the processor 1203 or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with a processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1203 or an instruction in the form of software. The processor 1203 described above can be a general processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and combines the hardware to complete the steps of the above method.
[0336] In the embodiments of the present application, the processor 1203 is configured to execute the method performed by the second terminal device shown in Fig. 7 of the preceding embodiments.
[0337] In another possible design, when the first terminal device or the second terminal device is a chip in the terminal, the chip includes a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute computer-executed instructions stored in a storage unit, so that the chip in the terminal performs the method in any one of the first aspect. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the terminal, such as a read-only memory (ROM) or another type of static storage device that can store static information and instructions, a random access memory (RAM), etc.
[0338] The processor mentioned in any one of the above can be a general central processor, a microprocessor, an ASIC, or one or more integrated circuits for controlling execution of the programs of the method in the first aspect or the second aspect.
[0339] In addition, it should be noted that the above-described apparatus embodiments are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, in the apparatus embodiment provided in the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0340] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits, or special circuits, etc. However, for the present application, software program implementation is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the embodiments of the present application.
[0341] In the embodiments described above, the entire or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, the entire or part of the embodiments can be implemented in the form of a computer program product.
[0342] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the entire or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
Claims
1. A method of negotiation of an audio codec, characterized in that, The method is applied to a first terminal device, and the method comprises: obtaining audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; determining audio encoding / decoding configuration information of the first terminal device according to a mode control parameter of the first terminal device, the mode control parameter of the first terminal device at least comprising a user description parameter generated by the first terminal device; determining a coding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device, the coding / decoding negotiation result comprising a first coding / decoding parameter of the first terminal device for encoding / decoding an audio signal to be processed, the first coding / decoding parameter being adapted to the audio encoding / decoding capability of the second terminal device.
2. The method of claim 1, wherein, The mode control parameter of the first terminal device is used to indicate configuration information of a terminal device running mode used by the first terminal device when performing audio encoding / decoding.
3. The method according to claim 1 or 2, characterized in that, The determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises: determining the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device and an audio encoding / decoding capability of the first terminal device.
4. The method of claim 1, wherein, The mode control parameter of the first terminal device comprises at least one of a battery endurance mode of the first terminal device, an audio call user experience of the first terminal device, and a power distribution mode of the first terminal device.
5. The method of claim 1, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises: when the mode control parameter of the first terminal device satisfies a preset first software configuration condition, determining first audio encoding / decoding configuration information according to the mode control parameter of the first terminal device; when the mode control parameter of the first terminal device does not satisfy the preset first software configuration condition, determining second audio encoding / decoding configuration information according to the mode control parameter of the first terminal device; wherein an audio encoding / decoding capability corresponding to the first audio encoding / decoding configuration information is higher than an audio encoding / decoding capability corresponding to the second audio encoding / decoding configuration information.
6. The method of claim 1, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises: The mode control parameter of the first terminal device comprises a first battery power collected by the first terminal device in a user-permitted case, and the audio encoding / decoding configuration information of the first terminal device is updated according to the first battery power collected by the first terminal device in the user-permitted case to obtain updated audio encoding / decoding configuration information of the first terminal device.
7. The method of claim 6, wherein, The updating of the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device in the user-permitted case comprises: determining a first power interval to which the first battery power belongs; determining a first audio codec mode corresponding to the first power interval according to a correspondence between power intervals and audio codec modes of the first terminal device; adjusting audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the first audio codec mode.
8. The method of claim 6, wherein, The updating of the audio encoding / decoding configuration information of the first terminal device according to the first battery power collected by the first terminal device with user permission comprises: When the first battery power is lower than a preset first power threshold and the first terminal device is in a battery charging mode, maintaining or increasing the audio encoding / decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device.
9. The method of claim 1, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the mode control parameter of the first terminal device comprises: The mode control parameter of the first terminal device comprises a user habit parameter collected by the first terminal device with user permission, and the audio encoding / decoding configuration information of the first terminal device is determined according to the user habit parameter collected by the first terminal device with user permission.
10. The method of claim 9, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the user habit parameter collected by the first terminal device with user permission comprises: determining a first estimated standby duration of the first terminal device according to user schedule information of the first terminal device; determining a second audio codec mode corresponding to the first estimated standby duration according to a correspondence between standby durations and audio codec modes of the first terminal device; determining the audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the second audio codec mode.
11. The method of claim 9, wherein, The determining of the audio encoding / decoding configuration information of the first terminal device according to the user habit parameter collected by the first terminal device with user permission comprises: determining terminal capability allocation information according to a user operation habit information of the first terminal device; determining a third audio codec mode according to a correspondence between terminal capability allocation and audio codec modes of the first terminal device; determining the audio encoding / decoding configuration information of the first terminal device according to an audio encoding / decoding capability corresponding to the third audio codec mode.
12. The method according to any one of claims 1 to 11, characterized in that, The method further comprises: obtaining a network packet loss rate of the first terminal device in an audio call; updating the audio encoding / decoding configuration information of the first terminal device when the network packet loss rate is higher than a first network signal threshold.
13. The method according to any one of claims 1 to 12, characterized in that, The determining of the encoding / decoding negotiation result of the first terminal device according to the audio encoding / decoding configuration information of the first terminal device and the audio encoding / decoding configuration information of the second terminal device comprises: determining an audio encoding / decoding negotiation order according to an audio encoding / decoding negotiation priority strategy of the first terminal device, the audio encoding / decoding negotiation order comprising: performing audio encoding negotiation before audio decoding negotiation, or performing audio decoding negotiation before audio encoding negotiation; The audio encoding negotiation includes determining an encoding negotiation result of the first terminal device according to an audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and an audio decoding capability of the second terminal device. The audio decoding negotiation includes determining a decoding negotiation result of the first terminal device according to an audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and an audio encoding capability of the second terminal device.
14. The method of claim 13, wherein, The determining the encoding negotiation result of the first terminal device according to the audio encoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio decoding capability of the second terminal device includes: determining the encoding negotiation result of the first terminal device to include a first audio mode according to the first terminal device supporting a highest encoding mode being the first audio mode and the second terminal device supporting a decoding mode including the first audio mode; determining a first bit rate intersection between the first audio mode supported by the first terminal device and the first audio mode supported by the second terminal device; determining the encoding negotiation result of the first terminal device to include a maximum bit rate in the first bit rate intersection.
15. The method of claim 13, wherein, The determining the decoding negotiation result of the first terminal device according to the audio decoding capability indicated by the audio encoding / decoding configuration information of the first terminal device and the audio encoding capability of the second terminal device includes: determining the decoding negotiation result of the first terminal device to include a second audio mode according to the first terminal device supporting a highest decoding mode being the second audio mode and the second terminal device supporting an encoding mode including the second audio mode; determining a second bit rate intersection between the second audio mode supported by the first terminal device and the second audio mode supported by the second terminal device; determining the decoding negotiation result of the first terminal device to include a maximum bit rate in the second bit rate intersection.
16. The method of claim 13, wherein, A sum of a bit rate corresponding to the highest encoding mode supported by the first terminal device and a bit rate corresponding to an audio decoding mode performed by the first terminal device is less than a maximum bit rate supported by the audio encoding / decoding capability of the first terminal device.
17. The method of any one of claims 1 to 16, wherein, The method further includes: sending an audio negotiation request to the second terminal device.
18. The method of claim 17, wherein, The audio negotiation request includes the audio encoding / decoding configuration information of the first terminal device.
19. A terminal device, comprising: The terminal device is specifically a first terminal device, and the method includes: an obtaining module, configured to obtain audio encoding / decoding configuration information of a second terminal device, the audio encoding / decoding configuration information of the second terminal device being used to indicate an audio encoding / decoding capability of the second terminal device; a determining module, configured to determine the audio encoding / decoding configuration information of the first terminal device according to a mode control parameter of the first terminal device, the mode control parameter of the first terminal device at least including a user description parameter generated by the first terminal device. The negotiation module is configured to determine the coding / decoding negotiation result of the first terminal device according to the audio coding / decoding configuration information of the first terminal device and the audio coding / decoding configuration information of the second terminal device, wherein the coding / decoding negotiation result comprises first coding / decoding parameters of the first terminal device for coding / decoding audio signals to be processed, and the first coding / decoding parameters are adapted to the audio coding / decoding capability of the second terminal device. The sending module is configured to send the coding / decoding negotiation result of the first terminal device to the second terminal device.
20. A terminal device, comprising: The terminal device comprises at least one processor, which is coupled with a memory and reads and executes instructions in the memory to implement the method in any one of claims 1 to 18.
21. The terminal device of claim 20, wherein, The terminal device further comprises the memory.
22. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to carry out the method of any one of claims 1 to 18.
23. A computer readable storage medium comprising the coding / decoding negotiation result generated by the method of any one of claims 1 to 18.