Audio processing method and terminal device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请提供了一种音频处理方法及装置,目的在于解决如何提高宽窄带互通系统中宽带终端的语音通话音质的问题
[0036] The audio processing method and terminal device provided in the embodiments of this application encode the audio to be transmitted based on the fact that the called object includes both broadband terminals and narrowband terminals, respectively, using a first encoding method and a second encoding method. The first encoding method is the encoding method selected by the server for narrowband terminals that only support narrowband standards within the system, according to a preset voice encoding selection strategy. The second encoding method is the encoding method selected by the server for broadband terminals that support broadband standards, according to a preset voice encoding selection strategy. It can be seen that even if the called object includes narrowband terminals, the broadband terminals among the called objects can still encode the audio to be transmitted based on the encoding method under the broadband standard, thereby ensuring that the call between broadband terminals has better sound quality.
Smart Images

Figure CN117496985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to an audio processing method and terminal device. Background Technology
[0002] A broadband-narrowband interoperability system can be understood as a communication system that includes both broadband and narrowband terminals. Broadband terminals refer to terminals that communicate using public networks such as 3G, 4G, and Wi-Fi. Narrowband terminals refer to terminals that communicate using dedicated networks such as Police Digital Trunking (PDT), Digital Mobile Radio (DMR), and Terrestrial Trunked Radio (TETRA).
[0003] When public and private networks are integrated, both broadband and narrowband terminals can communicate with terminals within the same terminal group. For an individual terminal, the voice quality of a broadband terminal is generally better than that of a narrowband terminal. However, if the terminal group to which the broadband terminal belongs also includes narrowband terminals, then in the case of group calls, even if the broadband terminal communicates with other broadband terminals in the same group, it cannot obtain the superior voice quality that the broadband terminal possesses. Summary of the Invention
[0004] This application provides an audio processing method and apparatus, which aims to solve the problem of how to improve the voice call quality of broadband terminals in a broadband-narrowband interconnection system.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] An audio processing method, comprising:
[0007] A feedback message for receiving a call message, the feedback message including the type information of the called party indicated by the call message, the type information indicating: hybrid, broadband only, or narrowband only, broadband only means that all called parties are broadband terminals that support broadband standards, narrowband only means that all called parties are narrowband terminals that support narrowband standards, and hybrid means that the called parties include both broadband terminals that support broadband standards and narrowband terminals that support narrowband standards;
[0008] When the type information is mixed, the audio to be transmitted is encoded based on the first encoding method and the second encoding method respectively. The first encoding method is the encoding method selected by the server for narrowband terminals that only support narrowband standards in the system according to the preset voice encoding selection strategy. The second encoding method is the encoding method selected by the server for broadband terminals that support broadband standards according to the preset voice encoding selection strategy.
[0009] Optionally, the preset speech coding selection strategy includes:
[0010] Based on the communication standards supported by the terminals, the terminals in the system are divided into broadband groups and narrowband groups. An encoding method is selected based on the intersection of the encoding methods supported by each terminal member in each group and / or a preset priority.
[0011] Optionally, the preset encoding selection strategy includes:
[0012] The encoding methods of all terminals in the system are assigned to support terminal groups in descending order of preset priority. Each terminal can only be assigned an encoding method once, until all terminals have been assigned an encoding method.
[0013] Optionally, the encoding of the audio to be transmitted based on the first encoding method and the second encoding method respectively includes:
[0014] The audio to be transmitted is encoded based on the first encoding method to obtain a first audio data packet. A first value is added to the first audio data packet as an encoding type field. The first value indicates that the first audio data packet uses an encoding and decoding method supported by the narrowband terminal.
[0015] The audio to be transmitted is encoded based on the second encoding method to obtain a second audio data packet. A second value is added to the second audio data packet as an encoding type field. The second value indicates that the second audio data packet uses an encoding and decoding method supported by the broadband terminal.
[0016] Optional, also includes:
[0017] When the group type information is narrowband only, the audio to be transmitted is encoded based on the first encoding method to obtain an audio data packet, and a first value as the encoding type field is added to the audio data packet. The first value indicates that the audio data packet uses an encoding and decoding method supported by the narrowband terminal.
[0018] Alternatively, when the group type information is broadband only, the audio to be transmitted is encoded based on the second encoding method to obtain an audio data packet, and a second value is added to the audio data packet as an encoding type field. The second value indicates that the audio data packet uses an encoding and decoding method supported by the broadband terminal.
[0019] Optional, also includes:
[0020] Receive audio data packets and distribute the audio data packets based on the encoding type field in the audio data packets;
[0021] When the called party has a narrowband terminal, the audio data packet with the encoding type field set to the first value in the audio data packet will have the encoding type field deleted and will be sent to the narrowband communication system.
[0022] When there is a broadband terminal in the called party, the audio data packet with the encoding type field set to the second value in the audio data packet is sent to the broadband terminal in the called party.
[0023] Optional, also includes:
[0024] The system receives audio data packets forwarded by the server and selects the decoding method and audio parameters corresponding to the encoding method of the audio data packets from a set of preset decoding methods and audio parameters for decoding. The set of audio parameters corresponds one-to-one with the decoding method and each is different.
[0025] Optional, also includes:
[0026] In response to the fact that the sending end of the audio data packet is a narrowband terminal and the receiving end is a broadband terminal, the audio data packet is sent to the receiving end after adding a first value as the type field.
[0027] An audio processing method, the method comprising:
[0028] The system receives encoded first audio information sent by the calling terminal. Based on the group type information of the called party, it determines whether the first audio information needs to be decoded and re-encoded using a different encoding method. The group type information indicates: non-mixed or mixed. Non-mixed means that the called parties are all broadband terminals that support broadband standards, or all terminal groups that support only narrowband standards. Mixed means that the called parties include both broadband terminals that support broadband standards and terminal groups that support only narrowband standards.
[0029] When the group type information is mixed, if the first audio information uses the first encoding method, then the first encoding method is used to decode the audio information, and then the second encoding method is used to encode the decoded audio information to obtain the second audio information.
[0030] If the audio information uses the second encoding method, then the audio information is decoded using the second encoding method, and then the decoded audio information is encoded using the first encoding method to obtain the second audio information;
[0031] The first encoding method is an encoding method selected by the server for narrowband terminals that only support narrowband standards within the system, based on a preset voice encoding selection strategy. The second encoding method is an encoding method selected by the server for broadband terminals that support broadband standards, based on a preset voice encoding selection strategy.
[0032] Based on the encoding methods of the first and second audio information, the data is sent to the terminal of the called party that supports the corresponding encoding method.
[0033] A terminal device, comprising:
[0034] Memory, used to store computer programs;
[0035] A processor for running the computer program to implement the aforementioned audio processing method.
[0036] The audio processing method and terminal device provided in the embodiments of this application encode the audio to be transmitted based on the fact that the called object includes both broadband terminals and narrowband terminals, respectively, using a first encoding method and a second encoding method. The first encoding method is the encoding method selected by the server for narrowband terminals that only support narrowband standards within the system, according to a preset voice encoding selection strategy. The second encoding method is the encoding method selected by the server for broadband terminals that support broadband standards, according to a preset voice encoding selection strategy. It can be seen that even if the called object includes narrowband terminals, the broadband terminals among the called objects can still encode the audio to be transmitted based on the encoding method under the broadband standard, thereby ensuring that the call between broadband terminals has better sound quality. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1This is a structural example diagram of a terminal group in a broadband-narrowband interconnection system;
[0039] Figure 2 A flowchart illustrating the audio processing method for a scenario where a broadband terminal in a hybrid terminal group initiates a group call;
[0040] Figure 3 A flowchart illustrating the audio processing method for a scenario where a narrowband terminal in a mixed terminal group initiates a group call;
[0041] Figure 4 A flowchart illustrating the audio processing method for a broadband terminal calling a narrowband terminal in a mixed terminal group.
[0042] Figure 5 A flowchart illustrating the audio processing method for a narrowband terminal calling a broadband terminal in a mixed terminal group.
[0043] Figure 6 This is a flowchart of an audio processing method disclosed in an embodiment of this application;
[0044] Figure 7 This is a flowchart illustrating yet another audio processing method disclosed in an embodiment of this application;
[0045] Figure 8 This is a schematic diagram illustrating how voice data is obtained based on encoding / decoding method information and other information, as disclosed in an embodiment of this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more; "and / or" describes the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0047] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0048] The "multiple" mentioned in the embodiments of this application refers to two or more. It should be noted that in the description of the embodiments of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.
[0049] Figure 1 This is an example of the structure of a terminal group in a broadband-narrowband interoperability system. Figure 1 As can be seen from this, there are three types of terminal groups: consisting of broadband terminal 1, consisting of narrowband terminal 2, or consisting of both broadband terminal 1 and narrowband terminal 2.
[0050] Terminals in the same terminal group can communicate with each other.
[0051] The encoding and decoding methods supported by the broadband terminal 1 include, but are not limited to, Adaptive Multi-rate–Wideband (AMRWB) and Adaptive Multi-rate–Narrowband (AMRNB). Some broadband terminals can also simultaneously support narrowband network standards, i.e., dual-mode terminals, and can also support narrowband encoding and decoding methods, such as Narrowband Voice Coding (NVOC) and Advanced Multi-Band Excitation (AMBE). In this application, dual-mode terminals are generally referred to as broadband terminals. When necessary, broadband terminals that cannot support narrowband standards are referred to as broadband-only terminals.
[0052] Narrowband terminals typically support only one encoding / decoding method. Different narrowband terminals may support different encoding methods, and the sound quality obtained by this encoding method is worse than that obtained by broadband terminals. Encoding / decoding methods supported by narrowband terminals include AMBE, NVOC, and Sinusoidal excitation linear prediction (SELP).
[0053] Terminals within the same terminal group typically negotiate and adopt a codec method supported by all terminals in the group to encode and decode voice calls. It's understandable that for a terminal group consisting of both broadband and narrowband terminals, AMBE (Ampere Beam Encoding and Decoding) supported by the narrowband terminals is used. Conversely, for a terminal group consisting of broadband terminals, AMRWB or AMRNB (Ampere RNB Encoding and Decoding) is used.
[0054] The audio quality of speech streams generated by AMRWB or AMRNB is significantly better than that of speech streams generated by AMBE.
[0055] Therefore, for a terminal group consisting of broadband and narrowband terminals, because the encoding and decoding must be based on AMBE supported by the narrowband terminals, even if the broadband terminals in the terminal group support other encoding and decoding methods, AMBE must still be used for encoding and decoding in order to be compatible with the encoding method of the narrowband terminals. Therefore, even if the broadband terminals communicate with the broadband terminals in the terminal group, they cannot obtain better call quality.
[0056] To address the aforementioned problems, embodiments of this application disclose an audio processing method and apparatus to resolve voice quality issues during calls.
[0057] The audio processing methods will be explained below from the perspectives of broadband terminals, narrowband terminals, group calls, and individual calls. For ease of explanation, terminals consisting of broadband and narrowband terminals will be grouped into hybrid terminal groups. Broadband terminals other than those initiating the call will be referred to as remaining broadband terminals. Narrowband terminals other than those initiating the call will be referred to as remaining narrowband terminals.
[0058] Figure 2 This is an audio processing method for scenarios where broadband terminals in a mixed terminal group initiate group calls. In this scenario, the broadband terminal initiating the group call is referred to as the broadband terminal, and the other broadband terminals besides the broadband terminal initiating the group call are referred to as the remaining broadband terminals.
[0059] Figure 2 The process includes the following steps:
[0060] S101, the broadband terminal, other broadband terminals, and narrowband terminals send encoding and decoding capability information to the TMF module.
[0061] The codec capability information sent by any terminal is information about the codec methods it supports, intended to inform the network side of these supported methods. One example of codec capability information is a set of codec capability information (codec set). In this embodiment, the broadband terminal is a dual-mode terminal; other broadband terminals may include dual-mode terminals or terminals that only support broadband standards.
[0062] S102, the TMF module groups the terminals based on the encoding and decoding capability information and sends the corresponding encoding and decoding method information for each group to each terminal.
[0063] To distinguish it from the aforementioned terminal groups, the packets obtained based on encoding / decoding capability information are referred to here as protocol groups. The encoding / decoding method information corresponding to the protocol group indicates the encoding / decoding method used by the protocol group.
[0064] In some implementations, the protocol group is obtained based on the codec capability information as follows: according to the supported codec methods, terminals that only support AMBE (i.e., narrowband terminals) are divided into one group, called the narrowband group, and terminals that support multiple codec methods (i.e., broadband terminals) are divided into another group, called the broadband group.
[0065] In this scenario, for any protocol group, the codecs supported by that protocol group are negotiated based on the codecs supported by the terminals within that protocol group. In some implementations, the codecs negotiated by any protocol group are the intersection of all codecs supported by terminals in that protocol group. In other implementations, at least one codec is selected from the intersection based on a pre-configured priority. For example, the codec negotiated by the narrowband group might be AMBE. Another example is the intersection of codecs negotiated by the broadband group, which might be AMRWB, AMRNB, and AMBE; AMRWB, with the highest priority, is selected as the codec for the broadband group.
[0066] In other implementations, all codecs are obtained and sorted by priority to form a sequence. Terminals are then selected for each codec in the sequence in descending order of priority, resulting in the corresponding protocol group for each codec. It's understood that a terminal can only be selected if it supports a particular codec. A terminal may be selected by multiple codecs; in this case, the terminal belongs to the codec with the higher priority and is removed from the groups corresponding to the other codecs.
[0067] Understandably, this method can ultimately yield both narrowband and wideband groups. The encoding methods for these narrowband and wideband groups can also be obtained through the aforementioned negotiation, which will not be elaborated upon here.
[0068] It is understandable that S101-S102 can be regarded as preprocessing steps, which can be executed only once before the following steps, without needing to be executed repeatedly.
[0069] S103, The broadband terminal sends a call message to the TCF module.
[0070] The call message includes group call type information and terminal group identifier. The group call type information indicates that the initiated call type is a group call. The terminal group identifier indicates the terminal group to which the broadband terminal belongs, i.e., the terminal group of the group call.
[0071] The S104 and TCF modules send feedback messages to the broadband terminal.
[0072] The feedback message includes: the type information of the called object. In this step, the called object is a terminal group.
[0073] The type of terminal group is called the group type, which includes hybrid, broadband only, or narrowband only. Group type information indicates the type of terminal group. Broadband only means all terminals in the group are broadband terminals; narrowband only means all terminals are narrowband terminals; and hybrid means the group includes both broadband and narrowband terminals.
[0074] Because this embodiment assumes that the broadband terminal belongs to a mixed terminal group, the group type information indicates that it is mixed.
[0075] S105. The broadband terminal encodes the voice based on the feedback message to obtain a voice data packet.
[0076] Because the feedback message includes mixed type information. The broadband terminal starts two encoders. The first encoder encodes the speech based on the narrowband group's encoding method, resulting in a first speech data packet. The first speech data packet includes a first value as a type field, which indicates the narrowband (e.g., the encoding / decoding method used by the narrowband terminal). In some implementations, the first value is added to the header of the first speech data packet.
[0077] The second encoder encodes the speech based on the wideband encoding method to obtain a second speech data packet. The second speech data packet includes a second value as a type field, which indicates the wideband (such as the encoding / decoding method for wideband terminals). In some implementations, the second value is added to the header of the second speech data packet.
[0078] It is understandable that the encoding / decoding methods for narrowband and wideband groups are obtained through S102.
[0079] The first and second voice data packets are collectively referred to as voice data packets.
[0080] S106. The broadband terminal sends voice data packets to the TMF module.
[0081] S107 and TMF modules respond to the fact that the source of the voice data packet is a broadband terminal by recognizing the type field of the voice data packet and forwarding the voice data packet.
[0082] Specifically, the type field of the first voice data packet is set to the first value. After deleting the first value from the first voice data packet, the first voice data packet is sent to the narrowband system. The type field of the second voice data packet is set to the second value, and the second voice data packet is sent to the other broadband terminals.
[0083] In some implementations, the first and second voice data packets are sent using a cross-transmission method.
[0084] S108. The narrowband system parses the first voice data packet, obtains the voice data, and sends the voice data to the narrowband terminals in the mixed terminal group.
[0085] S109. After receiving voice data, the narrowband terminal decodes the voice data based on the encoding and decoding methods supported by the narrowband terminal.
[0086] S110 and the other broadband terminals, after receiving the second voice data packet, decode the second bitstream based on the encoding and decoding method of the broadband group.
[0087] Understandably, the remaining broadband terminals obtain the second value from the second bitstream. This second value represents the encoding / decoding method of the broadband group. Therefore, the second bitstream is decoded based on the encoding / decoding method of the broadband group. The encoding / decoding method of the broadband group is obtained based on S102.
[0088] Figure 3 This is an audio processing method for scenarios where narrowband terminals in a mixed terminal group initiate group calls. In this scenario, the narrowband terminal initiating the group call is referred to as the narrowband terminal, and the other narrowband terminals besides the narrowband terminal initiating the group call are referred to as the remaining narrowband terminals.
[0089] Figure 3 The process includes the following steps:
[0090] S201, broadband terminals, other broadband terminals, and narrowband terminals send encoding / decoding capability information to the TMF module.
[0091] For details, please refer to S101.
[0092] The S202 and TMF modules group the terminals based on the encoding and decoding capability information and send the corresponding encoding and decoding capability information for each group to each terminal.
[0093] For details, please refer to S102.
[0094] S203, The narrowband terminal sends a call message to the TCF module.
[0095] The call message includes group call type information and the identifier of the terminal group. See S103 for details.
[0096] The S204 and TCF modules send feedback messages to the narrowband terminal.
[0097] The feedback message includes: terminal group type information (group type information). Because this embodiment assumes that the broadband terminal belongs to a mixed terminal group, the group type information indicates that it is mixed.
[0098] S205: The narrowband terminal encodes the voice data using its own supported encoding and decoding methods based on the feedback message, and then sends the voice data to the narrowband system.
[0099] S206. The narrowband system generates voice data packets based on voice data and sends the voice data packets to the TMF module.
[0100] S207 and TMF modules, in response to data sources originating from narrowband systems, add a first value as the type field to the voice data packet and then send the voice data packet to the broadband terminal.
[0101] In one embodiment, the TMF module also transmits voice data packets back to the narrowband system, which then forwards them to other narrowband terminals, including the following steps:
[0102] S208. The narrowband system parses voice data packets, obtains voice data, and sends voice data to the remaining narrowband terminals in the terminal group.
[0103] S209, and other narrowband terminals decode voice data based on their own supported encoding and decoding methods.
[0104] In another embodiment, when the narrowband system receives a call and voice information from a narrowband terminal, it will directly send it to other narrowband terminals without requiring the TMF module to transmit it back.
[0105] S210. After receiving the voice data packet, the broadband terminal decodes the voice data packet using the narrowband encoding and decoding method based on the first value.
[0106] Figure 4 This is an audio processing method for a scenario where a broadband terminal in a mixed terminal group calls a narrowband terminal. In this scenario, the broadband terminal that initiates the group call is called the broadband terminal, and the other broadband terminals besides the broadband terminal that initiates the group call are called the remaining broadband terminals.
[0107] Figure 4 The process includes the following steps:
[0108] S301, broadband terminals, other broadband terminals, and narrowband terminals send encoding / decoding capability information to the TMF module.
[0109] For details, please refer to S101.
[0110] The S302 and TMF modules group the terminals based on the encoding and decoding capability information and send the corresponding encoding and decoding capability information for each group to each terminal.
[0111] For details, please refer to S102.
[0112] S303, The broadband terminal sends a call message to the TCF module.
[0113] The call message includes call type information and a terminal identifier. The call type information indicates that the initiated call is a call. The terminal identifier indicates the called terminal belonging to the same hybrid terminal group as the broadband terminal. In the application scenario of this embodiment, the terminal identifier indicates the identifier of a narrowband terminal belonging to the same hybrid terminal group as the broadband terminal.
[0114] The S304 and TCF modules send feedback messages to the broadband terminal.
[0115] The feedback message includes: the type information of the called object. In this step, the called object is a terminal.
[0116] Terminal types include broadband and narrowband. Terminal type information is used to indicate the type of terminal. If the terminal type information indicates broadband, it means that the called terminal is a broadband terminal; if the terminal type information indicates narrowband, it means that the called terminal is a narrowband terminal.
[0117] Because this embodiment assumes that the broadband terminal initiates a single call to a narrowband terminal in the same terminal group, the terminal type information indicates narrowband.
[0118] S305. The broadband terminal encodes the voice based on the feedback message to obtain a voice data packet.
[0119] It is understood that, in the application scenario of this embodiment, the type information included in the feedback message indicates that the called terminal is a narrowband terminal. The broadband terminal encodes the voice based on the narrowband group's encoding method to obtain a voice data packet. The voice data packet includes a first value as a type field, which indicates the encoding / decoding method sent to the narrowband terminal and the narrowband group.
[0120] S306. The broadband terminal sends voice data packets to the TMF module.
[0121] The S307 and TMF modules respond that the source of the voice data packet is a broadband terminal. Based on the first value in the type field of the voice data packet, they delete the first value in the voice data packet and then send the voice data packet to the narrowband system.
[0122] S308: The narrowband system parses the voice data packets, obtains the voice data, and sends the voice data to the narrowband terminal represented by the terminal's identifier.
[0123] S309. After receiving voice data, the narrowband terminal decodes the voice data based on the encoding and decoding methods supported by the narrowband terminal.
[0124] Figure 5 This is an audio processing method for a scenario where a narrowband terminal in a mixed terminal group makes a single call to a broadband terminal. In this scenario, the narrowband terminal that initiates the single call is called the narrowband terminal, and the other narrowband terminals besides the narrowband terminal that initiates the single call are called the remaining narrowband terminals.
[0125] Figure 5 The process includes the following steps:
[0126] S401, broadband terminals, other broadband terminals, and narrowband terminals send encoding / decoding capability information to the TMF module.
[0127] For details, please refer to S101.
[0128] The S402 and TMF modules group the terminals based on their encoding and decoding capability information and send the corresponding encoding and decoding capability information for each group to each terminal.
[0129] For details, please refer to S102.
[0130] S403, the narrowband terminal sends a call message to the TCF module.
[0131] The call message includes call type information and a terminal identifier. The call type information indicates that the initiated call is a call. The terminal identifier indicates the called terminal belonging to the same hybrid terminal group as the narrowband terminal. In the application scenario of this embodiment, the terminal identifier indicates the identifier of a broadband terminal belonging to the same hybrid terminal group as the narrowband terminal.
[0132] The S404 and TCF modules send feedback messages to the narrowband terminal.
[0133] Based on S404, the feedback message includes: the type information of the called terminal. In this step, the type information of the called object indicates broadband.
[0134] S405: The narrowband terminal encodes the voice data using its own supported encoding and decoding methods based on the feedback message, and then sends the voice data to the narrowband system.
[0135] S406 The narrowband system generates voice data packets based on voice data and sends the voice data packets to the TMF module.
[0136] The S407 and TMF modules, in response to a narrowband system as the data source and a broadband terminal as the receiving end, add the first value as the type field to the voice data packet and then send the voice data packet to the broadband terminal.
[0137] S408. After receiving the voice data packet, the broadband terminal decodes the voice data packet using the narrowband encoding and decoding method based on the first value.
[0138] The embodiments of this application introduce a new codec negotiation mechanism, employing different encoding methods for different called parties. This ensures that the audio sent to the broadband terminal uses a codec with better sound quality, while the audio sent to the narrowband terminal uses a codec supported by the narrowband terminal. Theoretically, the final result is that in a broadband-narrowband interoperability system, the broadband terminal receives audio with superior sound quality without needing to reduce sound quality for compatibility with narrowband terminals.
[0139] The above embodiments can be summarized as follows: Figure 6 The audio processing method shown includes the following steps:
[0140] S51, Receive the feedback message for the call message.
[0141] The feedback message includes the type information of the called party indicated by the call message. The type information indicates: hybrid, broadband only, or narrowband only. Broadband only means that all called parties are broadband terminals that support broadband standards. Narrowband only means that all called parties are narrowband terminals that support narrowband standards. Hybrid means that the called parties include a terminal group consisting of both broadband terminals that support broadband standards and narrowband terminals that support narrowband standards.
[0142] S52. When the type information is mixed, the audio to be transmitted is encoded based on the first encoding method and the second encoding method respectively.
[0143] The first encoding method is an encoding method selected by the server for narrowband terminals that only support narrowband standards within the system, based on a preset voice encoding selection strategy. The second encoding method is an encoding method selected by the server for broadband terminals that support broadband standards, based on a preset voice encoding selection strategy.
[0144] In some implementations, the preset voice coding selection strategy includes: dividing the terminals in the system into broadband groups and narrowband groups based on the communication standards supported by the terminals, and selecting a coding method based on the intersection of the coding methods supported by each terminal member in each group and / or a preset priority.
[0145] In other implementations, the encoding methods of all terminals in the system are divided into supported terminal groups according to a preset priority from high to low. Each terminal can only be divided once, until all terminals have been assigned an encoding method.
[0146] The above embodiments can also be summarized as follows: Figure 7 The audio processing method shown is executed by a server on the network side, which includes the TCF module and / or TMF module described in the above embodiments.
[0147] Figure 7 The process includes the following steps:
[0148] S61. Receive the encoded first audio information sent by the calling terminal, and based on the group type information of the called party, determine whether it is necessary to re-encode the first audio information using a different encoding method after decoding.
[0149] The group type information indicates: non-mixed and mixed. Non-mixed means that the called parties are all broadband terminals that support broadband standards, or all narrowband terminals that only support narrowband standards. Mixed means that the called parties include both broadband terminals that support broadband standards and narrowband terminals that only support narrowband standards.
[0150] S62. When the group type information is mixed, if the first audio information uses the first encoding method, then the first encoding method is used to decode the audio information, and then the second encoding method is used to encode the decoded audio information to obtain the second audio information.
[0151] S63. When the group type information is mixed, if the audio information uses the second encoding method, then the audio information is decoded using the second encoding method, and then the decoded audio information is encoded using the first encoding method to obtain the second audio information.
[0152] The first encoding method is an encoding method selected by the server for narrowband terminals that only support narrowband standards within the system, based on a preset voice encoding selection strategy. The second encoding method is an encoding method selected by the server for broadband terminals that support broadband standards, based on a preset voice encoding selection strategy.
[0153] S64. Based on the encoding methods of the first audio information and the second audio information, respectively send them to the terminal of the called party that supports the corresponding encoding method.
[0154] Figure 6 and Figure 7The process shown enables broadband terminals to receive audio with superior sound quality when the called terminal includes both broadband and narrowband terminals.
[0155] like Figure 8 As shown, the terminal's audio recording and playback are implemented by calling audio class API interfaces according to specifications and passing in relevant parameters. There are three ways for the terminal to select the codec format: one is that the application directly sets and passes in the codec format information, such as in a recording APK; another is that the application provides information about the audio file and obtains the codec format based on the internal information of the audio file, such as in a music playback APK; and the third is a communication-related method where the server and terminal negotiate and determine the codec format, such as in public network calls and PoC calls.
[0156] Typically, codec information is used to select the corresponding codec, while other parameters set by the application are used to select the audio device, thereby selecting the corresponding audio parameters. Once the application is determined by the parameters, the device is also determined, and then the sound quality is ensured by adjusting the audio parameters corresponding to the device.
[0157] In the embodiments provided in this application, in order to obtain better sound effects, in addition to encoding and decoding based on the receiving end, different sound effect parameters are configured for different encoding and decoding methods.
[0158] Combination Figure 8 As shown, the call application running in the terminal transmits other information to the audio parameter module through the HAL layer to select the corresponding audio parameters. Different audio parameters are used by different algorithm modules in the DSP. For example, the first audio parameter is used for algorithm module one, and the second audio parameter is used for algorithm module two, etc. In the embodiments of this application, different audio parameters are configured for different encoding and decoding methods. Specifically, for the first encoding and decoding method and the second encoding and decoding method, the audio parameters used for at least one algorithm module are different. That is, the audio parameters used for algorithm module one to algorithm module N by the first encoding and decoding method are A1, A2, ..., AN, and the audio parameters used for algorithm module one to algorithm module N by the second encoding and decoding method are A1, B2, ..., BN.
[0159] If the functions described in the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computing device readable storage medium. Based on this understanding, the parts of the embodiments of this application that contribute to the prior art or the technical solutions can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computing device (which may be a personal computer, server, mobile computing device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0160] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
Claims
1. An audio processing method, characterized in that, include: A feedback message for receiving a call message, the feedback message including the type information of the called party indicated by the call message, the type information indicating: hybrid, broadband only, or narrowband only, broadband only means that all called parties are broadband terminals that support broadband standards, narrowband only means that all called parties are narrowband terminals that support narrowband standards, and hybrid means that the called parties include a terminal group consisting of both broadband terminals that support broadband standards and narrowband terminals that support narrowband standards; When the type information is mixed, the audio to be transmitted is encoded based on the first encoding method and the second encoding method respectively. The first encoding method is the encoding method selected by the server for narrowband terminals that only support narrowband standards in the system according to the preset voice encoding selection strategy. The second encoding method is the encoding method selected by the server for broadband terminals that support broadband standards according to the preset voice encoding selection strategy. When the type information is narrowband only, the audio to be transmitted is encoded based on the first encoding method to obtain an audio data packet, and a first value as the encoding type field is added to the audio data packet. The first value indicates that the audio data packet uses an encoding and decoding method supported by the narrowband terminal. When the type information is broadband only, the audio to be transmitted is encoded based on the second encoding method to obtain an audio data packet, and a second value is added to the audio data packet as an encoding type field. The second value indicates that the audio data packet uses an encoding and decoding method supported by the broadband terminal.
2. The audio processing method according to claim 1, characterized in that, The preset speech coding selection strategy includes: Based on the communication standards supported by the terminals, the terminals in the system are divided into broadband groups and narrowband groups. An encoding method is selected based on the intersection of the encoding methods supported by each terminal member in each group and / or a preset priority.
3. The audio processing method according to claim 1, characterized in that, The preset speech coding selection strategy includes: The encoding methods of all terminals within the system are assigned to support terminal members in descending order of preset priority. Each terminal can only be assigned an encoding method once, until all terminals have been assigned an encoding method.
4. The method according to claim 1, characterized in that, The encoding of the audio to be transmitted based on the first encoding method and the second encoding method respectively includes: The audio to be transmitted is encoded based on the first encoding method to obtain a first audio data packet. A first value is added to the first audio data packet as an encoding type field. The first value indicates that the first audio data packet uses an encoding and decoding method supported by the narrowband terminal. The audio to be transmitted is encoded based on the second encoding method to obtain a second audio data packet. A second value is added to the second audio data packet as an encoding type field. The second value indicates that the second audio data packet uses an encoding and decoding method supported by the broadband terminal.
5. The method according to claim 1 or 4, characterized in that, Also includes: Receive audio data packets and distribute the audio data packets based on the encoding type field in the audio data packets; When the called party has a narrowband terminal, the audio data packet with the encoding type field set to the first value in the audio data packet will have the encoding type field deleted and will be sent to the narrowband communication system. When there is a broadband terminal in the called party, the audio data packet with the encoding type field set to the second value in the audio data packet is sent to the broadband terminal in the called party.
6. The method according to claim 1 or 4, characterized in that, Also includes: The system receives audio data packets forwarded by the server and selects the decoding method and audio parameters corresponding to the encoding method of the audio data packets from a set of preset decoding methods and audio parameters for decoding. The set of audio parameters corresponds one-to-one with the decoding method and each is different.
7. The method according to claim 6, characterized in that, Also includes: In response to the fact that the sending end of the audio data packet is a narrowband terminal and the receiving end is a wideband terminal, the audio data packet is sent to the receiving end after adding a first value as the encoding type field.
8. An audio processing method, characterized in that, The method includes: The system receives encoded first audio information sent by the calling terminal. Based on the group type information of the called party, it determines whether the first audio information needs to be decoded and re-encoded using a different encoding method. The group type information indicates: non-mixed or mixed. Non-mixed means that the called parties are all broadband terminals that support broadband standards, or all terminal groups that support only narrowband standards. Mixed means that the called parties include both broadband terminals that support broadband standards and terminal groups that support only narrowband standards. When the group type information is mixed, if the first audio information uses the first encoding method, then the first encoding method is used to decode the audio information, and then the second encoding method is used to encode the decoded audio information to obtain the second audio information. If the audio information uses the second encoding method, then the audio information is decoded using the second encoding method, and then the decoded audio information is encoded using the first encoding method to obtain the second audio information; the first encoding method is the encoding method selected by the server for narrowband terminals that only support narrowband standards in the system according to a preset voice encoding selection strategy, and the second encoding method is the encoding method selected by the server for broadband terminals that support broadband standards according to a preset voice encoding selection strategy. Based on the encoding methods of the first and second audio information, the data is sent to the terminal of the called party that supports the corresponding encoding method.
9. A terminal device, characterized in that, include: Memory, used to store computer programs; A processor for running the computer program to implement the audio processing method as described in claim 1, 4 or 6.
Citation Information
Patent Citations
Voice communication method and system in broadband and narrowband intercommunication environment
CN112770269A
Voice coding method and device, electronic equipment and storage medium
CN114898760A