Voice processing method, terminal device and storage medium
By encoding the voice signal in a weak signal environment, the signal distortion and unclear call caused by traditional voice encoding methods is solved, and clear voice transmission and decoding at extremely low code rates is achieved, which improves call quality.
Patent Information
- Application Number
- CN202011568861.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-12-25
AI Technical Summary
When the communication environment is poor, traditional voice encoding methods lead to large distortion of the signal at the receiving end, intermittent calls, noise or even silent, and the call effect is very poor.
In a weak signal environment, the single-word phonetic symbols, tone and duration in the voice signal are extracted for encoding, and encoding information of the preset bit length is generated, and decoded at the receiving end to restore the pronunciation, and playback using the preset sound.
Clear voice restoration is achieved in a weak signal environment, improving the call effect, allowing the receiver to judge semantics based on pronunciation, and solving the problem of intermittent call and noise.
Smart Images

Figure CN114694662B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of communication technologies, and in particular to a voice processing method, a terminal device, and a storage medium. Background Art
[0002] When a user uses a terminal device to make a voice call, the terminal device on the voice sender side performs voice encoding on the voice signal, and correspondingly, on the voice receiver side, the terminal device performs voice decoding on the received data and restores it to voice.
[0003] Currently, terminal devices use traditional voice coding methods such as waveform coding. However, when the communication environment is poor, using traditional voice coding methods will cause significant signal distortion at the receiving end, resulting in intermittent calls, noise, or even silence, and poor call quality. Summary of the Invention
[0004] The embodiments of the present application provide a voice processing method, a terminal device, and a storage medium. When the communication environment of the terminal device is poor, the two parties in a call can clearly communicate the pronunciation of words, and the user can judge the semantics through the pronunciation, thereby improving the call quality.
[0005] In a first aspect, a voice processing method is provided, which is applied to a first terminal device, and the first terminal device is in a call state with a second terminal device, the method comprising: obtaining a voice signal input by a user; when it is determined that the current environment is weak and it is determined that both the first terminal device and the second terminal device support phonetic codec, extracting multiple groups of information from the voice signal, each group of information including the phonetic symbol, tone and duration of a single word; phonetic symbol codec refers to encoding and decoding the phonetic symbol, tone and duration; obtaining the coding information corresponding to each group of information according to a coding table; the coding table stores the correspondence between the phonetic symbol, tone and duration and the coding information; and sending the coding information to the second terminal device.
[0006] The voice processing method provided in the first aspect can be applied to the voice sender making a voice call in a weak signal environment. At the voice sender, the phonetic symbol, tone, and duration of each word in the user's voice are extracted and encoded to obtain encoding information of a preset bit length corresponding to each word. Since the encoding information is of a preset bit length, it can be transmitted at an extremely low bit rate in a weak signal environment. By encoding the pronunciation and duration of the words, the receiver can clearly restore the pronunciation of the sender user, so that the receiver user obtains the semantics based on the played pronunciation, thereby improving the call quality.
[0007] In one possible implementation, the coding information corresponding to each group of information is obtained according to the coding table, including: for each group of information, determining whether the commonly used index table includes the phonetic symbols in the group of information; if the commonly used index table includes the phonetic symbols in the group of information, obtaining the coding information according to the commonly used index table; if the commonly used index table does not include the phonetic symbols in the group of information, obtaining the coding information according to the global index table.
[0008] In this implementation, a commonly used index table is generated based on the user's commonly used words. The number of commonly used index values in the commonly used index table is much smaller than the number of global index values in the global index table. The commonly used index table is searched and encoded first, which reduces the amount of search data and improves encoding efficiency.
[0009] In one possible implementation, determining that the current environment is a weak signal environment includes: if it is determined that the target parameter meets the first preset condition, sending a first request message to the second terminal device, the first request message is used to instruct the second terminal device to use phonetic codec, and the target parameter is used to indicate the signal state of the communication environment currently located by the first terminal device; receiving a first response message sent by the second terminal device, the first response message is used to instruct the second terminal device to use phonetic codec.
[0010] In a possible implementation, the first request message includes a hysteresis timer, and the hysteresis timer is used to indicate a delay time for the second terminal device to use phonetic symbol encoding and decoding.
[0011] In this implementation, by setting a hysteresis timer, time is reserved for switching the voice codec mode. After the hysteresis timer times out, the first terminal device and the second terminal device use the phonetic codec at the same time, thereby improving the switching effect of the codec mode.
[0012] In one possible implementation, determining that the current environment is a weak signal includes: receiving a second request message sent by a second terminal device, the second request message is used to instruct the first terminal device to use phonetic codec; sending a second response message to the second terminal device, the second response message is used to instruct the first terminal device to use phonetic codec.
[0013] In one possible implementation, the method also includes: if it is determined that the target parameter meets the second preset condition, sending a third request message to the second terminal device, the third request message is used to instruct the second terminal device to use waveform coding and decoding, and the target parameter is used to indicate the signal status of the communication environment currently located by the first terminal device; receiving a third response message sent by the second terminal device, the third response message is used to instruct the second terminal device to use waveform coding and decoding.
[0014] In one possible implementation, the method also includes: receiving a fourth request message sent by the second terminal device, the fourth request message is used to instruct the first terminal device to use waveform coding and decoding; if it is determined that the target parameter meets the second preset condition, sending a fourth response message to the second terminal device, the fourth response message is used to instruct the first terminal device to use waveform coding and decoding, and the target parameter is used to indicate the signal status of the communication environment currently located by the first terminal device.
[0015] In the above implementation, when the communication environment of the first terminal device and the second terminal device is not a weak signal environment, it is possible to switch to a traditional voice coding and decoding method to improve the call quality.
[0016] In one possible implementation, determining that the first terminal device supports phonetic symbol encoding and decoding includes: if it is determined that the first terminal device does not have the phonetic symbol encoding and decoding function turned on, generating and outputting a prompt message; receiving a first instruction from the user; and turning on the function according to the first instruction.
[0017] In one possible implementation, determining that both the first terminal device and the second terminal device support phonetic symbol codecs includes: sending first capability information to the second terminal device, the first capability information being used to indicate that the first terminal device supports phonetic symbol codecs; receiving first capability response information sent by the second terminal device, the first capability response information being used to indicate that the second terminal device supports phonetic symbol codecs; or, receiving second capability information sent by the second terminal device, the second capability information being used to indicate that the second terminal device supports phonetic symbol codecs; and sending second capability response information to the second terminal device, the second capability response information being used to indicate that the first terminal device supports phonetic symbol codecs.
[0018] In a possible implementation, the method further includes: displaying a settings interface; receiving an operation of a user in the settings interface; and enabling a phonetic symbol encoding and decoding function in response to the operation.
[0019] In the second aspect, a voice processing method is provided, which is applied to a second terminal device, and the second terminal device is in a call state with the first terminal device. The method includes: receiving a first message sent by the first terminal device; decoding the first message according to a coding table to obtain multiple groups of information; each group of information includes the phonetic symbol, tone and duration of a single word, and the coding table stores the correspondence between the phonetic symbol, tone and duration and the coding information; generating a voice signal based on the multiple groups of information; and playing the voice signal using a preset sound.
[0020] The voice processing method provided in the second aspect can be applied to the voice receiver of a voice call in a weak signal environment. After the coded information sent by the voice sender is transmitted through the channel, the received information is decoded at the voice receiver to obtain the phonetic symbol, tone and duration of each word, thereby generating a complete and fluent voice signal, and playing it with a preset sound. Since the coded information is of a preset bit length, it can be transmitted at an extremely low bit rate in a weak signal environment. By encoding the pronunciation and duration of a single word, the receiver can clearly restore the pronunciation of the sender user, so that the receiver user obtains the semantics based on the played pronunciation, thereby improving the call effect.
[0021] In one possible implementation, the first information is decoded according to a coding table to obtain multiple groups of information, including: obtaining multiple encoded information from the first information in sequence, where the length of the encoded information is a preset bit length; for each encoded information, obtaining the phonetic symbol, tone and duration corresponding to the encoded information according to the coding table.
[0022] In one possible implementation, before receiving the first information sent by the first terminal device, it also includes: receiving a first request message sent by the first terminal device, the first request message is used to instruct the second terminal device to use phonetic codec; phonetic codec refers to encoding and decoding phonetic symbols, tones and duration; sending a first response message to the first terminal device, the first response message is used to instruct the second terminal device to use phonetic codec.
[0023] In a possible implementation, the first request message includes a hysteresis timer, and the hysteresis timer is used to indicate a delay time for the second terminal device to use phonetic symbol encoding and decoding.
[0024] In one possible implementation, before receiving the first information sent by the first terminal device, it also includes: if it is determined that the target parameter meets the first preset condition, sending a second request message to the first terminal device, the second request message is used to instruct the first terminal device to use phonetic codec, phonetic codec refers to encoding and decoding phonetic symbols, tones and duration, and the target parameters are used to indicate the signal status of the communication environment currently located by the second terminal device; receiving a second response message sent by the first terminal device, the second response message is used to instruct the first terminal device to use phonetic codec.
[0025] In one possible implementation, the method also includes: receiving a third request message sent by the first terminal device, the third request message is used to instruct the second terminal device to use waveform coding and decoding; if it is determined that the target parameter meets the second preset condition, sending a third response message to the first terminal device, the third response message is used to instruct the second terminal device to use waveform coding and decoding, and the target parameter is used to indicate the signal status of the communication environment currently located by the first terminal device.
[0026] In one possible implementation, the method also includes: if it is determined that the target parameter meets the second preset condition, sending a fourth request message to the first terminal device, the fourth request message is used to instruct the first terminal device to use waveform coding and decoding; receiving a fourth response message sent by the first terminal device, the fourth response message is used to instruct the first terminal device to use waveform coding and decoding.
[0027] In one possible implementation, before receiving the first information sent by the first terminal device, it also includes: receiving first capability information sent by the first terminal device, the first capability information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding; sending first capability response information to the first terminal device, the first capability response information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding; phonetic symbol encoding and decoding refers to encoding and decoding phonetic symbols, tones and duration; or, sending second capability information to the first terminal device, the second capability information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding; receiving second capability response information sent by the first terminal device, the second capability response information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding.
[0028] In one possible implementation, the method further includes: displaying a settings interface; receiving a user operation in the settings interface; and in response to the operation, turning on a phonetic symbol encoding and decoding function, where phonetic symbol encoding and decoding refers to encoding and decoding phonetic symbols, tones, and durations.
[0029] In a third aspect, an apparatus is provided, comprising: a unit or means for executing each step in any of the above aspects.
[0030] In a fourth aspect, a terminal device is provided, comprising a processor, a memory and a transceiver, wherein the transceiver is used to communicate with other devices, and the processor is used to call a program stored in the memory to execute the method provided in any of the above aspects.
[0031] In a fifth aspect, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are executed on a computer or a processor, the method provided in any of the above aspects is implemented.
[0032] In a sixth aspect, a program product is provided, which includes a computer program, wherein the computer program is stored in a readable storage medium, at least one processor of a device can read the computer program from the readable storage medium, and the at least one processor executes the computer program so that the device implements the method provided in any of the above aspects.
[0033] In any of the above aspects, in a possible implementation manner, the encoded information includes a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration.
[0034] In a possible implementation, the first information component includes a first information sub-component corresponding to the phonetic symbol and a second information sub-component corresponding to the tone.
[0035] In one possible implementation, the coding table includes a global index table and a commonly used index table. The commonly used index table is generated based on the number of times a user uses a single word within a preset time period. The global index table includes phonetic symbols and global index values of the phonetic symbols. The phonetic symbols included in the commonly used index table have commonly used index values and the global index values of the phonetic symbols in the global index table.
[0036] In a possible implementation, the target parameter includes at least one of the following: location information of the terminal device, a cell identifier of a cell currently accessed by the terminal device, signal strength of a signal received by the terminal device, or a voice packet loss rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a diagram of application scenarios applicable to the embodiments of the present application;
[0038] Figure 2 Schematic diagram of the principle of voice call for terminal equipment;
[0039] Figure 3 This is a diagram showing the effect of a call using traditional voice codec in a weak signal environment.
[0040] Figure 4 A schematic diagram of the call effect provided by an embodiment of the present application in a weak signal environment;
[0041] Figure 5 A message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0042] Figure 6 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0043] Figure 7 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0044] Figure 8 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0045] Figure 9 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0046] Figure 10 An interface diagram for setting the phonetic symbol encoding and decoding method provided in an embodiment of the present application;
[0047] Figure 11 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0048] Figure 12 Another message interaction diagram of the voice processing method provided in an embodiment of the present application;
[0049] Figure 13 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0050] Figure 14 Another structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following describes the embodiments of the present application with reference to the accompanying drawings.
[0052] For example, Figure 1 This is an application scenario diagram applicable to the embodiment of this application. Figure 1 As shown, user A uses terminal device 100 and user B uses terminal device 200, and user A and user B can conduct a voice call. The present embodiment does not limit the type of terminal device. For example, some examples of terminal devices may include: mobile phones, tablet computers, PDAs, wearable devices, etc.
[0053] Figure 2 This is a schematic diagram of the principle of a voice call on a terminal device. Figure 1 and Figure 2 As shown, when user A speaks and user B listens, terminal device 100 is the voice sender and terminal device 200 is the voice receiver. Terminal device 100 receives the voice signal input by user A, performs voice encoding on it, and generates coded information. After the coded information is transmitted through the channel, it is received by terminal device 200. Terminal device 200 performs voice decoding on the received information, restores the generated voice signal, and outputs it to user B.
[0054] Need to explain, Figure 2 The voice encoding and decoding part of the voice call process is shown, and other processing processes are not limited.
[0055] The concepts in the embodiments of the present application are explained below.
[0056] 1. Speech Coding
[0057] Speech coding has both broad and narrow meanings. The broad meaning refers to a coding method that includes speech encoding at the sender and speech decoding at the receiver. The narrow meaning refers to speech encoding at the sender. For the sake of distinction, in the embodiments of this application, coding in the broad sense is referred to as coding and decoding.
[0058] The purpose of speech coding and decoding is to digitize speech signals, compress the transmission bandwidth of speech signals, and increase the transmission rate of the channel.
[0059] There are many ways to implement speech codecs, for example, traditional speech codecs such as waveform codecs, feature codecs, and parameter codecs, and also include the phonetic codecs in the embodiments of this application.
[0060] For ease of explanation, the traditional speech coding and decoding method in the embodiment of the present application is described using waveform coding and decoding as an example.
[0061] 2. Weak signal environment
[0062] When the terminal device is in a weak signal environment, the signal quality is poor. If the traditional voice codec method is used, the transmitted voice bit rate will be reduced, and the voice signal restored at the receiving end will be distorted, resulting in intermittent, noisy or even silent conditions.
[0063] It should be noted that in the embodiment of the present application, the terminal device is in a weak signal environment, which means that any one of the two terminal devices in the call is in a weak signal environment. Figure 1 Taking the terminal device 100 as an example, the terminal device 100 is in a weak signal environment, including the following three scenarios: scenario 1, the terminal device 100 is in a weak signal environment; scenario 2, the terminal device 200 talking with the terminal device 100 is in a weak signal environment; scenario 3, both the terminal device 100 and the terminal device 200 are in a weak signal environment.
[0064] Among them, the terminal device can obtain its own target parameters and determine whether the terminal device is in a weak signal environment based on whether the target parameters meet the preset conditions. In an embodiment of the present application, when the target parameters meet the first preset condition, it is determined that the terminal device is in a weak signal environment, and when the target parameters meet the second preset condition, it is determined that the terminal device is not in a weak signal environment. Different target parameters have different corresponding first preset conditions and second preset conditions. Optionally, the target parameters may include at least one of the following: the location information of the terminal device, the cell identifier of the cell currently accessed by the terminal device, the signal strength of the signal received by the terminal device, or the voice packet loss rate.
[0065] Optionally, the target parameter is the location information of the terminal device, and the weak signal geographical range can be pre-recorded. When the location information of the terminal device is within the weak signal geographical range, it is determined that the terminal device is in a weak signal environment. Conversely, when the location information of the terminal device is not within the weak signal geographical range, it is determined that the terminal device is not in a weak signal environment. The embodiments of the present application do not limit the weak signal geographical range, for example, some areas where it is difficult to deploy stations, such as mountainous areas and bridges.
[0066] Optionally, the target parameter is the cell identifier of the cell currently accessed by the terminal device, and the weak signal cell identifier can be pre-recorded. When the cell identifier of the cell currently accessed by the terminal device is a weak signal cell identifier, it is determined that the terminal device is in a weak signal environment. On the contrary, when the cell identifier of the cell currently accessed by the terminal device is not a weak signal cell identifier, it is determined that the terminal device is not in a weak signal environment. The embodiment of the present application does not limit the weak signal cell identifier. For example, in chain distribution scenarios such as high-speed railways, subways or highways, the signal coverage blind spots are relatively fixed, and the identifiers of cells with poor signals are provided.
[0067] Optionally, the target parameter is the signal strength of the signal received by the terminal device. A first threshold and a second threshold may be pre-set, and the second threshold is greater than or equal to the first threshold. When the signal strength of the signal received by the terminal device is less than or equal to the first threshold, it is determined that the terminal device is in a weak signal environment. When the signal strength of the signal received by the terminal device is greater than or equal to the second threshold, it is determined that the terminal device is not in a weak signal environment. The embodiments of the present application do not limit the values of the first threshold and the second threshold.
[0068] Optionally, the target parameter is the voice packet loss rate of the terminal device. A third threshold and a fourth threshold may be pre-set, with the third threshold being greater than or equal to the fourth threshold. When the voice packet loss rate is greater than or equal to the fourth threshold, the terminal device is determined to be in a weak signal environment. When the voice packet loss rate is less than or equal to the third threshold, the terminal device is determined to be not in a weak signal environment. This embodiment of the application does not limit the values of the third and fourth thresholds.
[0069] 3. Phonetic encoding and decoding
[0070] In the embodiment of the present application, phonetic codec refers to encoding and decoding the phonetic symbol, tone and duration of a single word according to a coding table. The embodiment of the present application does not limit the name of the phonetic codec, for example, it can also be called weak signal high-definition voice coding.
[0071] The phonetic symbol can be in units of characters, words, sentences or other units. The embodiment of the present application is described with characters as an example, and each single character has three pieces of information, namely, phonetic symbol, tone and duration.
[0072] The embodiments of the present application do not limit the classification criteria for duration. For example, duration can include three categories: short, medium, and long. Short means less than 0.5 seconds, medium means greater than or equal to 0.5 seconds and less than 2 seconds, and long means greater than or equal to 2 seconds. For another example, duration can include four categories, 1 to 4. 1 means less than 0.5 seconds, 2 means greater than or equal to 0.5 seconds and less than 1 second, 3 means greater than or equal to 1 second and less than 2 seconds, and 4 means greater than or equal to 2 seconds.
[0073] At the voice sender, the phonetic symbol, pitch, and duration of each word are phonetically encoded according to a coding table to generate encoded information of a preset bit length. This embodiment of the present application does not limit the value of the preset bit length; for example, 22 bits are used. After transmission through a channel, the encoded information reaches the voice receiver, which decodes the received information according to the coding table to obtain the phonetic symbol, pitch, and duration of the word.
[0074] Among them, in different languages, the definitions of phonetic symbols and tones are different, and the embodiments of the present application do not limit this. For example, in Chinese, phonetic symbols refer to the pinyin of a Chinese character. For example, "hello" corresponds to 2 phonetic symbols, namely the pinyin "ni" for "you" and the pinyin "hao" for "good". Tones in Chinese include 4 kinds, namely: yinping (first tone), yangping (second tone), shangsheng (third tone) and qusheng (fourth tone). For another example, in English, phonetic symbols refer to an English word. For example, "Good morning" corresponds to 2 phonetic symbols, namely "good" and "morning". Tones in English can include but are not limited to at least two of the following: affirmative, interrogative, rising tone, falling tone, rising first and then falling, falling first and then rising, flat tone, high tone and low tone.
[0075] 4. Coding table, coding information, global index table and common index table
[0076] A coding table is stored in a terminal device, storing the correspondence between phonetic symbols, tones, durations, and coded information. The terminal device completes phonetic symbol encoding or decoding by searching the coding table. Optionally, the correspondence between phonetic symbols, tones, durations, and coded information may include at least one of the following: a correspondence between phonetic symbols and coded information, a correspondence between tones and coded information, a correspondence between durations and coded information, a correspondence between a combination of phonetic symbols and tones and coded information, or a correspondence between a combination of phonetic symbols, tones, and durations and coded information. The coding table may include one table or at least two tables, depending on the correspondence. Depending on the correspondence, the coded information of a preset bit length may include information components corresponding to different combinations of phonetic symbols, tones, and durations.
[0077] Optionally, the encoded information may include a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration.
[0078] Optionally, the first information component includes a first information sub-component corresponding to the phonetic symbol and a second information sub-component corresponding to the tone. That is, the phonetic symbol, the tone, and the duration each correspond to an information component.
[0079] Optionally, considering that the user's commonly used vocabulary is limited, in order to save lookup time and improve the efficiency of phonetic symbol encoding and decoding, the encoding table may include a global index table and a commonly used index table. The global index table is used for encoding by the voice sender and decoding by the voice receiver, and the commonly used index table is used for encoding by the voice sender. Among them, the global index table includes phonetic symbols and global index values of the phonetic symbols. The global index table can be understood as a complete set of single words or phonetic symbols in a certain language, and the range of global index values is relatively large. The commonly used index table is generated based on the number of times the user uses a single word within a preset time period. The embodiment of the present application does not limit the value of the preset time period, for example, 3 months, 6 months or 1 year. The phonetic symbols included in the commonly used index table have commonly used index values and the global index value of the phonetic symbol in the global index table.
[0080] When performing phonetic encoding and decoding based on the code table, the terminal device can first search the common index table. Because the common index table is generated based on the user's frequently used characters, it has fewer common index values and reduces table lookup time, improving phonetic encoding and decoding efficiency. If no result is found in the common index table, the device then searches the global index table for phonetic encoding and decoding.
[0081] The embodiment of the present application does not limit the value range of the global index value and the phonetic symbol sorting. For example, commonly used characters are sorted first.
[0082] The embodiment of the present application does not limit the range of commonly used index values and the order of phonetic symbols. For example, the number of phonetic symbols can be 5000, and the commonly used index value can be 13 bits in length, representing a maximum of 8192 phonetic symbols. The words can be sorted in descending order based on the frequency of use of the words in a preset time period.
[0083] Optionally, the global index table and the common index table may be updated periodically, and the embodiment of the present application does not limit the update period.
[0084] The following uses Chinese as an example to illustrate the encoding table, encoding information, global index table and common index table.
[0085] Optionally, in one implementation, the encoding table includes: a Chinese global index table, a duration table, and a commonly used index table. The encoding information includes a first information component and a second information component. The Chinese global index table is used to indicate the correspondence between the combination of single words, phonetic symbols, and tones and the first information component (global index value) in the encoding information. The global index value can be 20 bits, representing a maximum of 1.04 million single words. The duration table is used to indicate the correspondence between the duration and the second information component in the encoding information. The commonly used index table includes single words, phonetic symbols, tones, duration, global index values, and commonly used index values. For example, Tables 1 to 3 are used for explanation. Assume that the encoding information is 22 bits in length, wherein the first information component (global index value) is 20 bits in length and the second information component is 2 bits in length. As shown in Table 1, the Chinese global index table can be understood as the complete set of Chinese single words, and each word corresponds to a global index value. Among them, the tone values are 1 to 4, representing the first to fourth tones respectively. As shown in Table 2, durations are classified into three categories: short, medium, and long. The corresponding index value is 2 bits, which is the second information component. As shown in Table 3, commonly used index values can uniquely distinguish the combination of single words, phonetic symbols, tones, and durations.
[0086] Table 1 Chinese global index table
[0087] single word Phonetic Symbols tone Global index value (20 bits) good hao 3 0x00001 bad huai 4 0x00002 many duo 1 0x00003 few Shao 3 0x00004
[0088] Table 2 Duration
[0089] Duration Index value (2 bits) (binary) short 00 middle 01 long 10
[0090] Table 3 Common index table
[0091] single word Phonetic Symbols tone Duration Global index value (20 bits) Common index value (13 bits) good hao 3 short 0x00001 0x0001 good hao 3 middle 0x00001 0x0002 bad huai 4 long 0x00002 0x0003 many duo 1 short 0x00003 0x0004 few Shao 3 long 0x00004 0x0005
[0092] Optionally, in another implementation, the encoding table includes: a Chinese global index table, a duration table, and a commonly used index table. The encoded information includes a first information component and a second information component. The Chinese global index table can be found in Table 1, and the duration table can be found in Table 2. The commonly used index table can be found in Table 4. As shown in Table 4, the commonly used index values can uniquely distinguish combinations of single words, phonetic symbols, and tones.
[0093] In this implementation, the encoded information includes a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration. Because the duration is encoded separately, differences in duration can be ignored in the commonly used index table, further reducing the number of commonly used index values and improving the search speed of the commonly used index table.
[0094] Table 4 Commonly used index table
[0095] single word Phonetic Symbols tone Global index value (20 bits) Common index value (13 bits) good hao 3 0x00001 0x0001 bad huai 4 0x00002 0x0002 many duo 1 0x00003 0x0003 few Shao 3 0x00004 0x0004
[0096] Optionally, in another implementation, the coding table includes: a Chinese global index table, a tone table, a duration table and a commonly used index table. The coding information includes a first information sub-component, a second information sub-component and a third information component. Exemplarily, as shown in Table 5, the Chinese global index table is used to indicate the correspondence between the phonetic symbol and the first information sub-component (global index value) in the coding information. Exemplarily, as shown in Table 6, the tone table is used to indicate the correspondence between the tone and the second information sub-component (2-bit index value in Table 6) in the coding information. Among them, the tone takes values 1 to 4, representing the first to fourth tones respectively. The duration table can be found in Table 2. Exemplarily, as shown in Table 7, the commonly used index values can uniquely distinguish different phonetic symbols.
[0097] In this implementation, the phonetic symbols, tones and duration are encoded and decoded separately without considering the differences between different words, which further reduces the number of common index values and the number of global index values, increases the rate of searching the common index table and the global index table, and improves the encoding and decoding efficiency.
[0098] Table 5 Chinese global index table
[0099] Phonetic Symbols Global index value (18 bits) hao 0x00001 huai 0x00002 duo 0x00003 Shao 0x00004
[0100] Table 6 Tone table
[0101] tone Index value (2 bits) (binary) 1 00 2 01 3 10 4 11
[0102] Table 7 Commonly used index table
[0103] Phonetic Symbols Global index value (18 bits) Common index values (12 bits) hao 0x00001 0x0001 huai 0x00002 0x0002 duo 0x00003 0x0003 Shao 0x00004 0x0004
[0104] It should be noted that when there are multiple languages, each language corresponds to a global index table. Optionally, the multiple languages may include but are not limited to at least two of the following: Chinese, English, German, French, Japanese, Korean or dialects.
[0105] For example, the encoding table includes three global index tables: a Chinese global index table, an English global index table, and a dialect global index table. The Chinese global index table may include 380,000 phonetic symbols, of which 100,000 are commonly used, sorted first. The English global index table may include 280,000 phonetic symbols, of which 35,000 are commonly used, sorted first. The dialect global index table may include 100,000 phonetic symbols.
[0106] It should be noted that the embodiments of the present application do not limit the value range of the global index value in each global index table. Optionally, all global index tables can be uniformly numbered to facilitate quick search when encoding and decoding phonetic symbols. For example, the encoding table includes a Chinese global index table and a dialect global index table. The value range of the global index value in the Chinese global index table is 1 to 100, with a maximum of 100 phonetic symbols. The global index values in the dialect global index table can be numbered starting from 101.
[0107] Optionally, in order to ensure the efficiency of phonetic symbol encoding and decoding, the total number of phonetic symbols in all global index tables is less than a preset value. The embodiment of the present application does not limit the preset value, for example, 1 million.
[0108] 5. The terminal device supports phonetic codec
[0109] Optionally, in one implementation, the terminal device supports the phonetic symbol encoding and decoding function by default, and there is no related setting switch, and no user setting is required, then the terminal device supports the phonetic symbol encoding and decoding.
[0110] Optionally, in another implementation, the terminal device has the function of phonetic symbol encoding and decoding, and there is a related setting switch that needs to be set by the user. The terminal device supports phonetic symbol encoding and decoding, which means that the terminal device currently turns on the phonetic symbol encoding and decoding function through user settings. If the terminal device currently turns off the phonetic symbol encoding and decoding function, then the terminal device does not support phonetic symbol encoding and decoding. The embodiment of the present application does not limit the way in which the user turns on or off the phonetic symbol encoding and decoding function of the terminal device. For example, it can be controlled by any one of the following: voice control, preset gesture control, and control by touch operation in the relevant interface.
[0111] Currently, when a terminal device makes a voice call, it usually uses a traditional voice codec, such as waveform codec or feature codec. When the signal in the communication environment of the terminal device deteriorates, the transmitted voice bit rate decreases and the waveform features become sparser, resulting in a large distortion between the waveform restored by the receiving end and the original waveform, resulting in intermittent calls, noise or even silence, and poor call quality. For example, Figure 3 The following is a diagram showing the effect of a call using traditional voice codec in a weak signal environment. Figure 3 As shown, user A and user B are currently in a weak signal environment. For user B, the call is intermittent and the effect is very poor.
[0112] The voice processing method provided in the embodiment of the present application can obtain coding information of a preset bit length by encoding the phonetic symbols, pitch and duration of the words spoken by the user at the voice sender when any one of the two terminal devices conducting a voice call is in a weak signal environment. After the encoded information is transmitted through the channel, the received information is decoded at the voice receiver to obtain the phonetic symbols, pitch and duration of the words, thereby generating a complete and smooth voice signal, and playing it with a preset sound. The voice processing method provided in the embodiment of the present application can be transmitted at an extremely low bit rate in a weak signal environment. Moreover, by encoding the pronunciation and duration of the words, the receiver can clearly restore the pronunciation of the sender user, and then obtain the semantics based on the played pronunciation, thereby solving the problems of intermittent calls, noise or even silence, and improving the call quality. For example, Figure 4 This is a diagram of the call effect provided by the embodiment of the present application in a weak signal environment. Figure 4 As shown, user A and user B are currently in a weak signal environment. User B can hear the pinyin pronunciation played by the terminal device using the preset sound, which clearly restores the pronunciation of the sender user A. User B can then judge the semantics through the pronunciation, understand user A's intention, and improve the call quality.
[0113] The following are some examples of application scenarios.
[0114] Optionally, in an application scenario, user A and user B are familiar with each other. When user A or user B is in a weak signal environment and needs to communicate, the voice processing method provided in the embodiment of the present application uses fuzzy matching of homophones or similar pronunciations. The receiving user does not need to recognize the exact meaning of the word emitted by the sound source. Instead, the receiving user can rely on the familiarity between the two parties and accurately understand what the other party really wants to say based on the pronunciation played by the terminal device, thereby improving the call quality.
[0115] Optionally, in another application scenario, user A is in an emergency or dangerous environment and the signal is poor. After user A talks with user B, user A can convey short and important words to user B through the voice processing method provided in the embodiment of the present application. User B can accurately understand user A's intentions based on the pronunciation played by the terminal device, thereby improving the call quality.
[0116] The technical solution of the present application is described in detail below through specific embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0117] The terms "first", "second", "third", "fourth", etc. (if any) in the embodiments of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0118] Figure 5 A message interaction diagram of the voice processing method provided in the embodiment of the present application. This embodiment involves a first terminal device and a second terminal device, the first terminal device and the second terminal device are in a call state, the first terminal device is the voice sender, and the second terminal device is the voice receiver. Figure 5 As shown, the speech processing method provided in this embodiment may include:
[0119] S501. The first terminal device obtains a voice signal input by a user.
[0120] For example, in Figure 4 In the example, the voice signal input by the user is the voice of user A "Hello, I'm on the mountain" which is processed by the first terminal device and corresponds to the voice signal.
[0121] S502: When the first terminal device determines that it is currently in a weak signal environment and that both the first terminal device and the second terminal device support phonetic symbol encoding and decoding, extract multiple sets of information from the voice signal, wherein each set of information includes the phonetic symbol, tone, and duration of a single word.
[0122] The first terminal device determining that it is currently in a weak signal environment may include: the first terminal device is in a weak signal environment, or the second terminal device is in a weak signal environment, or both the first terminal device and the second terminal device are in a weak signal environment. Regarding the weak signal environment and the implementation method for determining whether the terminal device is in a weak signal environment, please refer to the above description and will not be repeated here.
[0123] Among them, the terminal device supports phonetic codec, which can be found in the above description and will not be repeated here.
[0124] Optionally, multiple groups of information can be extracted from the speech signal by waveform comparison. Specifically, the waveform of the speech signal is segmented. For each waveform segment, the waveform or waveform feature of the segment is compared with the waveforms or waveform features of all locally pre-stored words, and the word with the greatest similarity in waveform or waveform feature among all words is determined as the word corresponding to the waveform segment. Optionally, the waveform of the speech signal is segmented, and waveform extraction can be performed by word, or segmented according to a specified length. This embodiment does not limit the value of the specified length, for example, 1 second.
[0125] Since the phonetic encoding and decoding adopted in the embodiment of the present application only requires that the pronunciations of the characters are similar, whether the meanings of the characters are consistent is not considered. The efficiency of extracting multiple groups of information can be improved by waveform comparison.
[0126] Optionally, multiple sets of information can be extracted from the speech signal using a neural network model or a machine model. Optionally, the neural network model or machine model is used for semantic recognition, outputting corresponding words based on the input speech signal. Optionally, the neural network model or machine model is used for speech recognition, outputting corresponding phonetic symbols and tones based on the input speech signal.
[0127] Alternatively, the duration of a single word can be determined by energy detection. For example, the level of duration and the energy threshold corresponding to each level can be pre-set, and the duration of a single word can be determined by comparing with multiple energy thresholds. Exemplary, the level of duration can be found in Table 2.
[0128] For example, in Figure 4 In this example, six sets of information can be extracted from the speech signal: the phonetic symbol, pitch, and duration of the words "ni," "hao," "wo," "zai," "shan," and "shang." For the word "ni," the phonetic symbol is ni, the pitch is the third tone, and the duration is assumed to be short. For the word "hao," the phonetic symbol is hao, the pitch is the third tone, and the duration is assumed to be medium.
[0129] S503. The first terminal device obtains the coding information corresponding to each group of information according to the coding table.
[0130] The coding table stores the correspondence between the phonetic symbol, tone, duration and coding information. The coding information is a preset bit length. Please refer to the above description and will not be repeated here.
[0131] S504: The first terminal device sends coded information to the second terminal device.
[0132] Correspondingly, the coded information is received by the second terminal device after being transmitted through the channel. In this embodiment, the information received by the second terminal device from the channel is referred to as first information.
[0133] For example, in Figure 4 In this example, assuming the preset bit length of the encoded information is 22 bits, the first terminal device sends six pieces of encoded information to the second terminal device. These pieces correspond to the characters "you," "good," "I," "in," "mountain," and "up," totaling 22 * 6 = 132 bits. Accordingly, after the encoded information is transmitted over the channel, the second terminal device can receive the 132-bit first information from the channel.
[0134] S505: The second terminal device decodes the first information according to the coding table to obtain multiple groups of information, each group of information including the phonetic symbol, tone and duration of a single word.
[0135] Optionally, decoding the first information according to the coding table to obtain multiple groups of information may include:
[0136] A plurality of coded information is sequentially obtained from the first information, where the length of the coded information is a preset bit length.
[0137] For each piece of coded information, the phonetic symbol, tone and duration corresponding to the coded information are obtained according to the coding table.
[0138] For example, in Figure 4 The first information is 132 bits long. First, the first 22-bit encoded information is obtained from the first information. The phonetic symbol, tone, and duration of the word corresponding to the encoded information are obtained according to the encoding table. Then, the second 22-bit encoded information is obtained from the first information. The phonetic symbol, tone, and duration of the word corresponding to the encoded information are obtained according to the encoding table. This process is repeated until the first information is decoded and the phonetic symbol, tone, and duration of the six words are obtained.
[0139] S506. The second terminal device generates a voice signal according to the multiple groups of information.
[0140] Since each set of information includes the phonetic symbol, tone and duration of a single word, the pronunciation and duration of each word can be restored, thereby synthesizing a complete and smooth speech signal.
[0141] S507: The second terminal device plays the voice signal using a preset sound.
[0142] The present embodiment does not limit the preset voice, for example, it can be a male voice or a female voice.
[0143] It can be seen that the voice processing method provided by this embodiment can be applied to two terminal devices making voice calls in a weak signal environment. At the voice sender, the phonetic symbol, tone and duration of each word in the user's voice are extracted and encoded to obtain the coding information of the preset bit length corresponding to each word. After the coded information is transmitted through the channel, the received information is decoded at the voice receiver accordingly to obtain the phonetic symbol, tone and duration of each word, thereby generating a complete and smooth voice signal, and playing it using the preset sound. The voice processing method provided by the embodiment of the present application can be transmitted at an extremely low bit rate in a weak signal environment because the coded information is of a preset bit length. By encoding and decoding the pronunciation and duration of the words, the receiver can clearly restore the pronunciation of the sender user, and then obtain the semantics based on the played pronunciation, solving the problems of intermittent calls, noise or even silence when using traditional voice encoding and decoding, and improving the call effect.
[0144] Optionally, in S503, obtaining the coding information corresponding to each group of information according to the coding table may include:
[0145] For each group of information, it is determined whether the commonly used index table includes the phonetic symbols in the group of information.
[0146] If the commonly used index table includes the phonetic symbols in the group of information, the encoding information is obtained according to the commonly used index table.
[0147] If the phonetic symbols in the group of information are not included in the common index table, the encoding information is obtained according to the global index table.
[0148] Since the frequently used index table is generated based on the user's frequently used characters, the number of frequently used index values in the frequently used index table is much smaller than the number of global index values in the global index table. Therefore, the frequently used index table is searched for encoding first, which reduces the amount of search data. If the frequently used index table is not found, the global index table is searched for encoding again, which improves encoding efficiency.
[0149] The implementation of the coding table is different, and the phonetic symbol encoding and decoding methods are different. S503 and S505 are explained below with examples.
[0150] Optionally, in one implementation, as shown in Tables 1 to 3 above, the encoding table may include a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration. In S503, for each group of information, the first terminal device may first search in the common index table based on the phonetic symbol, tone, and duration of the single word. If found, the corresponding 20-bit global index value is used as the first information component. If not found, the Chinese global index table is searched based on the phonetic symbol and tone of the single word, and the corresponding 20-bit global index value is used as the first information component. Then, the duration table is searched based on the duration of the single word, and the corresponding 2-bit index value is used as the second information component, thereby obtaining 22 bits of encoding information. Accordingly, in S505, the second terminal device obtains 22 bits of encoding information, the first 20 bits being the first information component, and the last 2 bits being the second information component. The phonetic symbol and tone of the single word are obtained by searching in the Chinese global index table based on the first 20 bits. The duration of the word is obtained by searching the duration table based on the last 2 bits.
[0151] Alternatively, in another implementation, such as the encoding tables shown in Tables 1, 2, and 4 above, this implementation differs from the above implementation in that when the first terminal device searches the commonly used index table, it searches the commonly used index table based on individual words, their phonetic symbols, and tones. In this implementation, because the commonly used index table does not take duration into account, the number of commonly used index values is further reduced, the search speed is faster, and the encoding efficiency is improved.
[0152] Optionally, in another implementation, as shown in Tables 2, 5, 6, and 7 above, the encoding tables may include a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration. The first information component includes a first information sub-component corresponding to the phonetic symbol and a second information sub-component corresponding to the tone. In S503, for each group of information, the first terminal device may first search the common index table based on the phonetic symbol of the single word. If found, the corresponding 18-bit global index value is used as the first information sub-component. If not found, the first terminal device may search the Chinese global index table based on the phonetic symbol of the single word and use the corresponding 18-bit global index value as the first information sub-component. Then, the first terminal device may search the tone table based on the tone of the single word and use the corresponding 2-bit index value as the second information sub-component. Thereafter, the first terminal device may search the duration table based on the duration of the single word and use the corresponding 2-bit index value as the second information component, thereby obtaining 18+2+2=22 bits of encoded information. Accordingly, in S505, the second terminal device obtains 22 bits of encoded information, with the first 18 bits representing the first information subcomponent, the middle 2 bits representing the second information subcomponent, and the last 2 bits representing the second information subcomponent. The first 18 bits are searched in the Chinese global index table to obtain the phonetic symbol of the word. The middle 2 bits are searched in the tone table to obtain the tone of the word. The last 2 bits are searched in the duration table to obtain the duration of the word.
[0153] In this implementation, the phonetic symbols, tones, and durations are encoded and decoded separately, the number of common index values and the number of global index values are further reduced, the search speed is faster, and the encoding and decoding efficiency is improved.
[0154] Optionally, in another embodiment of the present application, in the above Figure 5 Based on the embodiment shown, an implementation method for determining that the terminal device is in a weak signal environment in S502 is provided. Through negotiation between the first terminal device and the second terminal device, it is determined that the phonetic codec can be used when the terminal device is currently in a weak signal environment.
[0155] Optionally, in one implementation, as Figure 6 As shown, the first terminal device determines that it is currently in a weak signal environment, which may include:
[0156] S601. If the first terminal device determines that the target parameter of the first terminal device meets the first preset condition, it sends a first request message to the second terminal device, where the first request message is used to instruct the second terminal device to use phonetic codec.
[0157] The target parameter of the first terminal device is used to indicate the signal state of the communication environment that the first terminal device is currently in. The target parameter and the first preset condition can be found in the above description of this application and will not be repeated here.
[0158] Correspondingly, the second terminal device receives the first request message.
[0159] S602. The second terminal device sends a first response message to the first terminal device, where the first response message is used to instruct the second terminal device to use phonetic codec.
[0160] In this implementation, the first terminal device, which is the voice sender, determines that it is currently in a weak signal environment based on its own target parameters, and actively initiates negotiation for switching the codec mode to the second terminal device, thereby ensuring the timely adoption of phonetic codecs and improving call quality.
[0161] Optionally, the first request message may include a first indication field for indicating the use of phonetic codec. The embodiment of the present application does not limit the name of the first indication field.
[0162] Optionally, the first request message may include a hysteresis timer, where the hysteresis timer is used to indicate a delay time for the second terminal device to use phonetic symbol encoding and decoding.
[0163] Typically, terminal devices use the traditional voice codec by default. By setting a hysteresis timer, time is reserved for switching between voice codecs. After the hysteresis timer expires, the first and second terminal devices simultaneously use the phonetic codec, improving the switching effect of the codec.
[0164] Optionally, in another implementation, such as Figure 7 As shown, the first terminal device determines that it is currently in a weak signal environment, which may include:
[0165] S701. If the second terminal device determines that the target parameter of the second terminal device meets the first preset condition, it sends a second request message to the first terminal device, where the second request message is used to instruct the first terminal device to use phonetic codec.
[0166] The target parameter of the second terminal device is used to indicate the signal state of the communication environment that the second terminal device is currently in. The target parameter and the first preset condition can be found in the above description of this application and will not be repeated here.
[0167] Correspondingly, the first terminal device receives the second request message.
[0168] Optionally, the second request message may include a first indication field. Please refer to the above description of the first indication field and will not be repeated here.
[0169] Optionally, the second request message may include a hysteresis timer, where the hysteresis timer is used to indicate a delay time for the first terminal device to use phonetic symbol encoding and decoding.
[0170] S702. The first terminal device sends a second response message to the second terminal device, where the second response message is used to instruct the first terminal device to use phonetic codec.
[0171] In this implementation, the second terminal device, which serves as the voice receiver, determines that it is currently in a weak signal environment based on its own target parameters, and actively initiates negotiation for switching the codec mode to the first terminal device, thereby ensuring timely adoption of the phonetic codec mode and improving call quality.
[0172] Optionally, in the above process, if the first response message or the second response message is not received successfully, it can be resent. Optionally, the number of resends can be set, and this embodiment does not limit the specific value.
[0173] Optionally, in another embodiment of the present application, in the above Figure 5 Based on the illustrated embodiment, an implementation method for determining in S502 that both the first terminal device and the second terminal device support the phonetic codec is provided. Through capability negotiation between the first terminal device and the second terminal device, it is determined that both parties in the call support the phonetic codec, and the phonetic codec can be used in a weak signal environment.
[0174] Optionally, in one implementation, as Figure 8 As shown, the first terminal device determines that both the first terminal device and the second terminal device support phonetic codec, which may include:
[0175] S801. A first terminal device sends first capability information to a second terminal device, where the first capability information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding.
[0176] Correspondingly, the second terminal device receives the first capability information sent by the first terminal device.
[0177] S802. The second terminal device sends first capability response information to the first terminal device, where the first capability response information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding.
[0178] In this implementation, the first terminal device, which is the voice sender, actively initiates capability negotiation with the second terminal device, thereby ensuring timely adoption of the phonetic codec method and improving call quality.
[0179] Optionally, in another implementation, such as Figure 9 As shown, the first terminal device determines that both the first terminal device and the second terminal device support phonetic codec, which may include:
[0180] S901. The second terminal device sends second capability information to the first terminal device, where the second capability information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding.
[0181] Correspondingly, the first terminal device receives the second capability information sent by the second terminal device.
[0182] S902. The first terminal device sends second capability response information to the second terminal device, where the second capability response information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding.
[0183] In this implementation, the second terminal device, which is the voice receiver, actively initiates capability negotiation with the first terminal device, thereby ensuring timely adoption of the phonetic codec method and improving call quality.
[0184] It should be noted that the first capability information and the second capability information can be separate messages or carried in an existing message. This embodiment does not limit the time of the capability negotiation process. For example, after the first terminal device and the second terminal device establish a connection, the capability negotiation can be performed during the altering message. The altering message can include the Newaudiocodec capability field to indicate whether the terminal device supports phonetic codec.
[0185] Optionally, in the above process, if the first capability response information or the second capability response information is not received successfully, it can be resent. Optionally, the number of resends can be set, and this embodiment does not limit the specific value.
[0186] Optionally, in a scenario where the user can set the terminal device to turn on or off the phonetic symbol codec function, the terminal device supporting phonetic symbol codec means that the phonetic symbol codec function is currently turned on by the terminal device through user settings.
[0187] Optionally, if the terminal device currently does not have the phonetic symbol encoding and decoding function turned on, the user can set it up. The speech processing method provided in this embodiment may also include:
[0188] The settings interface is displayed.
[0189] Receive user operations in the settings interface.
[0190] In response to the operation, the phonetic symbol encoding and decoding function is enabled.
[0191] For example, Figure 10 This is an interface diagram for setting the phonetic codec mode provided in the embodiment of the present application. Figure 10 As shown in (a), the terminal device currently displays the setting interface 1001, which includes the function option "weak signal high-definition voice encoding", that is, the phonetic codec function in the embodiment of the present application. The status of the control 1010 can show whether the current terminal device has the phonetic codec function turned on. Figure 10In (a), the phonetic codec function is in a closed state. The user can click on the control 1010, and accordingly, the terminal device responds to the click operation and turns on the phonetic codec function, such as Figure 10 As shown in (b) in .
[0192] It should be noted that this embodiment does not limit the time when the user sets the phonetic symbol encoding and decoding function.
[0193] Optionally, if the terminal device currently does not enable the phonetic symbol encoding and decoding function, and the terminal device determines that it is currently in a weak signal environment, the voice processing method provided in this embodiment may further include:
[0194] Generate and output prompt information.
[0195] A first instruction from a user is received.
[0196] Enable the phonetic symbol encoding and decoding function according to the first instruction.
[0197] This embodiment does not limit the implementation of the prompt information. For example, it can be playing a prompt voice, playing prompt music, or popping up a prompt box or prompt information in the interface currently displayed on the terminal device.
[0198] Optionally, in another embodiment of the present application, an implementation method for switching from phonetic codec to traditional voice codec is provided on the basis of the above embodiment. The communication environment in which the two terminal devices of the call are located changes in real time and will not always be in a weak signal environment. The phonetic codec in the embodiment of the present application is more suitable for weak signal environments and meets basic communication needs. When the communication environment improves, it should be switched back to the traditional voice codec in time to enhance the user's call experience. Through negotiation between the first terminal device and the second terminal device, it is determined that the traditional voice codec can be used when the current environment is not a weak signal environment.
[0199] Optionally, in one implementation, as Figure 11 As shown, the speech processing method provided in this embodiment may further include:
[0200] S1101. If the first terminal device determines that the target parameter of the first terminal device meets the second preset condition, the third request message is sent to the second terminal device, where the third request message is used to instruct the second terminal device to use waveform coding and decoding.
[0201] The target parameter of the first terminal device is used to indicate the signal state of the communication environment that the first terminal device is currently in. The target parameter and the second preset condition can be found in the above description of this application and will not be repeated here.
[0202] Correspondingly, the second terminal device receives the third request message.
[0203] Optionally, the third request message may include a second indication field for indicating the use of a traditional voice codec. This embodiment of the application does not limit the name of the second indication field. For example, the name may be "back to HD." Optionally, the second indication field may also indicate the start time of using the traditional voice codec.
[0204] S1102: If the second terminal device determines that the target parameter of the second terminal device meets the second preset condition, it sends a third response message to the first terminal device, where the third response message is used to instruct the second terminal device to use waveform coding and decoding.
[0205] The target parameter of the second terminal device is used to indicate the signal state of the communication environment that the second terminal device is currently in. The target parameter and the second preset condition can be found in the above description of this application and will not be repeated here.
[0206] In this implementation, the first terminal device, which is the voice sender, actively initiates negotiation on switching the codec mode to the second terminal device when it determines that its signal environment has improved. When the second terminal device also determines that it is not currently in a weak signal environment, it returns a response message to ensure that the traditional voice codec mode is adopted in a timely manner when the communication environment is better, thereby improving the call quality.
[0207] Optionally, in another implementation, such as Figure 12 As shown, the speech processing method provided in this embodiment may further include:
[0208] S1201. If the second terminal device determines that the target parameter of the second terminal device meets the second preset condition, it sends a fourth request message to the first terminal device, where the fourth request message is used to instruct the first terminal device to use waveform coding and decoding.
[0209] Correspondingly, the first terminal device receives the fourth request message.
[0210] Optionally, the fourth request message may include a second indication field. Please refer to the above description of the second indication field and will not be repeated here.
[0211] S1202. If the first terminal device determines that the target parameter of the first terminal device meets the second preset condition, it sends a fourth response message to the second terminal device, where the fourth response message is used to instruct the first terminal device to use waveform coding and decoding.
[0212] In this implementation, the second terminal device, which serves as the voice receiver, actively initiates negotiation on switching the codec mode to the first terminal device when it determines that its signal environment has improved. When the first terminal device also determines that it is not currently in a weak signal environment, it returns a response message to ensure that the traditional voice codec mode is adopted in a timely manner when the communication environment is better, thereby improving the call quality.
[0213] Optionally, in the above process, if the third response message or the fourth response message is not received successfully, it can be resent. Optionally, the number of resends can be set, and this embodiment does not limit the specific value.
[0214] It is understandable that, in order to implement the above functions, the terminal device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of this application.
[0215] The embodiment of the present application can divide the functional modules of the terminal device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. It should be noted that the names of the modules in the embodiment of the present application are schematic and are not limited in actual implementation.
[0216] In the case of dividing each functional module into corresponding functional modules, Figure 13 A schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 13 As shown, the terminal device may include: a sending module 1301, a processing module 1302 and a receiving module 1303.
[0217] The sending module 1301 is configured to send data to other devices, for example, coding information, a first request message, a second response message, a third request message, a fourth response message, first capability information, or second capability response information.
[0218] The receiving module 1303 is configured to receive data from other devices, for example, first information, second request message, first response message, fourth request message, third response message, second capability information, or first capability response information.
[0219] Processing module 1302 is used to obtain the voice signal input by the user, extract multiple groups of information from the voice signal, and obtain the encoding information corresponding to each group of information according to the encoding table; decode the first information according to the encoding table to obtain multiple groups of information, generate a voice signal based on the multiple groups of information, and play the voice signal using a preset sound, etc.
[0220] It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0221] Please refer to Figure 14 , which shows another structure of the terminal device provided in an embodiment of the present application, the terminal device includes: a processor 1401, a receiver 1402, a transmitter 1403, a memory 1404 and a bus 1405. The processor 1401 includes one or more processing cores, and the processor 1401 executes various functional applications and information processing by running software programs and modules. The receiver 1402 and the transmitter 1403 can be implemented as a communication component, which can be a baseband chip. The memory 1404 is connected to the processor 1401 via the bus 1405. The memory 1404 can be used to store at least one program instruction, and the processor 1401 is used to execute at least one program instruction to implement the technical solution of the above embodiment. Its implementation principle and technical effects are similar to those of the above method-related embodiments, and will not be repeated here.
[0222] When the terminal is powered on, the processor reads the software program in the memory, interprets and executes the program's instructions, and processes the program's data. When data needs to be sent via the antenna, the processor performs baseband processing on the data to be transmitted and outputs the baseband signal to the control circuit within the control circuit. The control circuit then performs radio frequency processing on the baseband signal and transmits the radio frequency signal as electromagnetic waves via the antenna. When data is sent to the terminal, the control circuit receives the radio frequency signal via the antenna, converts the radio frequency signal into a baseband signal, and outputs the baseband signal to the processor, which converts the baseband signal into data and processes the data.
[0223] Those skilled in the art will understand that for ease of explanation, Figure 14 Only one memory and processor are shown. In an actual terminal, there may be multiple processors and memories. The memory may also be referred to as a storage medium or storage device, etc., which is not limited in the embodiments of the present application.
[0224] As an optional implementation, the processor may include a baseband processor and a central processing unit (CPU). The baseband processor is primarily responsible for processing communication data, while the CPU is primarily responsible for executing software programs and processing data from these programs. Those skilled in the art will appreciate that the baseband processor and the CPU may be integrated into a single processor or may be separate processors interconnected via a bus or other technology. Those skilled in the art will appreciate that a terminal may include multiple baseband processors to accommodate different network standards, multiple CPUs to enhance processing capabilities, and that the various components of the terminal may be connected via various buses. The baseband processor may also be referred to as a baseband processing circuit or a baseband processing chip. The CPU may also be referred to as a central processing circuit or a central processing chip. The functionality for processing communication protocols and communication data may be built into the processor or stored in a memory as a software program, which is then executed by the processor to implement the baseband processing functionality. The memory may be integrated into the processor or independent of the processor. The memory includes a cache to store frequently accessed data / instructions.
[0225] In the embodiments of the present application, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0226] In the embodiments of the present application, the memory may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SS), or a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, without limitation.
[0227] The memory in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data. The methods provided in the various embodiments of the present application may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DWD), or a semiconductor medium (e.g., an SSD), etc.
[0228] The present application provides a computer program product that, when executed on a terminal, enables the terminal to execute the technical solution in the above embodiment. The implementation principle and technical effects are similar to those of the above related embodiments and will not be described in detail here.
[0229] The embodiment of the present application provides a computer-readable storage medium on which program instructions are stored. When the program instructions are executed by a terminal, the terminal executes the technical solution of the above-mentioned embodiment. Its implementation principle and technical effect are similar to those of the above-mentioned related embodiments and will not be repeated here. In summary, the above embodiments are only used to illustrate the technical solution of the present application, rather than to limit it. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in this field should understand that it is still possible to modify the technical solutions described in the above-mentioned embodiments, or to replace some of the technical features therein with equivalents; and these modifications or replacements do not make the essence of the corresponding technical solution deviate from the scope of the technical solution of each embodiment of the present application.
Claims
1. A speech processing method, characterized in that: Applied to a first terminal device, where the first terminal device is in a call state with a second terminal device, the method includes: Obtain the voice signal input by the user; When it is determined that the current environment is a weak signal environment and it is determined that both the first terminal device and the second terminal device support phonetic symbol encoding and decoding, extracting multiple groups of information from the voice signal, each group of information including the phonetic symbol, tone, and duration of a single word; the phonetic symbol encoding and decoding refers to encoding and decoding the phonetic symbol, tone, and duration; Obtaining coding information corresponding to each set of information according to a coding table; the coding table stores the corresponding relationship between phonetic symbols, tones and durations and coding information; The encoded information is sent to the second terminal device via channel transmission.
2. The method according to claim 1, characterized in that The encoded information includes a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration.
3. The method according to claim 2, characterized in that The first information component includes a first information sub-component corresponding to the phonetic symbol and a second information sub-component corresponding to the tone.
4. The method according to claim 1, wherein The coding table includes a global index table and a commonly used index table. The commonly used index table is generated based on the number of times a user uses a single word within a preset time period. The global index table includes phonetic symbols and global index values of the phonetic symbols. The phonetic symbols included in the commonly used index table have commonly used index values and the global index values of the phonetic symbols in the global index table.
5. The method according to claim 4, characterized in that The obtaining, according to the coding table, the coding information corresponding to each group of information includes: For each group of information, determining whether the commonly used index table includes the phonetic symbols in the group of information; If the commonly used index table includes the phonetic symbols in the group of information, obtaining the encoding information according to the commonly used index table; If the commonly used index table does not include the phonetic symbols in the group of information, the encoding information is obtained according to the global index table.
6. The method according to any one of claims 1 to 5, characterized in that The determining that the current signal environment is weak includes: If it is determined that the target parameter meets the first preset condition, sending a first request message to the second terminal device, where the first request message is used to instruct the second terminal device to use the phonetic symbol codec, and the target parameter is used to indicate the signal state of the communication environment currently located by the first terminal device; A first response message sent by the second terminal device is received, where the first response message is used to instruct the second terminal device to use the phonetic symbol codec.
7. The method according to claim 6, characterized in that The first request message includes a hysteresis timer, and the hysteresis timer is used to indicate the delay time for the second terminal device to use the phonetic symbol codec.
8. The method according to any one of claims 1 to 5, characterized in that The determining that the current signal environment is weak includes: receiving a second request message sent by the second terminal device, where the second request message is used to instruct the first terminal device to use the phonetic symbol codec; A second response message is sent to the second terminal device, where the second response message is used to instruct the first terminal device to use the phonetic symbol codec.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: If it is determined that the target parameter meets the second preset condition, sending a third request message to the second terminal device, where the third request message is used to instruct the second terminal device to use waveform codec, and the target parameter is used to indicate the signal state of the communication environment currently located by the first terminal device; A third response message sent by the second terminal device is received, where the third response message is used to instruct the second terminal device to use the waveform codec.
10. The method according to any one of claims 1 to 8, characterized in that The method further comprises: receiving a fourth request message sent by the second terminal device, where the fourth request message is used to instruct the first terminal device to use waveform coding and decoding; If it is determined that the target parameter meets the second preset condition, a fourth response message is sent to the second terminal device, where the fourth response message is used to instruct the first terminal device to use the waveform codec, and the target parameter is used to indicate the signal status of the communication environment in which the first terminal device is currently located.
11. The method according to claim 6, 9 or 10, characterized in that: The target parameter includes at least one of the following: location information of the first terminal device, a cell identifier of a cell currently accessed by the first terminal device, signal strength or voice packet loss rate of a signal received by the first terminal device.
12. The method according to any one of claims 1 to 11, characterized in that Determining that the first terminal device supports phonetic symbol encoding and decoding includes: If it is determined that the first terminal device does not have the phonetic symbol encoding and decoding function turned on, generating and outputting a prompt message; receiving a first instruction from the user; The function is enabled according to the first instruction.
13. The method according to any one of claims 1 to 11, characterized in that The determining that both the first terminal device and the second terminal device support phonetic symbol encoding and decoding includes: Sending first capability information to the second terminal device, where the first capability information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding; Receiving first capability response information sent by the second terminal device, where the first capability response information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding; or, receiving second capability information sent by the second terminal device, where the second capability information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding; Send second capability response information to the second terminal device, where the second capability response information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding.
14. The method according to any one of claims 1 to 11, characterized in that The method further comprises: Display settings interface; Receiving an operation of the user in the setting interface; In response to the operation, the phonetic symbol encoding and decoding function is enabled.
15. A speech processing method, characterized in that: Applied to a second terminal device, where the second terminal device is in a call state with the first terminal device, the method includes: Receiving first information sent by the first terminal device via channel transmission; Decoding the first information according to a coding table to obtain multiple groups of information; each group of information includes the phonetic symbol, tone and duration of a single word, and the coding table stores the corresponding relationship between the phonetic symbol, tone and duration and the coding information; generating a speech signal according to the plurality of sets of information; The voice signal is played using a preset sound.
16. The method according to claim 15, characterized in that The decoding of the first information according to the coding table to obtain multiple groups of information includes: Sequentially acquiring a plurality of coded information from the first information, wherein the length of the coded information is a preset bit length; For each piece of the coding information, the phonetic symbol, tone and duration corresponding to the coding information are obtained according to the coding table.
17. The method according to claim 16, characterized in that The encoded information includes a first information component corresponding to the phonetic symbol and tone, and a second information component corresponding to the duration.
18. The method according to claim 17, characterized in that The first information component includes a first information sub-component corresponding to the phonetic symbol and a second information sub-component corresponding to the tone.
19. The method according to claim 15, characterized in that The coding table includes a global index table and a commonly used index table. The commonly used index table is generated based on the number of times a user uses a single word within a preset time period. The global index table includes phonetic symbols and global index values of the phonetic symbols. The phonetic symbols included in the commonly used index table have commonly used index values and the global index values of the phonetic symbols in the global index table.
20. The method according to any one of claims 15 to 19, characterized in that Before receiving the first information sent by the first terminal device, the method further includes: receiving a first request message sent by the first terminal device, wherein the first request message is used to instruct the second terminal device to use phonetic symbol encoding and decoding; the phonetic symbol encoding and decoding refers to encoding and decoding phonetic symbols, tones, and duration; A first response message is sent to the first terminal device, where the first response message is used to instruct the second terminal device to use the phonetic symbol codec.
21. The method according to claim 20, characterized in that The first request message includes a hysteresis timer, and the hysteresis timer is used to indicate the delay time for the second terminal device to use the phonetic symbol codec.
22. The method according to any one of claims 15 to 19, characterized in that Before receiving the first information sent by the first terminal device, the method further includes: If it is determined that the target parameter meets the first preset condition, sending a second request message to the first terminal device, where the second request message is used to instruct the first terminal device to use the phonetic symbol codec, where the phonetic symbol codec refers to encoding and decoding phonetic symbols, tones, and durations, and the target parameter is used to indicate a signal state of a communication environment in which the second terminal device is currently located; Receive a second response message sent by the first terminal device, where the second response message is used to instruct the first terminal device to use the phonetic symbol codec.
23. The method according to any one of claims 15 to 22, characterized in that The method further comprises: receiving a third request message sent by the first terminal device, where the third request message is used to instruct the second terminal device to use waveform coding and decoding; If it is determined that the target parameter meets the second preset condition, a third response message is sent to the first terminal device, where the third response message is used to instruct the second terminal device to use the waveform codec, and the target parameter is used to indicate the signal status of the communication environment in which the first terminal device is currently located.
24. The method according to any one of claims 15 to 22, characterized in that The method further comprises: If it is determined that the target parameter meets the second preset condition, sending a fourth request message to the first terminal device, where the fourth request message is used to instruct the first terminal device to use waveform coding and decoding; A fourth response message sent by the first terminal device is received, where the fourth response message is used to instruct the first terminal device to use the waveform codec.
25. The method according to any one of claims 22 to 24, characterized in that The target parameter includes at least one of the following: location information of the second terminal device, a cell identifier of a cell currently accessed by the second terminal device, signal strength or voice packet loss rate of a signal received by the second terminal device.
26. The method according to any one of claims 15 to 25, characterized in that Before receiving the first information sent by the first terminal device, the method further includes: receiving first capability information sent by the first terminal device, where the first capability information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding; Sending first capability response information to the first terminal device, where the first capability response information is used to indicate that the second terminal device supports phonetic symbol codec; the phonetic symbol codec refers to encoding and decoding phonetic symbols, tones, and durations; or, Second capability information sent to the first terminal device, where the second capability information is used to indicate that the second terminal device supports phonetic symbol encoding and decoding; Receive second capability response information sent by the first terminal device, where the second capability response information is used to indicate that the first terminal device supports phonetic symbol encoding and decoding.
27. The method according to any one of claims 15 to 25, characterized in that The method further comprises: Display settings interface; Receiving user operations in the setting interface; In response to the operation, a phonetic symbol encoding and decoding function is enabled, where the phonetic symbol encoding and decoding refers to encoding and decoding phonetic symbols, tones, and duration.
28. A terminal device, characterized in that: The system comprises a processor, a memory and a transceiver, wherein the transceiver is used to communicate with other devices, and the processor is used to call a program stored in the memory to execute the method according to any one of claims 1 to 27.
29. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a terminal device, the terminal device is caused to execute the method according to any one of claims 1 to 27.
Citation Information
Patent Citations
Voice processing method and related equipment
CN110913073A