Voice data transmission method

By selecting the target encoder based on uplink bandwidth and voice characteristics, the problem of uneven switching between voice codecs is solved, enabling smooth switching between codecs in wireless communication, improving user experience and adapting to encoding effects under different bandwidth conditions.

CN115527544BActive Publication Date: 2026-04-28NANJING BIG FISH SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING BIG FISH SEMICON CO LTD
Filing Date
2022-08-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing speech codecs cannot achieve dynamic and smooth codec switching, resulting in poor user experience, high resource consumption, long switching latency, and loss of high-frequency information.

Method used

The target encoder is selected based on uplink bandwidth and voice characteristics. The encoder-decoder is switched smoothly by sending decoding instruction information. The encoder suitable for the current bandwidth and voice characteristics is selected for encoding, reducing the impact of switching on user experience.

Benefits of technology

It achieves smooth switching of codecs in wireless communication, reduces the impact of switching on user experience, adapts to the optimal encoding effect under different bandwidth conditions, and does not require modification of the codec's internal implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527544B_ABST
    Figure CN115527544B_ABST
Patent Text Reader

Abstract

The application provides a voice data transmission method, and relates to the technical field of audio processing. The voice data transmission method comprises the following steps: determining a first target encoder from a first encoder and a second encoder according to uplink bandwidth and voice characteristics of to-be-sent voice data, encoding the to-be-sent voice data by using the first target encoder to obtain first voice coding data, and sending a first voice data packet comprising first decoding instruction information and the first voice coding data to a receiving party. According to the scheme, the first target encoder is determined by combining voice characteristics and uplink bandwidth of a physical layer during switching, and the switching time of voice coding and decoding is selected based on this, so that the influence of switching on user experience is reduced. The voice transmission method in the wireless voice communication scene where the physical layer bandwidth of wireless communication is dynamically adjusted according to the channel environment realizes smooth switching of voice coding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and more specifically, to a method for transmitting voice data. Background Technology

[0002] Voice codec functionality is a crucial factor affecting the quality of voice communication, while the bitrate of voice data is limited by the data transmission rate. Wireless communication is influenced by the channel environment, and the transmission rate is usually dynamic. It's generally impossible to achieve optimal performance at all rates using a single voice codec; therefore, it's necessary to switch between different voice codecs based on the bitrate.

[0003] However, current speech codecs cannot achieve dynamic and smooth codec switching. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of the prior art by providing a voice data transmission method to achieve dynamic and smooth codec switching in voice communication.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide a voice data transmission method, including:

[0007] Based on the uplink bandwidth and the voice characteristics of the voice data to be transmitted, a first target encoder is determined from the first encoder and the second encoder; the first encoder and the second encoder correspond to different bit rate ranges.

[0008] The voice data to be transmitted is encoded using the first target encoder to obtain first voice encoded data;

[0009] A first voice data packet is sent to the receiver. The first voice data packet includes: first decoding indication information and first voice encoding data. The first decoding indication information is used to instruct the receiver to decode the first voice encoding data using the first target decoder to obtain the voice data to be sent.

[0010] Optionally, if the bitrate of the first encoder is less than that of the second encoder, the bandwidth ranges corresponding to the first encoder and the second encoder overlap; the step of determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the speech features of the voice data to be transmitted includes:

[0011] If the uplink bandwidth is within the first bandwidth range, and the original encoder is the second encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the first bandwidth range is: greater than the minimum bandwidth of the bandwidth range corresponding to the second encoder, and less than or equal to the first preset low bandwidth threshold.

[0012] Optionally, the voice features include: mute status indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0013] If the duration of the original encoder being the second encoder exceeds the first preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder.

[0014] Alternatively, if the duration of the original encoder being the second encoder exceeds the second preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the first encoder is determined to be the first target encoder.

[0015] Optionally, the method further includes:

[0016] If the uplink bandwidth is within the first bandwidth range, and the duration during which the original encoder is the second encoder exceeds the third preset duration, then the first encoder is determined to be the second target encoder.

[0017] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0018] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0019] Optionally, determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes:

[0020] If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a first preset high bandwidth threshold, and less than or equal to the maximum bandwidth of the bandwidth range corresponding to the first encoder.

[0021] Optionally, the voice features include: mute status indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0022] If the duration of the original encoder being the first encoder exceeds the fourth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder.

[0023] Alternatively, if the duration of the original encoder being the first encoder exceeds a fifth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the second encoder is determined to be the first target encoder.

[0024] Optionally, the method further includes:

[0025] If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the sixth preset duration, then the second encoder is determined to be the second target encoder.

[0026] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0027] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0028] Optionally, if the bitrate of the first encoder is less than that of the second encoder, the bandwidth ranges corresponding to the first encoder and the second encoder do not overlap; determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the speech features of the voice data to be transmitted includes:

[0029] If the uplink bandwidth is within the first bandwidth range, and the original encoder is the second encoder, then based on the speech features, a first target encoder is determined from the first encoder and the second encoder, wherein the first bandwidth range is: the minimum bandwidth greater than the bandwidth range corresponding to the second encoder, and less than or equal to the second preset low bandwidth threshold.

[0030] Optionally, the voice features include: mute status indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0031] If the duration of the original encoder being the second encoder exceeds the seventh preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder.

[0032] Alternatively, if the duration of the original encoder being the second encoder exceeds an eighth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the first encoder is determined to be the first target encoder.

[0033] Optionally, the method further includes:

[0034] If the uplink bandwidth is within the first bandwidth range, and the duration for which the original encoder is the second encoder exceeds the ninth preset duration, then the first encoder is determined to be the second target encoder.

[0035] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0036] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0037] Optionally, determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes:

[0038] If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a second preset low bandwidth threshold, and less than or equal to a second preset high bandwidth threshold.

[0039] Optionally, the voice features include: mute status indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0040] If the duration of the original encoder being the first encoder exceeds the tenth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder.

[0041] Alternatively, if the duration of the original encoder being the first encoder exceeds the eleventh preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the second encoder is determined to be the first target encoder.

[0042] Optionally, the method further includes:

[0043] If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the twelfth preset duration, then the second encoder is determined to be the second target encoder.

[0044] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0045] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0046] The beneficial effects of this application are as follows: This application provides a voice data transmission method. Based on the uplink bandwidth and the voice characteristics of the voice data to be transmitted, a first target encoder is determined from a first encoder and a second encoder. The first target encoder is used to encode the voice data to be transmitted to obtain first voice encoded data. A first voice data packet including first decoding instruction information and the first voice encoded data is then sent to the receiver. This application provides a voice transmission method in a wireless voice communication scenario where the physical layer bandwidth of wireless communication dynamically adjusts with the channel environment, achieving smooth switching of the voice codec. According to the scheme of this application, the first target encoder is determined by combining voice characteristics and the uplink bandwidth of the physical layer during switching, and the switching time of the voice codec is selected based on this, thereby reducing the impact of switching on the user experience. In addition, the scheme of this application can select the voice encoder suitable for the optimal effect of the bandwidth when the physical layer is in different uplink bandwidth states, and the switching method does not require modification of the internal implementation of the codec, nor does it impose limitations on the frame length, implementation algorithm, sampling rate, etc. of the encoder. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This application provides a schematic diagram of a voice data transmission system structure according to an embodiment of the present application.

[0049] Figure 2 A flowchart of a voice data transmission method provided in one embodiment of this application;

[0050] Figure 3 A flowchart illustrating a voice data transmission method provided in another embodiment of this application;

[0051] Figure 4 A flowchart illustrating a voice data transmission method provided in another embodiment of this application;

[0052] Figure 5 A schematic diagram of a transmitting device provided in an embodiment of this application;

[0053] Figure 6 This is a schematic diagram of a receiving device provided in an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0055] In this application, unless otherwise expressly specified and limited, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one feature. In the description of this invention, "a plurality of" means at least two, such as two or three, unless otherwise expressly specified. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0056] Voice codec functionality is a crucial factor affecting the quality of voice communication, as the bitrate of voice data is limited by the data transmission rate. Different voice codes require different bitrates, and at different bitrates, the output voice quality will vary. Furthermore, the transmission rate of wireless communication, a vital method of voice data transmission, is typically influenced by the channel environment. Due to these multiple factors, using a single voice codec cannot achieve optimal performance at all data rates; therefore, it is necessary to switch between voice codecs based on the bitrate.

[0057] Currently, encoders that support dynamic bitrate mainly have the following problems: First, the adjustment range of some encoders that support dynamic bitrate is limited. For example, the AMR-NB codec supports dynamic bitrate, but its supported dynamic bitrate range is from 4.75kbps to 12.2kbps. It cannot be used for transmission bandwidths less than 4.75kbps, and it cannot fully utilize transmission bandwidths greater than 12.2kbps.

[0058] Second, some encoders that support a large adjustment range do not perform well in encoding parts within that range. For example, Opus encoding supports a large adjustment range, but at the same bitrate as AMR-NB, its performance is not as good as AMR-NB.

[0059] In addition to these issues, current voice data transmission also suffers from limited available computing resources, poor user experience, high latency during codec switching, and loss of high-frequency information. Specifically, these problems may manifest in the following ways:

[0060] Voice encoding and decoding are typically performed by embedded devices. Advanced encoding and decoding often require more resources, but embedded devices have limited available computing resources such as memory and CPU, making it impossible to support multiple encoders working simultaneously. How to dynamically and smoothly switch between encoding and decoding without adding extra resources is a challenging problem to solve.

[0061] The encoding and decoding switching does not take into account the characteristics of the speech itself, which may cause the switching point to be in the middle of the pronunciation, resulting in obvious differences in the pronunciation before and after, which seriously affects the user experience;

[0062] The initialization of the new codec and the release of the old codec take time, which may cause additional delays. If the time exceeds the voice buffer time, it may cause packet loss.

[0063] Due to the influence of the encoder sampling rate, high-frequency information will be lost when switching from a high sampling rate encoder to a low sampling rate encoder. If high-frequency information is present when switching is performed, the switching effect is abrupt and not smooth, which affects the user experience.

[0064] To address the aforementioned problems in current voice data transmission, this application provides various possible implementation methods to achieve dynamic and smooth codec switching in voice communication. These are explained below with reference to the accompanying drawings and several examples.

[0065] first, Figure 1 This application provides a schematic diagram of a voice data transmission system structure according to an embodiment of the present application, as shown below. Figure 1 As shown, the voice data transmission system includes a transmitting device 100 and a receiving device 300. The transmitting device 100 and the receiving device 300 are connected wirelessly, for example, the transmitting device 100 and the receiving device 300 can achieve wireless communication connection through radio waves, electromagnetic waves, etc. This application does not limit this.

[0066] It should be noted that the transmitting device 100, as the sender of voice data, has functions such as audio acquisition, audio encoding, and transmission of encoded audio. The transmitting device 100 can be implemented by an electronic device running the voice data transmission method of this application; such electronic device can be, for example, a terminal device or a server.

[0067] It should also be noted that the receiving device 300, as the receiver of voice data, receives the voice data sent by the sending device 100 and decodes it for use. The receiving device 300 can be implemented by an electronic device running the receiving voice data transmission method of this application, such as a terminal device or a server.

[0068] The above is merely an illustrative example. In actual implementation, the voice data transmission system may include multiple receiving devices 300 or multiple transmitting devices 100. There may also be other structural forms between the transmitting devices 100 and the receiving devices 300, which are not limited in this application.

[0069] Figure 2 A flowchart illustrating a voice data transmission method provided in an embodiment of this application is shown below. Figure 2 As shown, the method includes:

[0070] Step 110: Determine the first target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice characteristics of the voice data to be transmitted; the first encoder and the second encoder correspond to different bit rate ranges.

[0071] It should be noted that the first encoder and the second encoder support different dynamic bit rate ranges and have different suitable bandwidths.

[0072] It should also be noted that the uplink bandwidth information in this application is reported in real time by the physical layer of the sender (i.e., the sending device 100). In one specific implementation, the physical layer of the sender can adjust and smooth the reported uplink bandwidth value according to the channel quality fluctuations, for example, by reducing the reported value when the fluctuations are too large.

[0073] Furthermore, the speech features of the speech data to be transmitted are formal features that can distinguish speech data, such as frequency features and sound intensity features. The above is only an illustrative example; in actual implementation, speech features can take other forms, and this application does not limit them.

[0074] In one possible implementation, this application first determines a first target encoder based on the uplink bandwidth and the speech characteristics of the speech data to be transmitted in the first encoder and the second encoder. This first target encoder is the encoder among the first and second encoders that can adapt to the current uplink bandwidth and speech characteristics. Furthermore, since the first target encoder is directly used to encode the speech data after its determination, the conditions for determining the first target encoder can also be used to determine the specific switching timing; that is, the first target encoder can also determine a suitable switching timing based on speech characteristics (e.g., switching when the volume is lowest or when there is the least useful information in the sound). The above is merely an illustrative example; in actual implementation, the determination and switching of the first target encoder can take other forms, which this application does not limit.

[0075] It should also be noted that the first target encoder determined in step 110 can be the original encoder before step 110 was executed, or it can be a newly confirmed encoder. This application does not limit this.

[0076] Furthermore, although this application uses the first encoder and the second encoder as examples in the embodiments, the solution of this application can be extended to more application scenarios (more than 2 encoders), and this application does not limit the specific number of encoders.

[0077] Step 120: Encode the voice data to be sent using the first target encoder to obtain the first voice encoded data.

[0078] Step 130: Send a first voice data packet to the receiver. The first voice data packet includes: first decoding instruction information and first voice encoding data. The first decoding instruction information is used to instruct the receiver to decode the first voice encoding data using a first target decoder to obtain the voice data to be sent.

[0079] After determining the first target encoder, the voice data to be transmitted is encoded using the first target encoder to obtain the first voice encoded data.

[0080] In one possible implementation, in order for the receiving device 300 to smoothly switch the corresponding decoder according to the first target encoder switched by the transmitting device 100, the first voice data packet sent by the transmitting device 100 includes: first decoding indication information and first voice encoding data, wherein the first decoding indication information indicates the first target decoder used by the receiver.

[0081] In one specific implementation, the first decoding indication information may be the tagging information of the first target decoder corresponding to the first target encoder (i.e., the first decoding indication information can specify the first target decoder used by the receiving device 300). Alternatively, the first decoding indication information may be a parameter related to the first target encoder, which allows the receiving device 300 to select a suitable first target decoder. The above is merely an illustrative example; in actual implementations, other forms of the first decoding indication information may exist, and this application does not limit this. Furthermore, the first decoding indication information in this application may be sent in the form of packet header tags, and this application does not limit this as well.

[0082] In another specific implementation, for more efficient data transmission between the transmitting device 100 and the receiving device 300, the first decoding indication information can be added when switching encoders. That is, when the encoder used for the current voice encoded data is the same as the encoder used for the previously transmitted voice encoded data, the first voice data packet may only include the first voice encoded data; when the encoder used for the current voice encoded data is different from the encoder used for the previously transmitted voice encoded data, the first voice data packet may include both the first voice encoded data and the first decoding indication information. Thus, when the first voice data packet received by the receiving device 300 does not include the first decoding indication information, the receiving device 300 can continue to use the currently used decoder for decoding; only when the voice data packet includes the first decoding indication information is it necessary to switch decoders.

[0083] The above is merely an example; in actual implementation, there may be other forms of implementation, which this application does not limit.

[0084] Optional, continue to refer to Figure 2 This application also provides a possible implementation of a voice data transmission method, applied to a receiving device 300, the method comprising:

[0085] Step 310: Obtain the first voice data packet, which includes: first decoding instruction information and first voice encoding data. The first voice encoding data is obtained by the sender encoding the voice data to be sent using a first target encoder.

[0086] Step 320: According to the first decoding instruction information, the first target decoder is used to decode the first voice encoded data to obtain the voice data to be sent.

[0087] In one possible implementation, after the receiver (i.e., receiving device 300) receives the first voice data packet sent by the sender (sending device 100), it uses the first target decoder indicated by the first decoding indication information (or selects the first target decoder corresponding to the first target encoder according to the parameter information of the first target encoder contained in the first decoding indication information) to decode the first voice encoded data, thereby obtaining the voice data to be sent.

[0088] Optionally, based on the above embodiments, before using the first target decoder to decode the first voice encoded data according to the first decoding instruction information, the original decoder can also be used to decode a preset number of voice data; and the original decoder is released and the system switches to the first target decoder.

[0089] In one specific implementation, when the receiver detects a change in the first decoding indication information of the first voice data packet compared to the previously received first voice data packet (e.g., a change in the packet header marker), or (where the sender only sends the first decoding indication information when switching encoders), upon receiving the first decoding indication information of the first voice data packet, the receiver uses the original decoder to decode a preset number of voice data packets (e.g., using the original decoder to generate one or more PLC compensation packets), then releases the original decoder and switches to the first target decoder. Thus, if switching to the first target decoder causes a relatively long delay, this delay can be filled with pre-generated voice data (or compensation packets), thereby minimizing the impact of decoder switching on the user.

[0090] In summary, this application provides a voice data transmission method. Based on the uplink bandwidth and the voice characteristics of the voice data to be transmitted, a first target encoder is determined from a first encoder and a second encoder. The first target encoder is used to encode the voice data to be transmitted to obtain first voice encoded data. A first voice data packet, including first decoding instruction information and the first voice encoded data, is then sent to the receiver. This application provides a voice transmission method in a wireless voice communication scenario where the physical layer bandwidth of wireless communication dynamically adjusts with the channel environment, achieving smooth switching of the voice codec. According to the scheme of this application, the first target encoder is determined by combining voice characteristics and the uplink bandwidth of the physical layer during switching, and the switching time of the voice codec is selected based on this, thereby reducing the impact of switching on the user experience. Furthermore, the scheme of this application can select the voice encoder suitable for the optimal effect of the bandwidth when the physical layer is in different uplink bandwidth states, and the switching method does not require modification of the internal implementation of the codec, nor does it impose limitations on the frame length, implementation algorithm, sampling rate, etc. of the encoder.

[0091] Optionally, in the above Figure 2 Based on this, if the bit rate of the first encoder is less than the bit rate of the second encoder, the bandwidth ranges corresponding to the first encoder and the second encoder overlap; step 110, determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted, includes:

[0092] If the uplink bandwidth is within the first bandwidth range, and the original encoder is the second encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the first bandwidth range is: greater than the minimum bandwidth of the bandwidth range corresponding to the second encoder, and less than or equal to the first preset low bandwidth threshold.

[0093] For the convenience of illustration in subsequent embodiments, the present application sets a possible first implementation scenario here: In the first implementation scenario, there are two encoders, the first encoder and the second encoder; among them, the bandwidth range suitable for the first encoder is s1 - s5, and the bandwidth range suitable for the second encoder is s2 - s6 (where the bandwidth ranges are s1 < s2 < s3 < s4 < s5 < s6), and two threshold points are preset: s3 (the first preset low - bandwidth threshold), s4 (the first preset high - bandwidth threshold). It should be noted that the above - mentioned implementation scenario is only a possible implementation scenario provided for the convenience of subsequent examples. In the actual implementation of the present application, there may be other usage scenarios, and the present application does not limit this. It should also be noted that for the purpose of illustrating the method of the present application, the present application divides the bandwidth range into multiple intervals (i.e., s1, s2... s6), and each interval includes a certain bandwidth range. Regarding the specific interval division method and the specific bandwidth range corresponding to each interval, the user can set it according to the actual usage situation, and the present application does not limit this.

[0094] First of all, it should be noted that the selection of the threshold points in the present application includes but is not limited to the following methods:

[0095] When there is an overlap in the bandwidth ranges corresponding to the first encoder and the second encoder, the first preset low - bandwidth threshold and the first preset high - bandwidth threshold can be selected within the overlapping part of the bandwidth, or the boundaries of the range where the sound qualities of the two encoders are not significantly different can be selected as the two threshold points (the first preset low - bandwidth threshold and the first preset high - bandwidth threshold respectively), or the user can select appropriate threshold points according to the specific usage scenario and can adjust the threshold points during use. The above is only an example. In the actual implementation, there may be other threshold - selection methods, and the present application does not limit this.

[0096] It should also be noted that the code rate of the first encoder is less than that of the second encoder, and there is an overlap in the bandwidth ranges corresponding to the first encoder and the second encoder; where the code rate of the first encoder being less than that of the second encoder means that outside the overlapping code - rate range of the first encoder and the second encoder, the non - overlapping code - rate range of the first encoder is less than the non - overlapping code - rate range of the second encoder. In this case, the bandwidth range of the first encoder is less than that of the second encoder. For example, if the overlapping bandwidth range corresponding to the first encoder and the second encoder is s3 - s4, and if the non - overlapping bandwidth range of the first encoder is s2 - s3 and the non - overlapping bandwidth range of the second encoder is s4 - s5, then it can be considered that the bandwidth of the first encoder is less than the code - rate bandwidth of the second encoder.

[0097] In one possible implementation, the first bandwidth range is: greater than the minimum bandwidth of the bandwidth range corresponding to the second encoder, and less than or equal to a first preset low bandwidth threshold. For example, in the first implementation scenario of this application, the first bandwidth range is s2-s3.

[0098] In one specific implementation, in the first implementation scenario of this application, the first bandwidth range is s2-s3. Within this range, if the original encoder is the second encoder, the first target encoder can be determined from the first and second encoders based on the speech features. It should be noted that the first target encoder determined based on the speech features can be either the first encoder or the second encoder, and this application does not limit this.

[0099] Optionally, based on the above embodiments, this application also provides a possible implementation of the voice data transmission method, wherein if the uplink bandwidth is less than a first bandwidth range, a first encoder is used. Alternatively, if the uplink bandwidth is less than the first bandwidth range, the method switches to the first encoder.

[0100] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, wherein the voice features may include: silence state indication information, or volume indication information, and the above step: determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0101] If the duration of the original encoder being the second encoder exceeds the first preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder.

[0102] Alternatively, if the duration of the original encoder being the second encoder exceeds the second preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the first encoder is determined to be the first target encoder.

[0103] It should be noted that the speech features can be detected by the application layer of the sender. These speech features may include silence status indication information for each speech frame (e.g., Voice Activity Detection (VAD) silence status), volume indication information (e.g., volume level), frequency range, or energy (e.g., high-frequency energy), etc. This application does not limit the specific types of information included in the speech features.

[0104] In one possible implementation, if the duration for which the original encoder is the second encoder exceeds a first preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder. For example, in a first implementation scenario, if the original encoder is the second encoder, the time for using the second encoder exceeds the first preset duration (T1), and the current voice frame is a silent frame, then the first encoder is determined to be the first target encoder, and the system switches to the first target encoder. It should be noted that, based on the above embodiments, a judgment on the silence status can be added. For example, it can be set that if the duration for which the original encoder is the second encoder exceeds the first preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state and the duration of the silent state is greater than a preset silence duration threshold, then the first encoder is determined to be the first target encoder. Here, the duration for which the original encoder is the second encoder refers to the cumulative usage (or running) time of the current second encoder within the first bandwidth range, that is, the time length from the most recent start encoding time within the first bandwidth range to the current time.

[0105] In another possible implementation, if the duration of the original encoder being the second encoder exceeds a second preset duration, and the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, then the first encoder is determined to be the first target encoder. For example, in the first implementation scenario, if the original encoder is the second encoder, and the time spent using the second encoder exceeds the second preset duration (T2), and the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, then the first encoder is determined to be the first target encoder. Based on this, if the sampling rates of the first encoder and the second encoder are different, and if the duration of the second encoder exceeds the second preset duration, the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, and the high-frequency energy is below a preset energy threshold, then the first encoder is determined to be the first target encoder. It should be noted that the second preset duration is longer than the first preset duration.

[0106] Optionally, based on the above embodiments, this application also provides a possible implementation of the voice data transmission method. Figure 3 A flowchart illustrating a voice data transmission method provided in another embodiment of this application; as shown Figure 3 As shown, the method also includes:

[0107] Step 130: If the uplink bandwidth is within the first bandwidth range, and the duration of the original encoder being the second encoder exceeds the third preset duration, then the first encoder is determined to be the second target encoder.

[0108] Step 140: Encode the voice data to be transmitted using a second target encoder to obtain second voice encoded data;

[0109] Step 150: Send a second voice data packet to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to use a second target decoder to decode the second voice encoding data to obtain the voice data to be sent.

[0110] In one possible implementation, the uplink bandwidth is within a first bandwidth range, and if the duration for which the original encoder is the second encoder exceeds a third preset duration, then the first encoder is determined to be the second target encoder. For example, in a first implementation scenario, if the original encoder is the second encoder and the time spent using the second encoder exceeds the third preset duration (T3), then the first encoder is determined to be the second target encoder. It should be noted that the third preset duration is longer than the second preset duration.

[0111] After determining the second target encoder, the voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data, and then sent to the receiver as a second voice data packet. This second voice data packet includes: second decoding instruction information and the second voice encoded data. The second decoding instruction information instructs the receiver to decode the second voice encoded data using the second target decoder to obtain the voice data to be transmitted. The specific implementation method is the same as the implementation method for sending the first voice packet data and determining the first target decoder described above (see the description of steps 120 and 130), and will not be repeated here.

[0112] After the receiver receives the second voice packet data, the method includes:

[0113] Step 330: Obtain the second voice data packet; the second voice data packet includes: second decoding instruction information and second voice encoding data; wherein, the second voice encoding data is obtained by the sender encoding the voice data to be sent using a second target encoder.

[0114] Step 340: According to the second decoding instruction information, the second target decoder is used to decode the second speech encoded data to obtain the speech data to be sent.

[0115] The specific processing methods for the second voice data packet in steps 330 and 340 of the receiving device 300 are the same as the specific processing methods for the first voice data packet in steps 310 and 320, and will not be repeated here.

[0116] In summary, in this application, when the uplink bandwidth of the sender's physical layer is within a certain range, the switching of speech coding is triggered by the silence state of the speech VAD; if the silence state of the VAD is not detected for a certain period of time, the switching of speech coding is triggered when the volume and high-frequency energy of the speech are lower than a certain threshold; if none of the above switching conditions are met, the switching is performed directly when the waiting time exceeds the maximum duration.

[0117] Optionally, in the above Figure 2 Based on this, this application also provides a possible implementation of a voice data transmission method. If the uplink bandwidth is greater than a first preset low bandwidth threshold and less than or equal to a first preset high bandwidth threshold, since both encoders are applicable within this range, the original encoder can remain unchanged.

[0118] Optionally, in the above Figure 2 Based on this, this application also provides a possible implementation of a voice data transmission method, wherein determining the target encoder from the first encoder and the second encoder according to the uplink bandwidth and the voice features of the voice data to be transmitted includes:

[0119] If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a first preset high bandwidth threshold, and less than or equal to the maximum bandwidth of the bandwidth range corresponding to the first encoder.

[0120] In one possible implementation, the second bandwidth range is: greater than the first preset high bandwidth threshold, and less than or equal to the maximum bandwidth of the bandwidth range corresponding to the first encoder. For example, in the first implementation scenario of this application, the second bandwidth range is s4-s5.

[0121] In one specific implementation, in the first implementation scenario of this application, the second bandwidth range is s4-s5. Within this range, if the original encoder is the first encoder, the first target encoder can be determined from the first and second encoders based on the speech features. It should be noted that the first target encoder determined based on the speech features can be either the first encoder or the second encoder, and this application does not limit this.

[0122] Optionally, if the uplink bandwidth is greater than the second bandwidth range, the second encoder is used. Alternatively, if the uplink bandwidth is greater than the second bandwidth range, the system switches to the second encoder.

[0123] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, wherein the voice features include: silence state indication information, or volume indication information, and the step of determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0124] If the duration of the original encoder being the first encoder exceeds the fourth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder.

[0125] Alternatively, if the duration of the original encoder being the first encoder exceeds a fifth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the second encoder is determined to be the first target encoder.

[0126] In one possible implementation, if the duration of the original encoder being the first encoder exceeds a fourth preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder. For example, in the first implementation scenario, if the original encoder is the first encoder, the time of using the first encoder exceeds the fourth preset duration (T4), and the current voice frame is a silent frame, then the second encoder is determined to be the first target encoder, and the system switches to the first target encoder. It should be noted that, based on the above embodiments, judgments on silence status, etc., can be added, and this application does not limit this. The duration of the original encoder being the first encoder refers to the cumulative usage (or running) time of the current first encoder within the second bandwidth range, that is, the time length from the most recent start encoding time within the second bandwidth range to the current time.

[0127] In another possible implementation, if the duration of the original encoder (the first encoder) exceeds a fifth preset duration, and the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, then the second encoder is determined to be the first target encoder. For example, in the first implementation scenario, if the original encoder is the first encoder, the time of using the first encoder exceeds a fifth preset duration (T5), and the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, then the second encoder is determined to be the first target encoder. Based on this, if the sampling rates of the first encoder and the second encoder are different, if the duration of the first encoder exceeds a fifth preset duration, the volume indication information of the current speech frame indicates that the volume is below a preset volume threshold, and the high-frequency energy is below a preset energy threshold, then the second encoder is determined to be the first target encoder. It should be noted that the fifth preset duration is longer than the fourth preset duration.

[0128] Optionally, based on the above embodiments, this application also provides a possible implementation of the voice data transmission method. Figure 4 A flowchart illustrating a voice data transmission method provided in another embodiment of this application; as shown Figure 4 As shown, the method also includes:

[0129] Step 160: If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the sixth preset duration, then the second encoder is determined to be the second target encoder.

[0130] Step 170: Encode the voice data to be transmitted using a second target encoder to obtain second voice encoded data;

[0131] Step 180: Send a second voice data packet to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to use a second target decoder to decode the second voice encoding data to obtain the voice data to be sent.

[0132] In one possible implementation, the uplink bandwidth is within the second bandwidth range, and if the duration of the original encoder being the first encoder exceeds a sixth preset duration, then the second encoder is determined to be the second target encoder. For example, in a first implementation scenario, if the original encoder is the first encoder and the time spent using the first encoder exceeds a sixth preset duration (T6), then the second encoder is determined to be the second target encoder. It should be noted that the sixth preset duration is longer than the fifth preset duration.

[0133] After determining the second target encoder, the voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data, and then sent to the receiver as a second voice data packet. This second voice data packet includes: second decoding instruction information and the second voice encoded data. The second decoding instruction information instructs the receiver to decode the second voice encoded data using the second target decoder to obtain the voice data to be transmitted. The specific implementation method is the same as the implementation method for sending the first voice packet data and determining the first target decoder described above (see the description of steps 120 and 130), and will not be repeated here.

[0134] After the receiver receives the second voice packet data, the specific implementation of the receiving device 300 is the same as that of steps 330 and 340, and will not be described again in this application.

[0135] Optionally, in the above Figure 2Based on the above, the present application further provides a possible implementation manner of a voice data transmission method. If the code rate of the first encoder is less than the code rate of the second encoder, there is no overlap in the bandwidth ranges corresponding to the first encoder and the second encoder; determining a target encoder from the first encoder and the second encoder according to the uplink bandwidth and the voice characteristics of the voice data to be transmitted includes:

[0136] If the uplink bandwidth is within the first bandwidth range and the original encoder is the second encoder, then determine a first target encoder from the first encoder and the second encoder according to the voice characteristics, where the first bandwidth range is: greater than the minimum bandwidth corresponding to the second encoder and less than or equal to the second preset low bandwidth threshold.

[0137] For the convenience of giving examples in subsequent embodiments, the present application sets a possible second implementation scenario here: there are two encoders in the second implementation scenario, a first encoder and a second encoder; among them, the bandwidth range suitable for the first encoder is s1 - s2, and the bandwidth range suitable for the second encoder is s3 - s6 (where the bandwidth ranges s1 < s2 < s3 < s4 < s5 < s6), and two threshold points are preset: s4 (the second preset low bandwidth threshold), s5 (the second preset high bandwidth threshold). It should be noted that the above implementation scenario is only a possible implementation scenario provided for the convenience of subsequent examples, and there may be other usage scenarios in the actual implementation of the present application, and the present application does not limit this. It should also be noted that for the purpose of illustrating the method of the present application, the present application divides the bandwidth range into multiple intervals (i.e., s1, s2... s6), and each interval includes a certain bandwidth range. For the specific interval division method and the specific bandwidth range corresponding to each interval, the user can set according to the actual usage situation, and the present application does not limit this.

[0138] First of all, it should be noted that for the selection of the threshold point in the present application, the user can select an appropriate threshold point according to the specific usage scenario and can adjust the threshold point during use, and the present application does not limit this.

[0139] In a possible implementation manner, the first bandwidth range is: greater than the minimum bandwidth corresponding to the second encoder and less than or equal to the second preset low bandwidth threshold. For example, in the second implementation scenario of the present application, the first bandwidth range is s3 - s4.

[0140] In one specific implementation, in the second implementation scenario of this application, the first bandwidth range is s3-s4. Within this range, if the original encoder is the second encoder, the first target encoder can be determined from the first encoder and the second encoder based on the speech features. It should be noted that the first target encoder determined based on the speech features can be either the first encoder or the second encoder, and this application does not limit this.

[0141] Optionally, based on the above embodiments, this application also provides a possible implementation of the voice data transmission method, wherein if the uplink bandwidth is less than a first bandwidth range, a first encoder is used. Alternatively, if the uplink bandwidth is less than the first bandwidth range, the method switches to the first encoder.

[0142] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, wherein the voice features include: silence state indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0143] If the duration of the original encoder being the second encoder exceeds the seventh preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder.

[0144] Alternatively, if the duration of the original encoder being the second encoder exceeds an eighth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the first encoder is determined to be the first target encoder.

[0145] In one possible implementation, if the duration of the original encoder being the second encoder exceeds a seventh preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder. For example, in a second implementation scenario, if the original encoder is the second encoder, and the time spent using the second encoder exceeds a seventh preset duration (T7), and the current voice frame is a silent frame, then the first encoder is determined to be the first target encoder, and the system switches to the first target encoder. It should be noted that, based on the above embodiments, a judgment on the silence status can be added. For example, it can be set that if the duration of the original encoder being the second encoder exceeds a seventh preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state and the duration of the silent state is greater than a preset silence duration threshold, then the first encoder is determined to be the first target encoder.

[0146] In another possible implementation, if the duration of the second encoder (the original encoder) exceeds an eighth preset duration, and the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, then the first encoder is determined to be the first target encoder. For example, in a second implementation scenario, if the original encoder is the second encoder, and the time spent using the second encoder exceeds an eighth preset duration (T8), and the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, then the first encoder is determined to be the first target encoder. Furthermore, if the sampling rates of the first encoder and the second encoder are different, and if the duration of the second encoder exceeds an eighth preset duration, the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, and the high-frequency energy is below a preset energy threshold, then the first encoder is determined to be the first target encoder. It should be noted that the eighth preset duration is longer than the seventh preset duration.

[0147] It should be noted that if the encoder's bit rate is greater than the actual transmission bandwidth, it will cause audio stuttering. Therefore, even if the actual bandwidth may reach s3 in the second implementation scenario, the first encoder is chosen as the first target encoder because the bit rate range of the second encoder actually exceeds s3.

[0148] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, which further includes:

[0149] If the uplink bandwidth is within the first bandwidth range, and the duration for which the original encoder is the second encoder exceeds the ninth preset duration, then the first encoder is determined to be the second target encoder.

[0150] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0151] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0152] In one possible implementation, the uplink bandwidth is within a first bandwidth range, and if the duration for which the original encoder is the second encoder exceeds a ninth preset duration, then the first encoder is determined to be the second target encoder. For example, in a first implementation scenario, if the original encoder is the second encoder and the time spent using the second encoder exceeds a ninth preset duration (T9), then the first encoder is determined to be the second target encoder. It should be noted that the ninth preset duration is longer than the eighth preset duration.

[0153] After determining the second target encoder, the voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data, and then sent to the receiver as a second voice data packet. This second voice data packet includes: second decoding instruction information and the second voice encoded data. The second decoding instruction information instructs the receiver to decode the second voice encoded data using the second target decoder to obtain the voice data to be transmitted. The specific implementation method is the same as the implementation method for sending the first voice packet data and determining the first target decoder described above (see the description of steps 120 and 130), and will not be repeated here.

[0154] After the receiver receives the second voice packet data, the specific implementation of the receiving device 300 is the same as that of steps 330 and 340, and will not be described again in this application.

[0155] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, wherein determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes:

[0156] If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a second preset low bandwidth threshold, and less than or equal to a second preset high bandwidth threshold.

[0157] In one possible implementation, the second bandwidth range is: greater than a second preset low bandwidth threshold, and less than or equal to a second preset high bandwidth threshold. For example, in the first implementation scenario of this application, the second bandwidth range is s4-s5.

[0158] In one specific implementation, in the first implementation scenario of this application, the second bandwidth range is s4-s5. Within this range, if the original encoder is the first encoder, the first target encoder can be determined from the first and second encoders based on the speech features. It should be noted that the first target encoder determined based on the speech features can be either the first encoder or the second encoder, and this application does not limit this.

[0159] Optionally, if the uplink bandwidth is greater than the second bandwidth range, the second encoder is used. Alternatively, if the uplink bandwidth is greater than the second bandwidth range, the system switches to the second encoder.

[0160] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, wherein the voice features include: silence state indication information, or volume indication information, and determining the target encoder from the first encoder and the second encoder based on the voice features includes:

[0161] If the duration of the original encoder being the first encoder exceeds the tenth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder.

[0162] Alternatively, if the duration of the original encoder being the first encoder exceeds the eleventh preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the second encoder is determined to be the first target encoder.

[0163] In one possible implementation, if the duration of the original encoder being the first encoder exceeds a tenth preset duration, and the silence status indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder. For example, in a second implementation scenario, if the original encoder is the first encoder, the time of using the first encoder exceeds a tenth preset duration (T10), and the current voice frame is a silent frame, then the second encoder is determined to be the first target encoder, and the system switches to the first target encoder. It should be noted that, based on the above embodiments, further determinations regarding silence status can be added, and this application does not limit this.

[0164] In another possible implementation, if the duration of the original encoder (the first encoder) exceeds an eleventh preset duration, and the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, then the second encoder is determined to be the first target encoder. For example, in a second implementation scenario, if the original encoder is the first encoder, and the time spent using the first encoder exceeds an eleventh preset duration (T11), and the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, then the second encoder is determined to be the first target encoder. Based on this, if the sampling rates of the first encoder and the second encoder are different, and if the duration of the first encoder exceeds an eleventh preset duration, the volume indicator information of the current speech frame indicates that the volume is below a preset volume threshold, and the high-frequency energy is below a preset energy threshold, then the second encoder is determined to be the first target encoder. It should be noted that the eleventh preset duration is longer than the tenth preset duration.

[0165] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, which further includes:

[0166] If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the twelfth preset duration, then the second encoder is determined to be the second target encoder.

[0167] The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data;

[0168] A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoding data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoding data using the second target decoder to obtain the voice data to be sent.

[0169] In one possible implementation, the uplink bandwidth is within the second bandwidth range, and if the duration of the original encoder being the first encoder exceeds the twelfth preset duration, then the second encoder is determined to be the second target encoder. For example, in a second implementation scenario, if the original encoder is the first encoder and the time spent using the first encoder exceeds the twelfth preset duration (T12), then the second encoder is determined to be the second target encoder. It should be noted that the twelfth preset duration is longer than the eleventh preset duration.

[0170] After determining the second target encoder, the voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data, and then sent to the receiver as a second voice data packet. This second voice data packet includes: second decoding instruction information and the second voice encoded data. The second decoding instruction information instructs the receiver to decode the second voice encoded data using the second target decoder to obtain the voice data to be transmitted. The specific implementation method is the same as the implementation method for sending the first voice packet data and determining the first target decoder described above (see the description of steps 120 and 130), and will not be repeated here.

[0171] After the receiver receives the second voice packet data, the specific implementation of the receiving device 300 is the same as that of steps 330 and 340, and will not be described again in this application.

[0172] Optionally, according to the Harry Nyquist sampling theory, the highest sound frequency of PCM data with a 16K sampling rate can reach 8KHz, while the highest sound frequency of PCM data with an 8K sampling rate can only reach 4KHz. The sampling rate of voice data also affects the transmission quality. Therefore, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, which further includes: if the sampling rate of the original encoder is different from the sampling rate of the first target encoder, before step 120, which uses the first target encoder to encode the voice data to be transmitted to obtain the first voice encoded data, the method further includes:

[0173] Low-pass filtering is applied to the voice data to be transmitted;

[0174] When the sampling rate of the filtered voice data to be transmitted is the same as the sampling rate of the first target decoder, the first target decoder is used to encode the voice data to be transmitted.

[0175] It should be noted that in this application, when performing low-pass filtering on the voice data to be transmitted, multiple low-pass filters can be used to process the voice data to be transmitted, so that its sampling rate decreases step by step. This application does not limit the specific setting method of the filters.

[0176] In one specific implementation, in the first implementation scenario, such as the sampling rate of the first encoder being 4KHz and the sampling rate of the second encoder being 8Hz, multiple low-pass filters can be added (for example, multiple low-pass filters with filter thresholds ranging from 4KHz to 8KHz).

[0177] When the transmitter bandwidth is within the first bandwidth range, if the encoder is the second encoder, the sound is low-pass filtered. The filtering range is lowered by one level every thirteen preset durations (T13) until the sampling rate (4KHz) corresponding to the first encoder is reached, after which the encoder switches to the first encoder.

[0178] When the transmitter bandwidth is within the second bandwidth range, if the encoder is the second encoder, low-pass filtering is applied to the sound. The filtering range is increased by one level every fourteen preset durations (T14) until the sampling rate (8Hz) corresponding to the second encoder is reached. Then, the low-pass filtering is canceled and the encoder is switched to the second encoder.

[0179] In another specific implementation, in the second implementation scenario, such as the sampling rate of the first encoder being 4KHz and the sampling rate of the second encoder being 8Hz, multiple low-pass filters can be added (for example, multiple low-pass filters with filter thresholds ranging from 4KHz to 8KHz).

[0180] When the transmitter bandwidth is in the first bandwidth range (s3-s4), if the encoder is the second encoder, the sound is low-pass filtered. The filtering range is lowered by one level every fifteen preset durations (T15) until the sampling rate (4KHz) corresponding to the first encoder is reached, and then the signal is switched to the first encoder.

[0181] When the transmitter bandwidth is in the second bandwidth range (s4-s5), if the encoder is the second encoder, the sound is low-pass filtered. The filtering range is increased by one level every sixteen preset durations (T16) until the sampling rate (8Hz) corresponding to the second encoder is reached. Then the low-pass filter is canceled and the encoder is switched to the second encoder.

[0182] Optionally, based on the above embodiments, this application also provides a possible implementation of a voice data transmission method, which further includes:

[0183] When the uplink bandwidth of the sender decreases, the encoder is down-adjusted (switched from the second encoder to the first encoder), the time threshold is set to be short (e.g., the first preset duration, the second preset duration, the third preset duration), and various thresholds (energy threshold, etc.) are wide, thereby speeding up the down-adjustment speed;

[0184] When the uplink bandwidth of the transmitter increases, the encoder is adjusted upwards (e.g., switching from the first encoder to the second encoder), and the time threshold is set to be longer (e.g., the fourth preset time, the fifth preset time, the sixth preset time), and various thresholds are strict (energy threshold, etc.), thereby slowing down the adjustment speed; in addition, the encoder adjustment and down adjustment areas in this application do not overlap, thereby avoiding ping-pong switching.

[0185] The following describes the transmitting device, receiving device, and computer-readable storage medium used to execute the present application. The specific implementation process and technical effects are described above and will not be repeated below.

[0186] This application provides a possible implementation example of a transmitting device capable of executing the voice data transmission method provided in the above embodiments. Figure 5 This is a schematic diagram of a transmitting device provided in an embodiment of this application. The transmitting device can be integrated into a terminal device or a chip of a terminal device. The terminal can be a computing device with data processing capabilities.

[0187] The transmitting device includes a processor 501, a storage medium 502, and a bus. The storage medium stores program instructions executable by the processor. When the transmitting device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the aforementioned voice data transmission method. The specific implementation and technical effects are similar and will not be described in detail here.

[0188] This application provides a possible implementation example of a receiving device capable of executing the voice data transmission method provided in the above embodiments. Figure 6 This is a schematic diagram of a receiving device provided in an embodiment of this application. The receiving device can be integrated into a terminal device or a chip of a terminal device. The terminal can be a computing device with data processing capabilities.

[0189] The receiving device includes a processor 601, a storage medium 602, and a bus. The storage medium stores program instructions executable by the processor. When the receiving device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the aforementioned voice data transmission method for the sender. The specific implementation and technical effects are similar and will not be described in detail here.

[0190] This application provides a possible implementation example of a computer-readable storage medium capable of executing the voice data transmission method provided in the above embodiments. The storage medium stores a computer program, which, when run by a processor, executes the steps of the above-described voice data transmission method for the sender and / or receiver.

[0191] A computer program stored in a storage medium may include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0192] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0193] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0194] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0195] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0196] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice data transmission method, characterized in that, include: The first target encoder is determined from the first encoder and the second encoder based on the uplink bandwidth and the voice characteristics of the voice data to be transmitted. The first encoder and the second encoder correspond to different bit rate ranges; the speech features include: frequency features or sound intensity features; The voice data to be transmitted is encoded using the first target encoder to obtain first voice encoded data; Send a first voice data packet to the receiver. The first voice data packet includes: first decoding indication information and first voice encoding data. The first decoding indication information is used to instruct the receiver to decode the first voice encoding data using a first target decoder to obtain the voice data to be sent. Before the receiver decodes the first speech encoded data using the first target decoder, the process includes: The receiver uses the original decoder to decode a preset number of voice data, then releases the original decoder and switches to the first target decoder.

2. The method as described in claim 1, characterized in that, If the bit rate of the first encoder is less than the bit rate of the second encoder, the bandwidth ranges corresponding to the first encoder and the second encoder overlap. The step of determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes: If the uplink bandwidth is within the first bandwidth range, and the original encoder is the second encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the first bandwidth range is: greater than the minimum bandwidth of the bandwidth range corresponding to the second encoder, and less than or equal to the first preset low bandwidth threshold.

3. The method as described in claim 2, characterized in that, The voice features further include: mute status indication information, or volume indication information. Determining the target encoder from the first encoder and the second encoder based on the voice features includes: If the duration of the original encoder being the second encoder exceeds the first preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder. Alternatively, if the duration of the original encoder being the second encoder exceeds the second preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the first encoder is determined to be the first target encoder.

4. The method as described in claim 2, characterized in that, The method further includes: If the uplink bandwidth is within the first bandwidth range, and the duration during which the original encoder is the second encoder exceeds the third preset duration, then the first encoder is determined to be the second target encoder. The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data; A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoded data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoded data using a second target decoder to obtain the voice data to be sent.

5. The method as described in claim 2, characterized in that, The step of determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes: If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a first preset high bandwidth threshold, and less than or equal to the maximum bandwidth of the bandwidth range corresponding to the first encoder.

6. The method as described in claim 5, characterized in that, The voice features include: mute status indication information, or volume indication information. Determining the target encoder from the first encoder and the second encoder based on the voice features includes: If the duration of the original encoder being the first encoder exceeds the fourth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder. Alternatively, if the duration of the original encoder being the first encoder exceeds a fifth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the second encoder is determined to be the first target encoder.

7. The method as described in claim 5, characterized in that, The method further includes: If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the sixth preset duration, then the second encoder is determined to be the second target encoder. The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data; A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoded data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoded data using a second target decoder to obtain the voice data to be sent.

8. The method as described in claim 1, characterized in that, If the bit rate of the first encoder is less than the bit rate of the second encoder, the bandwidth ranges corresponding to the first encoder and the second encoder do not overlap; The step of determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes: If the uplink bandwidth is within the first bandwidth range, and the original encoder is the second encoder, then based on the speech features, a first target encoder is determined from the first encoder and the second encoder, wherein the first bandwidth range is: the minimum bandwidth greater than the bandwidth range corresponding to the second encoder, and less than or equal to the second preset low bandwidth threshold.

9. The method as described in claim 8, characterized in that, The voice features include: mute status indication information, or volume indication information. Determining the target encoder from the first encoder and the second encoder based on the voice features includes: If the duration of the original encoder being the second encoder exceeds the seventh preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the first encoder is determined to be the first target encoder. Alternatively, if the duration of the original encoder being the second encoder exceeds an eighth preset duration, and the volume indication information of the current voice frame indicates that the volume is less than a preset volume threshold, then the first encoder is determined to be the first target encoder.

10. The method as described in claim 8, characterized in that, The method further includes: If the uplink bandwidth is within the first bandwidth range, and the duration for which the original encoder is the second encoder exceeds the ninth preset duration, then the first encoder is determined to be the second target encoder. The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data; A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoded data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoded data using a second target decoder to obtain the voice data to be sent.

11. The method as described in claim 8, characterized in that, The step of determining the target encoder from the first encoder and the second encoder based on the uplink bandwidth and the voice features of the voice data to be transmitted includes: If the uplink bandwidth is within the second bandwidth range, and the original encoder is the first encoder, then the first target encoder is determined from the first encoder and the second encoder based on the speech features, wherein the second bandwidth range is: greater than a second preset low bandwidth threshold, and less than or equal to a second preset high bandwidth threshold.

12. The method as described in claim 11, characterized in that, The voice features include: mute status indication information, or volume indication information. Determining the target encoder from the first encoder and the second encoder based on the voice features includes: If the duration of the original encoder being the first encoder exceeds the tenth preset duration, and the silence state indication information of the current voice frame indicates that it is in a silent state, then the second encoder is determined to be the first target encoder. Alternatively, if the duration of the original encoder being the first encoder exceeds the eleventh preset duration, and the volume indication information of the current voice frame indicates that the volume is less than the preset volume threshold, then the second encoder is determined to be the first target encoder.

13. The method as described in claim 11, characterized in that, The method further includes: If the uplink bandwidth is within the second bandwidth range, and the duration of the original encoder being the first encoder exceeds the twelfth preset duration, then the second encoder is determined to be the second target encoder. The voice data to be transmitted is encoded using the second target encoder to obtain second voice encoded data; A second voice data packet is sent to the receiver. The second voice data packet includes: second decoding instruction information and second voice encoded data. The second decoding instruction information is used to instruct the receiver to decode the second voice encoded data using a second target decoder to obtain the voice data to be sent.

Citation Information

Patent Citations

  • Communication terminal

    JP2006195144A