Audio transmission method and device

By dynamically adjusting the multi-level encoding strategy at both the sending and receiving ends, the audio quality problem under low bandwidth and unstable network conditions was solved, thereby improving audio quality, making reasonable use of bandwidth, and enhancing the user experience.

CN116170422BActive Publication Date: 2025-10-31SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310000961.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-03
Publication Date
2025-10-31
Estimated Expiration
2043-01-03

AI Technical Summary

Technical Problem

Under conditions of low bandwidth and unstable network, existing technologies cannot effectively guarantee audio quality, resulting in audio stuttering and poor quality.

Method used

By dynamically adjusting the multi-level encoding strategy at the sending and receiving ends, low-level, medium-level, and high-level encoding methods are adopted according to the available bandwidth and network conditions to transmit audio data in layers, ensuring the best match between audio quality and bandwidth utilization.

Benefits of technology

Improving audio quality, reducing bandwidth consumption, and enhancing the overall user experience under low bandwidth and unstable network conditions, especially when the downlink network is poor, allows more bandwidth to be used for video and data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116170422B_ABST
    Figure CN116170422B_ABST
Patent Text Reader

Abstract

This application provides an audio transmission method and apparatus. The method includes: a first electronic device assessing the available bandwidth between itself and a server, and resetting an encoding strategy based on the assessed available bandwidth; and encoding audio data based on the new encoding strategy, and outputting multiple encoded audio streams to the server. This achieves dynamic adjustment of the encoding strategy, meeting different bandwidth requirements of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and more particularly to an audio transmission method and apparatus. Background Technology

[0002] With the development of the communications field, the application scenarios of terminals are becoming increasingly widespread. For example, users can use terminals to conduct online voice or video conferences among multiple users. However, in the use of conferencing systems, scenarios often arise where uplink and downlink networks are unstable or limited. Therefore, how to ensure audio quality under low bandwidth and unstable network conditions has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides an audio transmission method and apparatus that can effectively improve audio quality.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] In a first aspect, embodiments of this application provide an audio transmission method. This method is applied to a first electronic device and includes: determining a first available bandwidth between the device and a server based on preset evaluation conditions. The determined first available bandwidth is not equal to the currently available bandwidth. Based on the first available bandwidth, a first encoding strategy is determined. The first encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method. The first encoding method includes a low-end encoding method, the second encoding method includes a low-end encoding method and a mid-end encoding method, and the third encoding method includes a low-end encoding method, a mid-end encoding method, and a high-end encoding method. The bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different. Based on the first encoding strategy, first audio data is encoded to obtain N audio bitstreams; N is a positive integer greater than 0 and less than 4; wherein a single bitstream among the N audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method. The N audio bitstreams are sent to a server, causing the server to send a first target audio bitstream among the N audio bitstreams to a second electronic device. Thus, in this embodiment of the application, the first electronic device, as the transmitting end, can adjust the encoding method based on the available bandwidth. Through multi-level encoding, it can effectively cope with low bandwidth and unstable network scenarios, so that the audio bitstream transmitted by the first electronic device has better audio effect under low bandwidth and unstable network conditions.

[0006] For example, the preset evaluation conditions include, but are not limited to, at least one of the following: packet transmission conditions and / or bandwidth variation conditions. That is, the first electronic device can determine the available bandwidth based on packet transmission conditions (e.g., packet loss rate) and / or bandwidth variation conditions (e.g., bandwidth change rate).

[0007] In one possible implementation, the low-level encoding method in the first encoding strategy corresponds to a first bitrate, which falls within a first bitrate range. If the first encoding strategy includes a mid-level encoding method, this mid-level encoding method corresponds to a second bitrate, which falls within a second bitrate range. If the first encoding strategy includes a high-level encoding method, this high-level encoding method corresponds to a third bitrate, which falls within a third bitrate range. The minimum value of the third bitrate range is greater than the maximum value of the second bitrate range, and the minimum value of the second bitrate range is greater than the maximum value of the first bitrate range. Thus, in this embodiment, by adjusting the bitrate of different encoding levels, the bandwidth and audio quality occupied by audio transmission can be further precisely controlled.

[0008] In one possible implementation, the first audio data is encoded based on a first encoding strategy to obtain N audio bitstreams, including: if the first encoding strategy is a first encoding method, encoding the first audio data based on a low-level encoding method to obtain a first audio bitstream; if the first encoding strategy is a second encoding method, encoding the first audio data based on a low-level encoding method to obtain a first audio bitstream; and encoding the first audio data based on a mid-level encoding method to obtain a second audio bitstream; if the first encoding strategy is a second encoding method, encoding the first audio data based on a low-level encoding method to obtain a first audio bitstream; and encoding the first audio data based on a mid-level encoding method to obtain a second audio bitstream; and encoding the first audio data based on a high-level encoding method to obtain a third audio bitstream. Thus, in this embodiment of the application, the first electronic device, acting as the transmitting end, achieves layered audio transmission through multi-level encoding.

[0009] In one possible implementation, the method further includes: at a preset periodic trigger time, determining a second available bandwidth between the device and the server based on preset evaluation conditions; the second available bandwidth is not equal to the first available bandwidth; determining a second encoding strategy based on the second available bandwidth; the second encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method, and the second encoding strategy is different from the first encoding strategy; encoding the second audio data based on the second encoding strategy to obtain P audio bitstreams; P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; sending the P audio bitstreams to the server, so that the server sends the second target audio bitstream among the P audio bitstreams to a second electronic device. Thus, in this embodiment, the first electronic device, as the sending end, can adjust the available bandwidth based on the current network status and adjust the encoding method based on the available bandwidth.

[0010] In one possible implementation, the first encoding strategy being different from the second encoding strategy includes: the types of first, second, or third encoding methods included in the first and second encoding strategies being different; and / or, the bitrates of the first, second, or third encoding methods included in the first and second encoding strategies being different. Thus, in this embodiment, the first electronic device acting as the transmitter can adjust the combination of different encoding levels, and / or the bitrates corresponding to each encoding level, to achieve reasonable bandwidth utilization while improving sound quality.

[0011] In one possible implementation, the low-level encoding method in the second encoding strategy corresponds to the fourth bit rate, which falls within the first bit rate range. If the second encoding strategy includes a mid-level encoding method, this mid-level encoding method corresponds to the fifth bit rate, which falls within the second bit rate range. If the second encoding strategy includes a high-level encoding method, this high-level encoding method corresponds to the sixth bit rate, which falls within the third bit rate range. In this embodiment, the first electronic device, acting as the transmitting end, can adjust the bit rate corresponding to each encoding level to achieve reasonable bandwidth utilization while improving sound quality.

[0012] In one possible implementation, determining the first encoding strategy based on the first available bandwidth further includes: determining a redundancy depth value based on the first available bandwidth; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the first electronic device to the server; correspondingly, sending N audio streams to the server includes: sending N audio streams to the server according to the redundancy depth value. Thus, in this embodiment, the first electronic device, as the sending end, can also adjust the redundancy depth value to reasonably occupy bandwidth, thereby improving resource utilization.

[0013] In one possible implementation, the bandwidth occupied by the N audio streams is less than or equal to the first available bandwidth.

[0014] Secondly, embodiments of this application provide an audio transmission method. This method is applied to a server and includes: obtaining a first available bandwidth between the server and a second electronic device; determining a first encoding strategy corresponding to the second electronic device based on the first available bandwidth; the first encoding strategy includes one of a low-end encoding method, a mid-range encoding method, or a high-end encoding method, wherein the bitrates of the low-end, mid-range, and high-end encoding methods are different; receiving N audio bitstreams sent by the first electronic device, where N is a positive integer greater than 0 and less than 4; wherein a single bitstream among the N audio bitstreams corresponds to a low-end, mid-range, or high-end encoding method; and sending a first target audio bitstream among the N audio bitstreams to the second electronic device based on the first encoding strategy. In this way, the server can adjust the encoding type of the bitstream sent to the receiving end in a timely manner based on the available bandwidth between itself and the second electronic device as the receiving end, thereby dynamically adjusting the encoding strategy to meet the needs of users with different bandwidths.

[0015] In one possible implementation, the method further includes: acquiring a second available bandwidth with a second electronic device at a preset periodic trigger time; determining a second encoding strategy corresponding to the second electronic device based on the second available bandwidth; the second encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the second encoding strategy is different from the first encoding strategy; receiving P audio bitstreams sent by the first electronic device, where P is a positive integer greater than 0 and less than 4; wherein each of the P audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; and sending a second target audio bitstream from the P audio bitstreams to the second electronic device based on the second encoding strategy. In this way, the server dynamically adjusts the encoding strategy to accommodate the different bandwidth requirements of the same user and the needs of users with different bandwidth requirements.

[0016] In one possible implementation, determining a first encoding strategy corresponding to the second electronic device based on a first available bandwidth includes: determining a redundancy depth value based on the first available bandwidth; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the server to the second electronic device; correspondingly, sending a first target audio stream from N audio streams to the second electronic device based on the first encoding strategy includes: sending the first target audio stream from N audio streams to the second electronic device based on the first encoding strategy and the redundancy depth value.

[0017] In one possible implementation, the bandwidth occupied by the first target audio bitstream is less than or equal to the first available bandwidth.

[0018] Thirdly, embodiments of this application provide an audio transmission device applied to a first electronic device. The device includes: an evaluation module, configured to determine a first available bandwidth between itself and a server based on preset evaluation conditions; the first available bandwidth is not equal to the currently available bandwidth; a processing module, configured to determine a first encoding strategy based on the first available bandwidth; the first encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method; wherein the first encoding method includes a low-end encoding method, the second encoding method includes a low-end encoding method and a mid-end encoding method, and the third encoding method includes a low-end encoding method, a mid-end encoding method, and a high-end encoding method; the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; an encoding module, configured to encode first audio data based on the first encoding strategy to obtain N audio bitstreams; N is a positive integer greater than 0 and less than 4; wherein a single bitstream among the N audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; and a transmission module, configured to send the N audio bitstreams to a server, so that the server sends a first target audio bitstream among the N audio bitstreams to a second electronic device.

[0019] In one possible implementation, the low-end encoding method in the first encoding strategy corresponds to a first bit rate, which falls within a first bit rate range; if the first encoding strategy includes a mid-end encoding method, the mid-end encoding method in the first encoding strategy corresponds to a second bit rate, which falls within a second bit rate range; if the first encoding strategy includes a high-end encoding method, the high-end encoding method in the first encoding strategy corresponds to a third bit rate, which falls within a third bit rate range; wherein, the minimum value of the third bit rate range is greater than the maximum value of the second bit rate range, and the minimum value of the second bit rate range is greater than the maximum value of the first bit rate range.

[0020] In one possible implementation, the encoding module is specifically configured to: if the first encoding strategy is a first encoding method, encode the first audio data based on the low-end encoding method to obtain a first audio bitstream; if the first encoding strategy is a second encoding method, encode the first audio data based on the low-end encoding method to obtain a first audio bitstream; and encode the first audio data based on the mid-end encoding method to obtain a second audio bitstream; if the first encoding strategy is a second encoding method, encode the first audio data based on the low-end encoding method to obtain a first audio bitstream; and encode the first audio data based on the mid-end encoding method to obtain a second audio bitstream; and encode the first audio data based on the high-end encoding method to obtain a third audio bitstream.

[0021] In one possible implementation, the evaluation module is further configured to determine a second available bandwidth with the server based on preset evaluation conditions at a preset periodic trigger time; the second available bandwidth is not equal to the first available bandwidth; the processing module is further configured to determine a second encoding strategy based on the second available bandwidth; the second encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method, and the second encoding strategy is different from the first encoding strategy; the encoding module is further configured to encode the second audio data based on the second encoding strategy to obtain P audio bitstreams; P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the transmission module is further configured to send the P audio bitstreams to the server, so that the server sends Q audio bitstreams among the P audio bitstreams to the second electronic device.

[0022] In one possible implementation, the first encoding strategy being different from the second encoding strategy includes: the first encoding strategy and the second encoding strategy including different types of first encoding methods, second encoding methods, or third encoding methods; and / or, the first encoding strategy and the second encoding strategy including different bit rates of first encoding methods, second encoding methods, or third encoding methods.

[0023] In one possible implementation, the low-end encoding method in the second encoding strategy corresponds to the fourth bit rate, which belongs to the first bit rate range; if the second encoding strategy includes a mid-end encoding method, the mid-end encoding method in the second encoding strategy corresponds to the fifth bit rate, which belongs to the second bit rate range; if the second encoding strategy includes a high-end encoding method, the high-end encoding method in the second encoding strategy corresponds to the sixth bit rate, which belongs to the third bit rate range.

[0024] In one possible implementation, the processing module is further configured to: determine a redundancy depth value based on a first available bandwidth; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the first electronic device to the server; and the transmission module is further configured to send N audio streams to the server according to the redundancy depth value.

[0025] In one possible implementation, the bandwidth occupied by the N audio streams is less than or equal to the first available bandwidth.

[0026] Fourthly, embodiments of this application provide an audio transmission device. Applied to a server, the device includes: an acquisition module for acquiring a first available bandwidth between itself and a second electronic device; a processing module for determining a first encoding strategy corresponding to the second electronic device based on the first available bandwidth; the first encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method, wherein the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; a transmission module for receiving N audio bitstreams sent by the first electronic device, where N is a positive integer greater than 0 and less than 4; wherein a single bitstream among the N audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the transmission module is further configured to send a first target audio bitstream among the N audio bitstreams to the second electronic device based on the first encoding strategy.

[0027] In one possible implementation, the device further includes: an acquisition module, further configured to acquire a second available bandwidth between itself and a second electronic device at a preset periodic trigger time; a processing module, further configured to determine a second encoding strategy corresponding to the second electronic device based on the second available bandwidth; the second encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method; a transmission module, further configured to receive P audio bitstreams sent by the first electronic device, where P is a positive integer greater than 0 and less than 4; wherein a single bitstream among the P audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; and the transmission module, further configured to send a second target audio bitstream among the P audio bitstreams to the second electronic device based on the second encoding strategy.

[0028] In one possible implementation, the processing module is further configured to: determine a redundancy depth value based on a first available bandwidth; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the server to the second electronic device; the transmission module is further configured to send a first target audio stream from N audio streams to the second electronic device based on a first encoding strategy and the redundancy depth value.

[0029] In one possible implementation, the bandwidth occupied by the first target audio bitstream is less than or equal to the first available bandwidth.

[0030] Fifthly, embodiments of this application provide a computer-readable medium for storing a computer program, the computer program including instructions for performing the method in the first aspect or any possible implementation of the first aspect.

[0031] In a sixth aspect, embodiments of this application provide a computer-readable medium for storing a computer program, the computer program including instructions for performing the methods in the second aspect or any possible implementation thereof.

[0032] In a seventh aspect, embodiments of this application provide a computer program including instructions for performing the method in the first aspect or any possible implementation thereof.

[0033] Eighthly, embodiments of this application provide a computer program including instructions for performing the method in the second aspect or any possible implementation thereof.

[0034] Ninthly, embodiments of this application provide a chip including a processing circuit and transceiver pins. The transceiver pins and the processing circuit communicate with each other via an internal connection path. The processing circuit executes the method in the first aspect or any possible implementation of the first aspect to control the receiving pin to receive signals and to control the transmitting pin to transmit signals.

[0035] In a tenth aspect, embodiments of this application provide a chip including a processing circuit and transceiver pins. The transceiver pins and the processing circuit communicate with each other via an internal connection path. The processing circuit executes the method in the second aspect or any possible implementation of the second aspect to control the receiving pin to receive signals and to control the transmitting pin to transmit signals.

[0036] Eleventhly, embodiments of this application provide an audio transmission system, which includes the first electronic device, server, and second electronic device mentioned in the first and second aspects above. Attached Figure Description

[0037] Figure 1 This application provides a schematic diagram of a communication system according to an embodiment;

[0038] Figure 2 This is a schematic diagram illustrating the audio transmission process for an uplink user;

[0039] Figure 3 A flowchart illustrating the setup of an encoding strategy that increases bandwidth, as shown in the example.

[0040] Figure 4 A flowchart illustrating the setup of an exemplary bandwidth-reducing encoding strategy;

[0041] Figure 5 This is a schematic diagram illustrating the audio transmission process for a downlink user, as an example.

[0042] Figure 6 A flowchart illustrating the setup of an encoding strategy that increases bandwidth, as shown in the example.

[0043] Figure 7 A flowchart illustrating the setup of an exemplary bandwidth-reducing encoding strategy;

[0044] Figure 8This is a schematic diagram illustrating an application scenario;

[0045] Figure 9 This is a schematic diagram illustrating an application scenario;

[0046] Figure 10 This is a schematic diagram illustrating an application scenario;

[0047] Figure 11 This is a schematic diagram of the structure of an exemplary device;

[0048] Figure 12 This is a schematic diagram of the structure of an exemplary device;

[0049] Figure 13 This is a schematic diagram of the structure of an exemplary device. Detailed Implementation

[0050] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0052] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.

[0053] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0054] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.

[0055] First, a brief explanation of some terms used in the embodiments of this application:

[0056] Data Rate: refers to the amount of data used by a video or audio file per unit of time.

[0057] Sampling rate: The number of samples extracted from a continuous signal per second to form a discrete signal.

[0058] Before describing the technical solutions of the embodiments of this application, the communication system of the embodiments of this application will first be described with reference to the accompanying drawings. See also Figure 1 This is a schematic diagram of a communication system provided in an embodiment of this application. The communication system includes electronic device A, electronic device B, electronic device C, and a server. For example, Figure 1 The number and type of devices mentioned are merely examples of suitability and are not intended to limit the scope of this application. In the embodiments of this application, the server may be a Selective Forwarding Unit (SFU). Electronic devices (e.g., electronic devices A to C) may be mobile phones, computers, tablets, wearable devices, smart home devices (e.g., televisions, speakers, etc.), etc., and are not limited in this application.

[0059] In existing technical embodiments, taking electronic device A as the audio transmitter (also known as the uplink user) and electronic devices B and C as receivers (also known as downlink users) as an example, electronic device A encodes the audio at a fixed bitrate and sends the encoded bitstream to the SFU. The SFU then forwards the bitstream to electronic devices B and C. Electronic devices B and C decode the bitstream to obtain the audio data and play it. In existing technologies, due to the use of fixed bitrate encoding, the technology cannot meet the needs of users with different bandwidths, often resulting in stuttering issues for users with limited bandwidth. Especially in conferencing systems, electronic device A encodes at the optimal bitrate and sends the bitstream with the optimal redundancy depth, causing bandwidth waste and insufficient bandwidth for other functions such as video, resulting in video stuttering, poor audio quality, and other problems, affecting the user experience. It should be noted that the redundancy depth involved in this application embodiment is used to indicate the redundancy of the bitstream. For example, if the redundancy depth is 2, then each data packet not only contains the current audio frame but also the previous two audio frames. Even if two of the three data packets are lost, all audio frames can still be parsed.

[0060] Several possible audio transmission methods have been proposed in existing technologies. One possible implementation uses a Microcontroller Unit (MCU) mixing mode. In this mode, the audio for a general conference is mixed by the MCU server up to three times before being sent to downstream users, with adaptive encoding based on the downstream users' network conditions. This method has the following problems: the MCU server needs to perform mixing for each downstream user, resulting in high server costs, significant overhead, high latency, and performance bottlenecks in very large conferences. Another possible implementation uses a Single-Functional-Unit (SFU) pure forwarding mode: the SFU purely forwards the three downstream streams to the downstream terminal. The upstream has only one encoding level, possessing some network adaptability. When the upstream network is limited or unstable, it adaptively adjusts the redundancy depth and bitrate based on the current network. The SFU adjusts the redundancy depth and forwards data to downstream users based on the downstream network conditions. This method has the following problems: when the downstream network is limited or unstable, and the upstream network is good with a high bitrate, the SFU directly forwards the downstream stream, resulting in severe packet loss, leading to choppy audio and poor sound quality. Furthermore, when the downlink network is limited or unstable, downlink audio consumes a large amount of bandwidth, resulting in less bandwidth for video data and other data, thus leading to a poor overall experience.

[0061] This application provides an audio transmission method that employs a layered encoding scheme, enabling different users to obtain optimal sound quality under their respective bandwidth conditions. Furthermore, in situations with poor downlink network conditions, the same audio effect requires less downlink bandwidth, allowing more bandwidth to be allocated to video and data, thereby improving the overall experience. Additionally, by using independent uplink encoding and layered packetization, the receiving end can simply decode according to the standard, effectively improving the receiver's compatibility.

[0062] Figure 2 This is a schematic diagram illustrating the audio transmission process for an uplink user, please refer to... Figure 2 Specifically, including but not limited to the following steps:

[0063] S201, the first electronic device determines the first available bandwidth between itself and the server based on preset evaluation conditions.

[0064] The first electronic device involved in this application embodiment is an audio transmitter, also known as an uplink user, for example... Figure 1 Electronic device A in the middle.

[0065] For example, the first electronic device may provide a user interface that allows the user to set the maximum audio bandwidth. The maximum audio bandwidth refers to the maximum bandwidth occupied by the audio transmission (i.e., the bitstream described below, which may also be called an audio bitstream or audio data packet, etc., which are not limited in this application) between the first electronic device and the server. In other words, during subsequent transmission of the audio bitstream, the bandwidth occupied by the audio bitstream is less than or equal to the maximum audio bandwidth.

[0066] Optionally, users can adjust the maximum audio bandwidth at any time during the audio transmission process, which is not limited in this application.

[0067] Optionally, the first electronic device may also offer some maximum audio bandwidth settings, such as a recommended maximum audio bandwidth of 64kbps for mono in 16-24kHz voice scenarios and 128kbps for 48kHz music scenarios.

[0068] The first electronic device is equipped with bandwidth assessment criteria. These criteria are used to evaluate the network status and determine the available bandwidth between the device and the server.

[0069] Specifically, the first electronic device determines whether to adjust the currently available bandwidth based on bandwidth assessment conditions. The currently available bandwidth can also be referred to as the currently available audio bandwidth, which is the bandwidth used for audio data interaction between the first electronic device and the server. Optionally, adjusting the available bandwidth includes increasing the bandwidth (i.e., adding bandwidth) or decreasing the bandwidth.

[0070] The bandwidth evaluation conditions in the embodiments of this application may include, but are not limited to:

[0071] If the packet loss rate is within a first packet loss rate threshold, and / or the bandwidth change rate is within a first bandwidth change rate threshold, it is determined that the available bandwidth needs to be increased.

[0072] If the packet loss rate is within the second packet loss rate threshold, and / or the bandwidth change rate is within the second bandwidth change rate threshold, it is determined that the available bandwidth needs to be reduced.

[0073] If the packet loss rate is not within the range of the first packet loss rate threshold and the second packet loss rate threshold, and / or the bandwidth change rate is not within the range of the first bandwidth change rate and the second bandwidth change rate, it is determined that no adjustment of the available bandwidth is required.

[0074] It should be noted that the thresholds or threshold ranges involved in the embodiments of this application are merely examples of suitability and can be set according to actual needs. This application does not impose any limitations.

[0075] In one example, if the first electronic device determines, based on bandwidth assessment conditions, that no adjustment to the current available bandwidth is needed, the current process ends. The first electronic device then executes S201 again at a preset periodic trigger time. That is, in this embodiment of the application... Figure 2 The process shown is executed periodically. Optionally, the duration of the cycle can be set according to actual needs, and this application does not impose any limitations.

[0076] In another example, if the first electronic device determines, based on bandwidth assessment conditions, that the current available bandwidth needs to be increased or decreased, then the adjusted first available bandwidth is further determined, and S202 is executed.

[0077] For example, the first electronic device can be pre-set to increase and decrease the bandwidth coefficient. The increasing and decreasing bandwidth coefficients can be the same or different; this application does not impose any limitation. The specific values ​​of the increasing and decreasing bandwidth coefficients can be set according to actual needs, for example, both can be 10%; this application does not impose any limitation.

[0078] In this embodiment, the first electronic device can increase or decrease the bandwidth based on increasing or decreasing the bandwidth coefficient. For example, if the first electronic device determines that the bandwidth coefficient needs to be increased, it will increase the current bandwidth by 10% based on increasing the bandwidth coefficient (e.g., 10%), and the adjusted bandwidth will be the first available bandwidth. Decreasing the bandwidth is similar and will not be described in detail here.

[0079] In this embodiment of the application, as described above, the user sets a maximum audio bandwidth. Therefore, after the first electronic device increases the currently available bandwidth, the determined first available bandwidth must be less than or equal to the maximum audio bandwidth. That is, if the first available bandwidth determined by the first electronic device based on the bandwidth increase factor is greater than the maximum audio bandwidth, then the finally determined first available bandwidth is the maximum audio bandwidth.

[0080] S202, the first electronic device determines a first coding strategy based on the first available bandwidth.

[0081] In this embodiment, after the first electronic device determines the first available bandwidth, it can determine the corresponding encoding strategy based on the first available bandwidth, such as a first encoding strategy. In this embodiment, the encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method. Specifically, the first encoding method includes a low-end encoding method, the second encoding method includes a low-end encoding method and a mid-range encoding method, and the third encoding method includes a low-end encoding method, a mid-range encoding method, and a high-end encoding method. That is, the encoding method in this embodiment is a combination of a low-end encoding method, a mid-range encoding method, and a high-end encoding method.

[0082] In this embodiment, the electronic device is configured with different encoding levels, each corresponding to a different bitrate range. The bitrate affects the quality of the encoded audio. A higher bitrate results in greater bandwidth usage and higher audio quality, while a lower bitrate results in less bandwidth usage and lower audio quality.

[0083] In this embodiment, the first electronic device may optionally encode using an encoding method based on a 16kHz sampling rate, that is, ... The encoding bitrate based on a 16kHz sampling rate includes, but is not limited to: (7, 10, 12, 15, 18, 20, 23, 25, 28, 30) (unit: kHz). Correspondingly, in this embodiment, the bitrate range corresponding to the low-end encoding method is 12kHz. The bitrate range corresponding to the mid-range encoding method is 15kHz-25kHz (i.e., 15kHz, 18kHz, 20kHz, 25kHz). The bitrate range corresponding to the high-end encoding method is 20kHz-30kHz (i.e., 20kHz, 25kHz, 28kHz, 30kHz). The bitrate ranges described in this embodiment are merely examples of suitability and are not intended to limit the application. As mentioned above, a higher bitrate generally results in better audio quality; correspondingly, the audio quality of the high-end encoding method is higher than that of the mid-range encoding method, which is higher than that of the low-end encoding method. Furthermore, the bandwidth occupied by the high-end audio bitstream is less than that occupied by the mid-range audio bitstream, which is less than that occupied by the low-end audio bitstream.

[0084] It should be noted that this embodiment only uses a 16kHz sampling rate encoding bitrate as an example. In other embodiments, other sampling rates can be used, and the bitrate range corresponding to each level can be adjusted based on the sampling rate change; this application does not limit this.

[0085] Specifically, as described above, the first electronic device can determine whether to increase or decrease the available bandwidth. The encoding strategies corresponding to increasing and decreasing the available bandwidth are different, and the corresponding encoding strategy settings are described below:

[0086] 1. Increase bandwidth

[0087] Figure 3 The flowchart illustrates the setup process for an exemplary encoding strategy that increases bandwidth. Please refer to... Figure 3Specifically, after determining that the available bandwidth needs to be increased, the first electronic device readjusts the encoding strategy and redistributes the bandwidth of each encoding level within the determined first available bandwidth. For example, the first electronic device detects the current encoding method, as described above, which includes a first encoding method (i.e., low-level encoding method), a second encoding method (i.e., a combination of low-level and mid-level encoding methods), and a third encoding method (i.e., a combination of low-level, mid-level, and high-level encoding methods).

[0088] like Figure 3 As shown in one example, the first electronic device determines that the current encoding method is the third encoding method. Therefore, the first electronic device determines that the current redundancy depth value and the low-level encoding bitrate remain unchanged. The first electronic device increases the encoding bitrate of the mid-level encoding. Optionally, the first electronic device increases the mid-level encoding bitrate to the maximum value, which is the maximum value of the bitrate range corresponding to the mid-level encoding method described above, for example, 25kHz.

[0089] For example, during the process of increasing the mid-range encoding bitrate, the first electronic device calculates the bandwidth occupied by the bitstream corresponding to the adjusted third encoding method, that is, the total bandwidth occupied by the bitstream corresponding to the low-end encoding method bitrate, the bitstream corresponding to the mid-range encoding method with the adjusted mid-range encoding bitrate, and the bitstream corresponding to the current high-end encoding method bitrate. The total bandwidth must be less than or equal to the first available bandwidth.

[0090] In one possible implementation, if the total bandwidth is already equal to the first available bandwidth when the first electronic device has not adjusted the bit rate of the mid-range encoding method to the highest bit rate, then the current encoding strategy adjustment process ends, and the adjusted encoding strategy is: the low-range encoding method is the current bit rate, the mid-range encoding method is the adjusted bit rate, and the high-range encoding method is the current bit rate.

[0091] In another possible implementation, if the first electronic device adjusts the bitrate of the intermediate encoding method to the highest bitrate, or if the intermediate encoding method is already at its highest bitrate, the first electronic device may optionally reduce (i.e., decrease) the packetization time. The packetization time may optionally be the Real-time Transport Protocol (RTP) packetization time, i.e., how often an RTP packet is sent; typically, the packetization time is 20ms or 40ms. For example, the first electronic device may reduce the packetization time to 20ms. Optionally, if the first packetization time is already at its minimum, such as 20ms, no further processing is required.

[0092] Still refer to Figure 3The first electronic device increases the high-end coding bitrate. Optionally, the first electronic device increases the mid-range coding bitrate to its maximum value, which is the maximum value of the bitrate range corresponding to the mid-range coding method mentioned above, for example, 30kHz. During the process of increasing the mid-range coding bitrate, the first electronic device calculates the bandwidth occupied by the bitstream corresponding to the adjusted third coding method, that is, the total bandwidth occupied by the bitstream corresponding to the high-end coding method bitrate, the bitstream corresponding to the mid-range coding method with the adjusted mid-range coding bitrate, and the bitstream corresponding to the high-end coding method bitrate. The total bandwidth must be less than or equal to the first available bandwidth.

[0093] In one possible implementation, if the total bandwidth is already equal to the first available bandwidth even though the first electronic device has not adjusted the bit rate of the high-end encoding method to the highest bit rate, then the current encoding strategy adjustment process ends, and the adjusted encoding strategy is: the low-end encoding method is the current bit rate, the mid-end encoding method is the adjusted bit rate, and the high-end encoding method is the adjusted bit rate.

[0094] In this embodiment, the first electronic device can gradually adjust the medium-range or high-range encoding bitrate. That is, for each increase in bitrate level, the corresponding total bandwidth is calculated until the maximum value of the corresponding encoding method is reached. Optionally, the first electronic device can also directly adjust the encoding bitrate to the maximum value of the level. If the total bandwidth exceeds the first available bandwidth, the bitrate is reduced to the second-to-last level, and the total bandwidth is calculated again, and this process is repeated.

[0095] Please continue to refer to Figure 3 In another example, the first electronic device determines that the current encoding method is the second encoding method, which is a combination of low-end and mid-end encoding methods. The first electronic device can determine whether to switch to a third encoding method based on preset decision conditions.

[0096] Optionally, the decision condition is used to indicate whether the total bandwidth occupied by the bitstream corresponding to the adjusted third encoding method is less than or equal to the first available bandwidth.

[0097] For example, decision conditions could be:

[0098] 1) Keep the current bitrate of the low-end and mid-end encoding methods unchanged, and after adjusting the encoding method to the third encoding method, the total bandwidth is less than or equal to the first available bandwidth.

[0099] And / or,

[0100] 2) Keep the bit rate of the low-end encoding method unchanged, reduce the mid-end encoding method, and adjust the encoding method to the third encoding method. The total bandwidth is less than or equal to the first available bandwidth.

[0101] The above decision-making conditions are merely illustrative examples and can be set according to actual needs; this application does not impose any limitations.

[0102] For example, if the first electronic device determines that decision condition 1) is met, the first electronic device determines an encoding strategy, and the new encoding method is the third encoding method, which includes a low-level encoding method, a mid-level encoding method, and a high-level encoding method. The bitrates corresponding to the low-level and mid-level encoding methods remain the same as the current bitrates; that is, the bitrates of these two levels are not adjusted. The first electronic device can determine the bitrate of the high-level encoding method based on the difference between the first available bandwidth and the bandwidth occupied by the bitstreams of the low-level and mid-level encoding methods. That is, the bitrate of the high-level encoding method can be any bitrate within the bitrate range that is less than or equal to the difference between the first available bandwidth and the bandwidth occupied by the bitstreams of the low-level and mid-level encoding methods.

[0103] For example, if the first electronic device determines that condition 2) is met, the first electronic device determines an encoding strategy. The new encoding strategy is a third encoding method, including a low-level encoding method, a mid-level encoding method, and a high-level encoding method. The bitrate of the low-level encoding method remains unchanged, the bitrate of the mid-level encoding method is reduced, and the high-level encoding method is added. For example, the specific values ​​of the adjusted bitrates of the mid-level and high-level encoding methods can be set according to actual needs. For instance, the bitrate of the mid-level encoding method can be reduced by one level, and the remaining bandwidth of the first available bandwidth can be allocated to the high-level encoding method. In other words, the corresponding level can be selected based on the bandwidth occupied by the level, and the specific selection method can be set according to actual needs; this application does not limit this selection.

[0104] Optionally, the first electronic device can detect whether the decision condition is met by setting a new encoding method and then calculating whether the bandwidth corresponding to the encoding method is less than or equal to the first available bandwidth. Alternatively, the first electronic device can also detect whether the decision condition is met by adjusting the mid-range encoding method and then calculating whether the difference between the first available bandwidth and the bandwidth occupied by the mid-range and low-range encoding methods is greater than or equal to the minimum value of the high-range encoding method. The specific method can be set according to actual needs, and this application does not limit it.

[0105] Please continue to refer to Figure 3 In another example, the first electronic device determines that the current encoding method is the first encoding method, which is a low-level encoding method. The first electronic device can determine whether to switch to the second encoding method based on preset decision conditions.

[0106] Optionally, the decision condition is used to indicate whether the total bandwidth occupied by the bitstream corresponding to the adjusted second encoding method is less than or equal to the first available bandwidth.

[0107] For example, decision conditions could be:

[0108] Keeping the current bitrate of the low-level encoding method unchanged, after adjusting the encoding method to the second encoding method, the total bandwidth is less than or equal to the first available bandwidth.

[0109] The above decision-making conditions are merely illustrative examples and can be set according to actual needs; this application does not impose any limitations.

[0110] For example, if the first electronic device determines that the decision conditions are met, it determines an encoding strategy. The new encoding method is the second encoding method, which includes a low-end encoding method and a mid-end encoding method. The bitrate corresponding to the low-end encoding method remains the same as the current one; that is, the low-end encoding bitrate is not adjusted. The first electronic device can determine the bitrate of the mid-end encoding method based on the difference between the first available bandwidth and the bandwidth occupied by the low-end encoding method's bitstream. That is, the bitrate of the mid-end encoding method can be any bitrate within the bitrate range that is less than or equal to the difference between the first available bandwidth and the bandwidth occupied by the low-end encoding method's bitstream.

[0111] 2. Reduce bandwidth

[0112] Figure 4 A flowchart illustrating the exemplary bandwidth reduction encoding strategy is provided. Please refer to... Figure 4 Specifically, after determining that the available bandwidth needs to be increased, the first electronic device readjusts the encoding strategy and redistributes the bandwidth of each encoding level within the determined first available bandwidth. For example, the first electronic device detects the current encoding method, as described above, which includes a first encoding method (i.e., low-level encoding method), a second encoding method (i.e., a combination of low-level and mid-level encoding methods), and a third encoding method (i.e., a combination of low-level, mid-level, and high-level encoding methods).

[0113] Please refer to Figure 4 In one example, the first electronic device determines that the current encoding method is the first encoding method. Therefore, the first electronic device determines that the encoding strategy remains the first encoding method, and the bitrate remains unchanged. The first electronic device adjusts the redundancy depth value, thereby reducing the bandwidth occupied by the audio bitstream. That is, in this example, when the bandwidth occupied by the encoding method is already at its minimum, the bandwidth occupied by the audio bitstream can be further reduced by adjusting the redundancy depth value. Optionally, the first electronic device calculates a redundancy depth prediction value according to a preset rule (which can be set according to actual needs, and is not limited in this application). If the predicted value is greater than the actual value by a certain value, the redundancy depth is doubled; if the predicted value is less than the actual value by a certain value, the redundancy depth is halved.

[0114] Please continue to refer to Figure 4In another example, the first electronic device determines that the current encoding strategy is the second encoding method, which is a combination of low-end and mid-end encoding methods. Optionally, the first electronic device increases the packing time (the concept can be referred to above), for example, switching the packing time from 20ms to 40ms. Then, the first electronic device reduces the mid-end encoding bitrate to the minimum value of the mid-end encoding bitrate range or reduces the bitrate by a predetermined amount (e.g., reducing it by one bit). Optionally, if the mid-end encoding bitrate is already at its minimum value, the encoding method is switched to the first encoding method, that is, the combination of low-end and mid-end encoding methods is switched to low-end encoding method.

[0115] like Figure 4 As shown in another example, the first electronic device determines that the current encoding method is the third encoding method. The first electronic device reduces the bitrate of the high-end encoding method, which can be by reducing it by a certain amount (e.g., reducing it by one level) or reducing it to the minimum value. Optionally, if the bitrate of the high-end encoding method is already the minimum value of the high-end encoding method bitrate range, then the high-end encoding method is canceled and switched to the second encoding method, that is, a combination of the low-end encoding method and the mid-end encoding method. Optionally, the bitrate of the mid-end encoding method can remain unchanged, or it can be any value less than or equal to the difference between the first available bandwidth and the bandwidth occupied by the low-end encoding method bitstream.

[0116] In summary, the encoding strategy adjustments in this application include, but are not limited to, adjusting the combination of encoding methods and / or adjusting the bitrate corresponding to the encoding method. This application can adjust the bandwidth occupied by the corresponding bitstream based on network conditions by adjusting the combination of encoding methods in the encoding strategy. Furthermore, this application further refines the bandwidth occupation by adjusting the bitrate within the bitrate range corresponding to the encoding method, achieving the optimal combination of encoding methods that meets actual needs.

[0117] S203, the first electronic device encodes the first audio data based on the first encoding strategy to obtain N audio bitstreams.

[0118] For example, after determining a new encoding strategy (i.e., the first encoding strategy), the first electronic device can encode the first audio data based on the encoding strategy. The first audio data is audio data generated by an application of the first electronic device.

[0119] In one example, if the first encoding strategy includes a first encoding method, then the first electronic device encodes the first audio data based on the bitrate of the low-end encoding method to obtain a low-end audio bitstream.

[0120] In another example, if the first encoding strategy includes a second encoding method, then the first electronic device encodes the first audio data based on the bitrate of the low-end encoding method to obtain a low-end audio bitstream. Furthermore, the first electronic device encodes the first audio data based on the bitrate of the medium-end encoding method to obtain a medium-end audio bitstream.

[0121] In another example, if the first encoding strategy includes a third encoding method, then the first electronic device encodes the first audio data based on the bitrate of the low-end encoding method to obtain a low-end audio bitstream. Furthermore, the first electronic device encodes the first audio data based on the bitrate of the mid-end encoding method to obtain a mid-end audio bitstream. Finally, the first electronic device encodes the first audio data based on the bitrate of the high-end encoding method to obtain a high-end audio bitstream.

[0122] S204, the first electronic device sends N audio streams to the server.

[0123] The first electronic device sends the encoded bitstream to the SFU according to the redundancy depth. The bitstream can be a low-end audio bitstream, a combination of low-end and mid-end audio bitstreams, or a combination of low-end, mid-end, and high-end audio bitstreams.

[0124] Figure 5 This is a schematic diagram illustrating the audio transmission process for a downlink user, please refer to... Figure 5 Specifically, including but not limited to the following steps:

[0125] S501, the server obtains a second available bandwidth between itself and the second electronic device.

[0126] For example, the second electronic device (also referred to as the downlink user) acts as the audio receiver, and the user can set the maximum audio bandwidth on the device. The maximum audio bandwidth is the maximum bandwidth expected to be used for audio data interaction between the second electronic device and the server. That is, the bandwidth occupied by the audio bitstream transmitted between the server and the second electronic device is less than this maximum audio bandwidth.

[0127] For example, the second electronic device can periodically determine the available bandwidth for audio transmission with the server based on preset evaluation conditions and the maximum audio bandwidth set by the user. The method for confirming the available bandwidth can be referred to the relevant description in S201, and will not be repeated here.

[0128] The second electronic device sends an instruction to the server to indicate the available bandwidth for the next cycle (i.e., the second available bandwidth).

[0129] The server receives the instruction information sent by the second electronic device and obtains the second available bandwidth between the server and the second electronic device.

[0130] S502, the server determines a second encoding strategy corresponding to the second electronic device based on the second available bandwidth.

[0131] In this embodiment, after the server determines the second available bandwidth, it can determine the corresponding encoding strategy based on the second available bandwidth, such as a second encoding strategy. The second encoding strategy is used to indicate the encoding method corresponding to the bitstream sent to the electronic device. The encoding strategy corresponding to the receiving end can also be called a receiving strategy or a receiving encoding strategy, which is not limited in this application. For example, the second encoding strategy includes, but is not limited to: low-end encoding method, mid-end encoding method, or high-end encoding method.

[0132] The server can determine whether to increase or decrease the bandwidth based on the received bandwidth, and adjust it according to the encoding strategy settings for increasing or decreasing bandwidth described below.

[0133] In one possible implementation, the server can be configured with a bandwidth increase threshold and a bandwidth decrease threshold (these can be set according to actual needs, and this application does not limit them). In one example, when the server detects that the bandwidth increase is greater than or equal to the bandwidth increase threshold, it can adjust the encoding strategy corresponding to the bandwidth increase described below. If the increase is less than the bandwidth increase threshold, no action is taken, and the bitstream continues to be sent according to the current encoding strategy. In another example, when the server detects that the bandwidth decrease is greater than or equal to the bandwidth decrease threshold, it can adjust the encoding strategy corresponding to the bandwidth decrease described below. If the decrease is less than the bandwidth decrease threshold, no action is taken, and the bitstream continues to be sent according to the current encoding strategy.

[0134] Optionally, in other embodiments, the server may also set bandwidth increase or decrease conditions, such as bandwidth increasing for three consecutive cycles, or bandwidth increasing continuously with an increase greater than or equal to a bandwidth increase threshold. In such cases, adjustments can be made according to the encoding strategy described below; otherwise, no processing is performed. Specific conditions can be set according to actual needs, and this application does not limit them. The bandwidth decrease condition is similar and will not be described in detail here.

[0135] The following sections describe how to set the corresponding encoding strategies:

[0136] 1) Increase bandwidth

[0137] Figure 6 The flowchart illustrates the setup process for an exemplary encoding strategy that increases bandwidth. Please refer to... Figure 6 Specifically, after the server determines that the available bandwidth (referring to the available audio bandwidth) between itself and the second electronic device has been increased, it readjusts the encoding strategy corresponding to the second electronic device.

[0138] For example, the server detects the current encoding method of the second electronic device, as described above, which includes low-end encoding method, mid-end encoding method, or high-end encoding method.

[0139] like Figure 6 As shown in one example, if the server determines that the current encoding method of the second electronic device is a high-level encoding method, then there is no need to adjust the encoding strategy, and the high-level encoding method will remain unchanged. Next, the server reduces the redundancy depth value corresponding to the second electronic device; the reduction amount can be set according to actual needs, and this application does not limit it. In other words, when the network status of the second electronic device is excellent, there is no need to repeatedly send data packets, thereby improving bandwidth utilization.

[0140] Please continue to refer to Figure 6 In another example, the server determines that the current encoding method of the second electronic device is a mid-range encoding method, and then switches the mid-range encoding method to a high-range encoding method. Next, the server reduces the redundancy depth value corresponding to the second electronic device. The reduction amount can be set according to actual needs, and this application does not limit it. Optionally, before adjusting the encoding strategy, the server can detect whether the adjusted available bandwidth is greater than or equal to the bandwidth occupied by the high-range encoding method's bitstream. In one example, if the available bandwidth is greater than or equal to the bandwidth occupied by the high-range encoding method's bitstream, that is, the adjusted available bandwidth can support the transmission of the high-range encoding bitstream, then it is determined to adjust the mid-range encoding method to a high-range encoding method. In another example, if the available bandwidth is less than the bandwidth occupied by the high-range encoding method's bitstream, that is, the adjusted available bandwidth does not support the transmission of the high-range encoding bitstream, then no adjustment is performed, and the mid-range encoding method continues to be maintained.

[0141] Please continue to refer to Figure 6 In another example, the server determines that the current encoding method of the second electronic device is a low-level encoding method, and then switches the low-level encoding method to a medium-level encoding method. Next, the server reduces the redundancy depth value corresponding to the second electronic device. The reduction amount can be set according to actual needs, and this application does not limit it. Optionally, before adjusting the encoding strategy, the server checks whether the adjusted available bandwidth supports the transmission of the bitstream using the medium-level encoding method. The specific detection method can be referred to above, and will not be repeated here.

[0142] 2) Reduce bandwidth

[0143] Figure 7 A flowchart illustrating the exemplary bandwidth reduction encoding strategy is provided. Please refer to... Figure 7 Specifically, after the server determines that the available bandwidth (referring to the available audio bandwidth) between itself and the second electronic device has decreased, it readjusts the encoding strategy corresponding to the second electronic device.

[0144] For example, the server detects the current encoding method of the second electronic device, as described above, which includes low-end encoding method, mid-end encoding method, or high-end encoding method.

[0145] like Figure 7 As shown in one example, if the server determines that the current encoding method of the second electronic device is high-level encoding, the server will switch the high-level encoding method to medium-level encoding. Next, the server increases the redundancy depth value corresponding to the second electronic device; the increase can be set according to actual needs, and this application does not limit it. In other words, when the network condition of the second electronic device is poor, the redundancy of repeatedly sending data packets can be increased to improve the data packet reception accuracy at the receiving end.

[0146] Please continue to refer to Figure 7 In another example, the server determines that the current encoding method of the second electronic device is a medium-range encoding method. The server then switches the medium-range encoding method to a low-range encoding method to reduce bandwidth consumption. Next, the server increases the redundancy depth value corresponding to the second electronic device. The increase can be set according to actual needs, and this application does not limit it.

[0147] Please continue to refer to Figure 6 In another example, if the server determines that the current encoding method of the second electronic device is a low-level encoding method, then the encoding method remains unchanged and is still a low-level encoding method. Next, the server increases the redundancy depth value corresponding to the second electronic device. The increase can be set according to actual needs, and this application does not limit it.

[0148] S503, the server receives N audio streams sent by the first electronic device.

[0149] For example, as described above, in S204, the first electronic device sends N audio streams to the server. The streams can be low-end audio streams, a combination of low-end and mid-end audio streams, or a combination of low-end, mid-end, and high-end audio streams.

[0150] S504, the server sends the first target audio stream from N audio streams to the second electronic device based on the second encoding strategy.

[0151] For example, the server may send the first target audio stream from N audio streams to the second electronic device based on redundancy depth and a second encoding strategy.

[0152] Optionally, in this embodiment of the application, if the encoding mode level indicated by the second encoding strategy of the second electronic device is higher than the encoding mode level received by the server, the server will send an encoded bitstream with a lower level than that indicated by the second encoding strategy to the second electronic device.

[0153] For example:

[0154] In one example, if the server receives a low-resolution audio stream corresponding to a low-resolution encoding method sent by the first electronic device, and the second encoding strategy of the second electronic device indicates a low-resolution encoding method, the server sends a low-resolution audio stream (i.e., the first target audio stream) to the second electronic device, which will not be described again below.

[0155] In another example, if the server receives a low-level audio bitstream corresponding to a low-level encoding method sent by the first electronic device, and the second encoding strategy of the second electronic device indicates a medium-level or high-level encoding method, as mentioned above, the encoding method level indicated by the encoding strategy of the second electronic device is higher than the encoding method level received by the server, the server can send an encoding method lower than that indicated by the encoding strategy to the second electronic device, and correspondingly, the server sends a low-level audio bitstream to the second electronic device.

[0156] In another example, if the server receives a low-end audio stream corresponding to a low-end encoding method and a medium-end audio stream corresponding to a medium-end encoding method from the first electronic device, and the second encoding strategy of the second electronic device indicates the low-end encoding method, the server sends the low-end audio stream to the second electronic device.

[0157] In another example, if the server receives a low-end audio stream corresponding to a low-end encoding method and a stream corresponding to a mid-end encoding method from the first electronic device, and the second encoding strategy of the second electronic device indicates a mid-end encoding method or a high-end encoding method, the server sends a mid-end audio stream to the second electronic device.

[0158] In another example, if the server receives a low-end audio stream corresponding to a low-end encoding method, a medium-end audio stream corresponding to a medium-end encoding method, and a high-end audio stream corresponding to a high-end encoding method from the first electronic device, and the second encoding strategy of the second electronic device indicates the low-end encoding method, the server sends the low-end audio stream to the second electronic device.

[0159] In another example, if the server receives a low-end audio stream corresponding to a low-end encoding method, a mid-end audio stream corresponding to a mid-end encoding method, and a high-end audio stream corresponding to a high-end encoding method from the first electronic device, and the second encoding strategy of the second electronic device indicates a combination of mid-end encoding methods, the server sends a mid-end audio stream to the second electronic device.

[0160] In another example, if the server receives a low-end audio stream corresponding to a low-end encoding method, a mid-end audio stream corresponding to a mid-end encoding method, and a high-end audio stream corresponding to a high-end encoding method from the first electronic device, and the second encoding strategy of the second electronic device indicates the high-end encoding method, the server sends the high-end audio stream to the second electronic device.

[0161] The audio transmission method provided in this application embodiment will be described in detail below with specific examples. Figure 8 This is a schematic diagram illustrating an illustrative application scenario. Please refer to... Figure 8 Terminal A is an uplink user, while terminals B and C are downlink users. The network status between terminal A and the SFU is poor, as is that between terminal B and the SFU, while the network status between terminal C and the SFU is relatively good.

[0162] For example, terminal A is currently sending a low-resolution audio stream to SFU according to the first redundancy depth value.

[0163] SFU sends a low-resolution audio stream to terminal B according to the second redundancy depth value, and sends a low-resolution audio stream to terminal C according to the third redundancy depth value. The third redundancy depth is less than the second redundancy depth value.

[0164] Combination Figure 8 , Figure 9 This is a schematic diagram illustrating an illustrative application scenario. Please refer to... Figure 9 In this scenario, compared to Figure 8 In the application scenario, the network status between terminal A and SFU is good, the network status between terminal B and SFU is poor, and the network status between terminal C and SFU is good.

[0165] For example, based on preset evaluation conditions, terminal A determines that the available bandwidth between itself and the SFU needs to be increased. Accordingly, terminal A... Figure 3 During execution, terminal A's current encoding strategy is the third encoding method (i.e., low-end encoding method). Terminal A adjusts its encoding strategy to the second encoding method (i.e., a combination of low-end and mid-end encoding methods), and increases the bitrate of the mid-end encoding method to the maximum value within the adjustable range. Furthermore, terminal A determines that the redundancy depth remains the first redundancy depth.

[0166] For example, terminal A sends a low-end audio stream and a mid-end audio stream to SFU according to the first redundancy depth.

[0167] Based on preset evaluation conditions, Terminal B determines that the available bandwidth between itself and SFU remains unchanged, and sends an indication message to SFU to indicate the available bandwidth for the next cycle.

[0168] Based on the indication information, the SFU determines that the encoding strategy does not need to be adjusted; that is, the encoding strategy corresponding to terminal B remains the low-level encoding method. Accordingly, the SFU sends the low-level audio bitstream to terminal B according to the current redundancy depth, i.e., the second redundancy depth.

[0169] For terminal C, based on preset evaluation conditions, terminal C determines that the available bandwidth between itself and other terminals needs to be increased, and further determines the available bandwidth for the next cycle. Terminal C sends an indication message to SFU to indicate the available bandwidth for the next cycle.

[0170] Based on the increased available bandwidth, the SFU reallocates the encoding strategy. For example, in this case, the SFU adjusts the encoding strategy of terminal C to a mid-range encoding method and reduces the redundancy depth (currently the third redundancy depth) to the fourth redundancy depth. Accordingly, the SFU sends the mid-range audio bitstream to terminal C according to the fourth redundancy depth.

[0171] Thus, in this example, the user on terminal B can hear smooth audio from terminal A, but with lower sound quality (due to the low bitrate). The user on terminal C, however, can hear smooth audio from terminal A with better sound quality (due to the mid-range bitrate).

[0172] Combination Figure 9 , Figure 10 This is a schematic diagram illustrating an illustrative application scenario. Please refer to... Figure 10 In this scenario, compared to Figure 9 In the application scenario shown, the network status between terminal A and SFU is excellent, the network status between terminal B and SFU is poor, the network status between terminal C and SFU is good, and the network status between terminal D and SFU is excellent.

[0173] For example, terminal A assesses the network status based on preset evaluation conditions and determines to increase the available bandwidth. Based on the available bandwidth, terminal A determines the encoding strategy to be the third encoding method, which includes low-end encoding, mid-end encoding, and high-end encoding. The selection of the bit rate can be referred to above and will not be repeated here. The redundancy depth remains the current redundancy depth.

[0174] For example, terminal A sends low-end audio streams, mid-end audio streams, and high-end audio streams to SFU according to the first redundancy depth.

[0175] Terminals B, C, and D send the available bandwidth for the next cycle to the SFU based on preset evaluation conditions.

[0176] Based on the available bandwidth of terminal B, SFU determines that the encoding strategy of terminal B remains unchanged and is still the low-end encoding method.

[0177] Based on the available bandwidth of terminal C, SFU determines that the encoding strategy of terminal C remains unchanged and is still a mid-range encoding method.

[0178] Based on the available bandwidth of terminal D, SFU determines that the encoding strategy of terminal D is high-end encoding.

[0179] For example, the SFU sends a low-end audio stream to terminal B according to the first redundancy depth. The SFU sends a mid-end audio stream to terminal C according to the fourth redundancy depth. The SFU sends a high-end audio stream to terminal D according to the fifth redundancy depth (which can be the initial default value). In this embodiment, the SFU can be set with redundancy depth values ​​corresponding to different bandwidth gradients. The SFU can determine the corresponding redundancy depth value based on the available bandwidth, which is the initial default value. The way the redundancy depth is determined is only an illustrative example and is not limited in this application.

[0180] The above mainly describes the solution provided by the embodiments of this application from the perspective of interaction between various network elements. It is understood that, in order to achieve the above functions, the audio transmission device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0181] This application embodiment can divide the audio transmission device into functional modules according to the above method example. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0182] When each function is divided into its own modules, the remaining modules are defined according to their respective functions. Figure 11 A possible structural schematic diagram of the audio transmission device 1100 involved in the above embodiments is shown, such as... Figure 11As shown, the audio transmission device may include: an evaluation module 1101, a processing module 1102, an encoding module 1103, and a transmission module 1104. Evaluation module 1101 is used to determine the first available bandwidth between the server and the server based on preset evaluation conditions; the first available bandwidth is not equal to the current available bandwidth; processing module 1102 is used to determine the first encoding strategy based on the first available bandwidth; the first encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method; wherein, the first encoding method includes a low-end encoding method, the second encoding method includes a low-end encoding method and a mid-end encoding method, and the third encoding method includes a low-end encoding method, a mid-end encoding method, and a high-end encoding method; the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; encoding module 1103 is used to encode the first audio data based on the first encoding strategy to obtain N audio bitstreams; N is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the N audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; transmission module 1104 is used to send the N audio bitstreams to the server, so that the server sends the first target audio bitstream among the N audio bitstreams to the second electronic device.

[0183] Figure 9 A possible structural diagram of the server 1200 involved in the above embodiments is shown, such as... Figure 12 As shown, the server may include: an acquisition module 1201, a processing module 1202, and a transmission module 1203. The acquisition module 1201 is used to acquire a first available bandwidth between itself and the second electronic device; the processing module 1202 is used to determine a first encoding strategy corresponding to the second electronic device based on the first available bandwidth; the first encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method, wherein the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; the transmission module 1203 is used to receive N audio bitstreams sent by the first electronic device, where N is a positive integer greater than 0 and less than 4; wherein a single bitstream among the N audio bitstreams corresponds to a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the transmission module 1203 is also used to send a first target audio bitstream among the N audio bitstreams to the second electronic device based on the first encoding strategy.

[0184] In another example, Figure 13 The schematic block diagram illustrating an audio transmission device 1300 according to an embodiment of this application shows that the audio transmission device may include a processor 1301 and a transceiver / transceiver pin 1302, and optionally, a memory 1303. The processor 1301 can be used to execute the steps performed by the audio transmission device in the methods of the foregoing embodiments, and to control the receive pin to receive signals and control the transmit pin to transmit signals.

[0185] The various components of the audio transmission device 1300 are coupled together via a bus 1304, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 1304 in the figure.

[0186] Optionally, the memory 1303 can be used for storage instructions in the foregoing method embodiments.

[0187] It should be understood that the audio transmission device 1300 according to the embodiments of this application may correspond to the electronic device or server in the methods of the foregoing embodiments, and the above and other management operations and / or functions of each element in the audio transmission device 1300 are respectively for implementing the corresponding steps of the foregoing methods, which will not be described in detail here for the sake of brevity.

[0188] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0189] Based on the same technical concept, embodiments of this application also provide a computer-readable storage medium storing a computer program containing at least one piece of code that can be executed by an audio transmission device to control the audio transmission device to implement the above-described method embodiments.

[0190] Based on the same technical concept, this application also provides a computer program, which, when executed by an audio transmission device, is used to implement the above-described method embodiments.

[0191] The program may be stored, in whole or in part, on a storage medium packaged with the processor, or in part or in whole on a memory not packaged with the processor.

[0192] Based on the same technical concept, this application also provides a processor for implementing the above-described method embodiments. The processor can be a chip.

[0193] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0194] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0195] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio transmission method, characterized in that, Applied to a first electronic device, the method includes: Based on preset evaluation conditions, a first available bandwidth between the server and the server is determined; the first available bandwidth is not equal to the currently available bandwidth. Based on the first available bandwidth, a first encoding strategy is determined; the first encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method; wherein, the first encoding method includes a low-end encoding method, the second encoding method includes the low-end encoding method and a mid-end encoding method, and the third encoding method includes the low-end encoding method, the mid-end encoding method, and a high-end encoding method; the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; Based on the first encoding strategy, the first audio data is encoded to obtain N audio bitstreams; N is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the N audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; The N audio streams are sent to the server, which then sends the first target audio stream from the N audio streams to the second electronic device.

2. The method according to claim 1, characterized in that, The low-level encoding method in the first encoding strategy corresponds to the first bit rate, and the first bit rate belongs to the first bit rate range; If the first encoding strategy includes a mid-range encoding method, the mid-range encoding method in the first encoding strategy corresponds to the second bit rate, and the second bit rate belongs to the range of the second bit rate; If the first encoding strategy includes a high-level encoding method, the high-level encoding method in the first encoding strategy corresponds to a third code rate, and the third code rate belongs to the third code rate range; Wherein, the minimum value of the third bit rate range is greater than the maximum value of the second bit rate range, and the minimum value of the second bit rate range is greater than the maximum value of the first bit rate range.

3. The method according to claim 2, characterized in that, The process of encoding the first audio data based on the first encoding strategy to obtain N audio bitstreams includes: If the first encoding strategy is the first encoding method, the first audio data is encoded based on the low-level encoding method to obtain the first audio bitstream; If the first encoding strategy is the second encoding method, the first audio data is encoded based on the low-end encoding method to obtain a first audio bitstream; and the first audio data is encoded based on the mid-end encoding method to obtain a second audio bitstream. If the first encoding strategy is the third encoding method, the first audio data is encoded based on the low-end encoding method to obtain a first audio bitstream; and the first audio data is encoded based on the mid-end encoding method to obtain a second audio bitstream; and the first audio data is encoded based on the high-end encoding method to obtain a third audio bitstream.

4. The method according to claim 2, characterized in that, The method further includes: At a preset periodic trigger time, based on the preset evaluation conditions, a second available bandwidth between the server and the server is determined; the second available bandwidth is not equal to the first available bandwidth. Based on the second available bandwidth, a second encoding strategy is determined; the second encoding strategy includes one of the first encoding method, the second encoding method, or the third encoding method, and the second encoding strategy is different from the first encoding strategy. Based on the second encoding strategy, the second audio data is encoded to obtain P audio bitstreams; P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; The P audio streams are sent to the server, which then sends the second target audio stream from the P audio streams to the second electronic device.

5. The method according to claim 4, characterized in that, The first encoding strategy being different from the second encoding strategy includes: the first encoding strategy and the second encoding strategy including different types of the first encoding method, the second encoding method, or the third encoding method; and / or, the first encoding strategy and the second encoding strategy including different bit rates of the first encoding method, the second encoding method, or the third encoding method.

6. The method according to claim 4, characterized in that, The low-level encoding method in the second encoding strategy corresponds to the fourth code rate, which is within the range of the first code rate; If the second encoding strategy includes a mid-range encoding method, the mid-range encoding method in the second encoding strategy corresponds to the fifth code rate, and the fifth code rate belongs to the range of the second code rate; If the second encoding strategy includes a high-level encoding method, the high-level encoding method in the second encoding strategy corresponds to the sixth code rate, and the sixth code rate belongs to the third code rate range.

7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the first encoding strategy based on the first available bandwidth further includes: Based on the first available bandwidth, a redundancy depth value is determined; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the first electronic device to the server; Accordingly, sending the N audio streams to the server includes: The N audio streams are sent to the server according to the redundancy depth value.

8. The method according to any one of claims 1 to 6, characterized in that, The bandwidth occupied by the N audio streams is less than or equal to the first available bandwidth.

9. An audio transmission method, characterized in that, Applied to a server, the method includes: Obtain the first available bandwidth between the device and the second electronic device; Based on the first available bandwidth, a first encoding strategy corresponding to the second electronic device is determined; the first encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method, wherein the bit rates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; The system receives N audio streams sent by a first electronic device, where N is a positive integer greater than 0 and less than 4; wherein, a single stream among the N audio streams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method. Based on the first encoding strategy, the first target audio stream among the N audio streams is sent to the second electronic device.

10. The method according to claim 9, characterized in that, The method further includes: At a preset periodic trigger time, the second available bandwidth between the device and the second electronic device is obtained; Based on the second available bandwidth, a second encoding strategy corresponding to the second electronic device is determined; the second encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the second encoding strategy is different from the first encoding strategy. Receive P audio bitstreams sent by the first electronic device, where P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; Based on the second encoding strategy, the second target audio stream from the P audio streams is sent to the second electronic device.

11. The method according to claim 9 or 10, characterized in that, The step of determining the first encoding strategy corresponding to the second electronic device based on the first available bandwidth includes: Based on the first available bandwidth, a redundancy depth value is determined; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the server to the second electronic device; Accordingly, sending the first target audio stream from the N audio streams to the second electronic device based on the first encoding strategy includes: Based on the first encoding strategy and the redundancy depth value, the first target audio stream among the N audio streams is sent to the second electronic device.

12. The method according to claim 9 or 10, characterized in that, The bandwidth occupied by the first target audio bitstream is less than or equal to the first available bandwidth.

13. An audio transmission device, characterized in that, Applied to a first electronic device, the device includes: An evaluation module is used to determine a first available bandwidth with the server based on preset evaluation conditions; the first available bandwidth is not equal to the currently available bandwidth. The processing module is configured to determine a first encoding strategy based on the first available bandwidth; the first encoding strategy includes one of a first encoding method, a second encoding method, or a third encoding method; wherein the first encoding method includes a low-end encoding method, the second encoding method includes the low-end encoding method and a mid-end encoding method, and the third encoding method includes the low-end encoding method, the mid-end encoding method, and a high-end encoding method; the bitrates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; The encoding module is used to encode the first audio data based on the first encoding strategy to obtain N audio bitstreams; N is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the N audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; A transmission module is used to send the N audio streams to the server, so that the server sends the first target audio stream among the N audio streams to the second electronic device.

14. The apparatus according to claim 13, characterized in that, The low-level encoding method in the first encoding strategy corresponds to the first bit rate, and the first bit rate belongs to the first bit rate range; If the first encoding strategy includes a mid-range encoding method, the mid-range encoding method in the first encoding strategy corresponds to the second bit rate, and the second bit rate belongs to the range of the second bit rate; If the first encoding strategy includes a high-level encoding method, the high-level encoding method in the first encoding strategy corresponds to a third code rate, and the third code rate belongs to the third code rate range; Wherein, the minimum value of the third bit rate range is greater than the maximum value of the second bit rate range, and the minimum value of the second bit rate range is greater than the maximum value of the first bit rate range.

15. The apparatus according to claim 14, characterized in that, The encoding module is specifically used for: If the first encoding strategy is the first encoding method, the first audio data is encoded based on the low-level encoding method to obtain the first audio bitstream; If the first encoding strategy is the second encoding method, the first audio data is encoded based on the low-level encoding method to obtain the first audio bitstream; Furthermore, based on the aforementioned mid-range encoding method, the first audio data is encoded to obtain a second audio bitstream; If the first encoding strategy is the third encoding method, the first audio data is encoded based on the low-level encoding method to obtain the first audio bitstream; Furthermore, based on the aforementioned mid-range encoding method, the first audio data is encoded to obtain the second audio bitstream; Furthermore, based on the aforementioned high-end encoding method, the first audio data is encoded to obtain a third audio bitstream.

16. The apparatus according to claim 14, characterized in that, The evaluation module is further configured to determine a second available bandwidth with the server based on the preset evaluation conditions at a preset periodic trigger time; the second available bandwidth is not equal to the first available bandwidth; The processing module is further configured to determine a second encoding strategy based on the second available bandwidth; the second encoding strategy includes one of the first encoding method, the second encoding method, or the third encoding method, and the second encoding strategy is different from the first encoding strategy; The encoding module is further configured to encode the second audio data based on the second encoding strategy to obtain P audio bitstreams; P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; The transmission module is further configured to send the P audio streams to the server, so that the server sends the second target audio stream from the P audio streams to the second electronic device.

17. The apparatus according to claim 16, characterized in that, The first encoding strategy being different from the second encoding strategy includes: the first encoding strategy and the second encoding strategy including different types of the first encoding method, the second encoding method, or the third encoding method; and / or, the first encoding strategy and the second encoding strategy including different bit rates of the first encoding method, the second encoding method, or the third encoding method.

18. The apparatus according to claim 16, characterized in that, The low-level encoding method in the second encoding strategy corresponds to the fourth code rate, which is within the range of the first code rate; If the second encoding strategy includes a mid-range encoding method, the mid-range encoding method in the second encoding strategy corresponds to the fifth code rate, and the fifth code rate belongs to the range of the second code rate; If the second encoding strategy includes a high-level encoding method, the high-level encoding method in the second encoding strategy corresponds to the sixth code rate, and the sixth code rate belongs to the third code rate range.

19. The apparatus according to any one of claims 13 to 18, characterized in that, The processing module is further configured to: Based on the first available bandwidth, a redundancy depth value is determined; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the first electronic device to the server; The transmission module is further configured to send the N audio streams to the server according to the redundancy depth value.

20. The apparatus according to any one of claims 13 to 18, characterized in that, The bandwidth occupied by the N audio streams is less than or equal to the first available bandwidth.

21. An audio transmission device, characterized in that, Applied to a server, the device includes: The acquisition module acquires the first available bandwidth between itself and the second electronic device; The processing module is configured to determine a first encoding strategy corresponding to the second electronic device based on the first available bandwidth; the first encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method, wherein the bit rates of the low-end encoding method, the mid-end encoding method, and the high-end encoding method are different; The transmission module is used to receive N audio bitstreams sent by the first electronic device, where N is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the N audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method; The transmission module is further configured to send the first target audio stream among the N audio streams to the second electronic device based on the first encoding strategy.

22. The apparatus according to claim 21, characterized in that, The device further includes: The acquisition module is further configured to acquire the second available bandwidth between itself and the second electronic device at a preset periodic trigger time. The processing module is further configured to determine a second encoding strategy corresponding to the second electronic device based on the second available bandwidth; the second encoding strategy includes one of a low-end encoding method, a mid-end encoding method, or a high-end encoding method; the second encoding strategy is different from the first encoding strategy; The transmission module is further configured to receive P audio bitstreams sent by the first electronic device, where P is a positive integer greater than 0 and less than 4; wherein, a single bitstream among the P audio bitstreams corresponds to the low-end encoding method, the mid-end encoding method, or the high-end encoding method. The transmission module is further configured to send the second target audio stream from the P audio streams to the second electronic device based on the second encoding strategy.

23. The apparatus according to claim 21 or 22, characterized in that, The processing module is further configured to: Based on the first available bandwidth, a redundancy depth value is determined; the redundancy depth value is used to indicate the amount of redundancy in the data packets sent by the server to the second electronic device; The transmission module is further configured to send the first target audio stream among the N audio streams to the second electronic device based on the first encoding strategy and the redundancy depth value.

24. The apparatus according to claim 21 or 22, characterized in that, The bandwidth occupied by the first target audio bitstream is less than or equal to the first available bandwidth.

25. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-8, or cause the electronic device to perform the method as described in any one of claims 9-12.

26. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-8, or causes the computer to perform the method as described in any one of claims 9-12.

27. A chip, characterized in that, The device includes one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of the electronic device and send the signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the method of any one of claims 1-8, or the electronic device performs the method of any one of claims 9-12.

Citation Information

Patent Citations

  • Video code stream transmission method and device

    CN111836079A

  • Voice processing method and device, equipment, storage medium and program product

    CN115426342A