An audio transmission method, an audio transmission device and a storage medium
By adjusting the working modes of the sending and receiving ends in the audio transmission system according to network parameters, the problem of low audio data intelligibility when the network is unstable is solved, and the intelligibility and stability of audio data transmission are improved when the network parameters are poor.
Patent Information
- Application Number
- CN202610729517.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-23
AI Technical Summary
When the network is unstable, the intelligibility of audio data transmitted through the audio transmission system is low.
By controlling the switching of the working modes of the sending and receiving ends, and adjusting the encoding and decoding processes according to network parameters, the mode of the sending and receiving ends is kept consistent to form data packets that match the network parameters.
When network parameters are poor, it improves the intelligibility of audio data and ensures the stability and reliability of audio data transmission.
Smart Images

Figure CN122266375A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio transmission technology, and in particular to an audio transmission method, an audio transmission device, and a storage medium. Background Technology
[0002] With the development of technology, audio transmission systems, capable of transmitting audio data, have been widely adopted. Typically, an audio transmission system includes a transmitter, a receiver, and a broadcaster. The transmitter is responsible for collecting audio data and sending it to the receiver. The receiver then decodes the audio data and transmits the decoded data to the broadcaster, which then plays it back. However, in these technologies, the intelligibility of audio data transmitted through audio transmission systems is low when the network is unstable. Summary of the Invention
[0003] This application aims to provide an audio transmission method, audio transmission device, and storage medium, which at least solves the problem of low intelligibility of audio data transmitted through an audio transmission system when the network is unstable.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide an audio transmission method applied to an audio transmission system, the audio transmission system including a transmitting end and a receiving end, and the audio transmission system being connected to a network to transmit data via the network, the audio transmission method comprising: The transmitting end is controlled to acquire audio data and encode the audio data to form an encoded frame; The parameters of the network are determined, and the transmitting end is controlled to switch to the corresponding working mode according to the parameters of the network. The working mode includes a first mode and a second mode. The first mode is different from the second mode. The parameters of the network include the network speed change rate and bandwidth. The coded frames are packaged into data packets according to the current working mode of the sending end, and the sending end is controlled to send the data packets to the receiving end. The receiving end is controlled to switch to the target mode, so that the receiving end receives the data packet and decodes the data packet according to the target mode in which the receiving end is located, to obtain decoded data, and then sends the decoded data outward. The target mode is the same as the working mode currently in which the sending end is located.
[0005] Optionally, determining the parameters of the network and controlling the transmitter to switch to the corresponding operating mode based on the network parameters includes: Determine the network speed change rate and bandwidth; If the network speed change rate is less than or equal to the network speed change threshold, and the bandwidth is less than or equal to the set bandwidth threshold, the sending end is controlled to switch to the first mode. If the network speed change rate is greater than the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold, the sending end is controlled to switch to the second mode.
[0006] Optionally, the step of packaging the encoded frame into a data packet according to the current operating mode of the sending end, and controlling the sending end to send the data packet to the receiving end, includes: When the sending end is currently in the first mode, N consecutive encoded frames are packaged into the data packet, and the sending end is controlled to send the data packet to the receiving end; When the sending end is currently in the second mode, the current encoded frame and the encoded frames within a set time period prior to the current time are combined and packaged into the data packet, and the sending end is controlled to send the data packet to the receiving end.
[0007] Optionally, controlling the receiving end to switch to the target mode, enabling the receiving end to receive the data packet, and decoding the data packet according to the target mode of the receiving end to obtain decoded data, includes: When the sending end is in the first mode, the receiving end is controlled to switch to the first mode, so that the receiving end receives the data packet and directly decodes the data packet based on the first mode to obtain the decoded data; When the sending end is in the second mode, the receiving end is controlled to switch to the second mode, so that the receiving end receives the data packet and decodes the data packet based on the second mode to obtain the decoded data.
[0008] Optionally, the decoded data is obtained by decoding the data packet based on the second mode, including: Based on the second mode, the data packets are split according to a set duration to obtain multiple first split frame data; The sequence numbers of the multiple first split frame data are reassigned, and the multiple first split frame data are used to generate multiple new data packets, each of which corresponds to a sequence number; The multiple new data packets are sorted according to their sequence numbers to obtain a sorted sequence of new data packets; Remove duplicate new data packets from the sorted sequence of new data packets; The sequence of new data packets after removing duplicates is decoded to obtain the decoded data.
[0009] Optionally, the data packet is directly decoded based on the first mode to obtain the decoded data, including: Based on the first mode, the data packet is split into multiple second split frame data, and the multiple second split frame data are decoded to obtain the decoded data.
[0010] Optionally, controlling the transmitting end to acquire audio data includes: The transmitting end is controlled to collect audio data at a set acquisition period.
[0011] Optionally, the audio data is encoded, including: The audio data is encoded according to a target coding rate, wherein the target coding rate is less than or equal to a set coding rate threshold, and the coding rate threshold is less than or equal to 3kbps.
[0012] Secondly, embodiments of this application provide an audio transmission device applied to an audio transmission system, the audio transmission system including a transmitting end and a receiving end, and the audio transmission system connected to a network to transmit data through the network, the audio transmission device comprising: The first control module is used to control the transmitting end to collect audio data and encode the audio data to form an encoded frame; The determining module is used to determine the parameters of the network and control the transmitting end to switch to the corresponding working mode according to the parameters of the network. The working mode includes a first mode and a second mode, the first mode being different from the second mode. The parameters of the network include the network speed change rate and bandwidth. The packetization module is used to package the encoded frame into a data packet according to the current working mode of the sending end, and control the sending end to send the data packet to the receiving end; The second control module is used to control the receiving end to switch to the target mode, so that the receiving end receives the data packet, decodes the data packet according to the target mode of the receiving end, obtains decoded data, and sends the decoded data outward. The target mode is the same as the working mode of the sending end.
[0013] Optionally, the determining module includes: The first determining unit is used to determine the network speed change rate and bandwidth of the network; The first control unit is configured to control the transmitting end to switch to the first mode when the network speed change rate is less than or equal to the network speed change threshold and the bandwidth is less than or equal to a set bandwidth threshold. The second control unit is used to control the transmitting end to switch to the second mode when the network speed change rate is greater than the network speed change threshold and the bandwidth is less than or equal to a set bandwidth threshold.
[0014] Optionally, the packaging module includes: The first packetizing unit is configured to, when the sending end is currently in the first mode, package N consecutive encoded frames into the data packet, and control the sending end to send the data packet to the receiving end; The second packetizing unit is used to, when the sending end is currently in the second mode, combine the current encoded frame and the encoded frames within a set time period before the current time into the data packet, and control the sending end to send the data packet to the receiving end.
[0015] Optionally, the second control module includes: The third control unit is configured to control the receiving end to switch to the first mode when the sending end is in the first mode, so that the receiving end receives the data packet and directly decodes the data packet based on the first mode to obtain the decoded data. The fourth control unit is configured to control the receiving end to switch to the second mode when the sending end is in the second mode, so that the receiving end receives the data packet and decodes the data packet based on the second mode to obtain the decoded data.
[0016] Optionally, the fourth control unit is further configured to: Based on the second mode, the data packets are split according to a set duration to obtain multiple first split frame data; The sequence numbers of the multiple first split frame data are reassigned, and the multiple first split frame data are used to generate multiple new data packets, each of which corresponds to a sequence number; The multiple new data packets are sorted according to their sequence numbers to obtain a sorted sequence of new data packets; Remove duplicate new data packets from the sorted sequence of new data packets; The sequence of new data packets after removing duplicates is decoded to obtain the decoded data.
[0017] Optionally, the third control unit is further configured to: Based on the first mode, the data packet is split into multiple second split frame data, and the multiple second split frame data are decoded to obtain the decoded data.
[0018] Optionally, the first control module is also used for: The transmitting end is controlled to collect audio data at a set acquisition period.
[0019] Optionally, the first control module is also used for: The audio data is encoded according to a target coding rate, wherein the target coding rate is less than or equal to a set coding rate threshold, and the coding rate threshold is less than or equal to 3kbps.
[0020] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0021] Fourthly, embodiments of this application provide a storage medium storing a program or instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0022] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0023] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0024] In this embodiment, the transmitting end is controlled to collect audio data and encode the audio data to form encoded frames; network parameters are determined, and the transmitting end is controlled to switch to the corresponding working mode according to the network parameters. The working modes include a first mode and a second mode, which are different from the second mode. The network parameters include the network speed change rate and bandwidth; the encoded frames are packaged into data packets according to the current working mode of the transmitting end, and the transmitting end is controlled to send the data packets to the receiving end; the receiving end is controlled to switch to the target mode, so that the receiving end receives the data packets, and decodes the data packets according to the target mode of the receiving end to obtain decoded data, and sends the decoded data outward. The target mode is the same as the current working mode of the transmitting end. In other words, in this embodiment, the working mode of the transmitting end can be switched according to the network parameters, and the working mode of the receiving end can be switched to the same working mode as the transmitting end. Thus, when decoding the data packets received by the receiving end to obtain decoded data, it is equivalent to associating the decoded data with the network parameters, thereby effectively avoiding the problem of poor network parameters leading to low intelligibility of transmitted audio data. That is, in this application, the intelligibility of transmitted audio data can be improved when the network parameters are poor. Attached Figure Description
[0025] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an audio transmission method provided in an embodiment of this application; Figure 2 This diagram illustrates an audio transmission device provided in an embodiment of this application. Figure 3 This is a schematic diagram illustrating an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are only used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0027] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0028] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, "at least one of a, b, or c" can represent: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0029] This application provides an audio transmission method applied to an audio transmission system. The audio transmission system includes a transmitter and a receiver, and is connected to a network to transmit data, such as... Figure 1 As shown, the audio transmission method includes: Step 101: Control the sending end to collect audio data and encode the audio data to form an encoded frame.
[0030] The transmitting end is equipped with an audio acquisition element. By controlling the transmitting end, the audio acquisition element can collect audio data. The audio acquisition element includes, but is not limited to, a sound card and a microphone.
[0031] In addition, in this embodiment of the application, the audio data includes, but is not limited to, the user's voice data, ambient sounds, etc.
[0032] In addition, in some implementations, the method of controlling the sending end to collect audio data can be: controlling the sending end to collect audio data at a set collection period.
[0033] In controlling the transmitting end to collect audio data, the transmitting end can be made to collect audio data at a set collection period, that is, the audio collection element in the transmitting end collects audio data according to the set collection period.
[0034] It should be noted that the acquisition period can be set according to actual needs. For example, if the acquisition period is set to 40ms, the transmitting end will acquire audio data at a period of 40ms, that is, the transmitting end will acquire audio data once every 40ms. As another example, if the acquisition period is set to 20ms, the transmitting end will acquire audio data at a period of 20ms, that is, the transmitting end will acquire audio data once every 20ms. The specific value of the acquisition period is not limited in this embodiment.
[0035] In addition, in some implementations, the audio data can be encoded by encoding the audio data according to a target bitrate, where the target bitrate is less than or equal to a set bitrate threshold, and the bitrate threshold is less than or equal to 3kbps.
[0036] In the process of encoding audio data, the audio data is encoded according to the target code rate, that is, each frame of audio data is encoded according to the target code rate. After encoding each frame of audio data, it is equivalent to outputting N bytes of data for each frame of audio data.
[0037] It should be noted that the target coding rate can be any value less than or equal to 3 kbps. For example, a target coding rate of 1.6 kbps is equivalent to encoding each frame of audio data at 1.6 kbps, resulting in an output of 8 bytes of data per frame. Similarly, a target coding rate of 2.4 kbps is equivalent to encoding each frame of audio data at 2.4 kbps, resulting in an output of 6 bytes of data per frame. The specific value of the target coding rate is not limited in this embodiment.
[0038] Step 102: Determine the network parameters and, based on the network parameters, control the sending end to switch to the corresponding working mode. The working modes include the first mode and the second mode. The first mode is different from the second mode. The network parameters include the network speed change rate and bandwidth.
[0039] The audio transmission system connects to the network, allowing for real-time network monitoring and determination of network parameters, such as network speed variation rate and bandwidth. Once these parameters are determined, the transmitting end can be controlled to switch to the appropriate operating mode. Different network parameters result in different operating modes for the transmitting end, enabling it to selectively switch modes based on these parameters and facilitate subsequent data transmission according to these different modes.
[0040] In addition, in some implementations, determining network parameters and controlling the transmitter to switch to the corresponding working mode based on these parameters can be achieved by: determining the network speed change rate and bandwidth; controlling the transmitter to switch to the first mode when the network speed change rate is less than or equal to the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold; and controlling the transmitter to switch to the second mode when the network speed change rate is greater than the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold.
[0041] Among them, the network speed change rate can reflect the stability of the network. When the network speed change rate is large, it indicates that the network is unstable and is prone to problems such as packet loss and delay when transmitting audio data over the network. When the network speed change rate is small, it indicates that the network is relatively stable and is less prone to problems such as packet loss and delay when transmitting audio data over the network.
[0042] In addition, once the audio transmission system connects to the network, it can monitor network parameters to determine these parameters, namely the rate of change of network speed and bandwidth. When the rate of change of network speed is less than or equal to the network speed change threshold, and the bandwidth is less than or equal to the set bandwidth threshold, it indicates that the audio transmission system is currently performing narrowband transmission and the network is relatively stable. In this case, the transmitting end can be controlled to switch to the first mode, which can be narrowband mode, allowing the transmitting end to send audio data in narrowband mode. When the rate of change of network speed is greater than the rate of change of network speed, and the bandwidth is less than or equal to the bandwidth threshold, it indicates that the audio transmission system is currently performing narrowband transmission, but the network is unstable and prone to problems such as packet loss and latency. In this case, the transmitting end can be controlled to switch to the second mode, which can be a weak network mode, allowing the transmitting end to send audio data in a weak network mode.
[0043] It should be noted that the specific value of the network speed change rate can be set according to actual needs. For example, a network speed change rate of 30% means that the network transmission speed at the current moment has changed by 30% compared to the network transmission speed at the previous moment. In other words, when the network speed change exceeds 30%, the network is considered unstable. Specifically, if the network transmission speed at the current moment is 13 and the network transmission speed at the previous moment was 10, the network speed change rate is 30%. As another example, a network speed change rate of 25% means that the network transmission speed at the current moment has changed by 25% compared to the network transmission speed at the previous moment. In other words, when the network speed change exceeds 25%, the network is considered unstable. The specific value of the network speed change rate is not limited in this embodiment.
[0044] Furthermore, in this embodiment, the bandwidth threshold can be set according to actual needs, for example, a bandwidth threshold of 16kbps, or even 15kbps. The specific value of the bandwidth threshold is not limited in this embodiment.
[0045] Step 103: Pack the encoded frames into data packets according to the current working mode of the sending end, and control the sending end to send the data packets to the receiving end.
[0046] When the network parameters are different, the sending end will be controlled to switch to different working modes. When the sending end needs to send data, it can package the encoded frame into a data packet according to the current working mode of the sending end, so that the packaged data packet corresponds to the current working mode of the sending end, and then control the sending end to send the data packet to the receiving end.
[0047] In some implementations, step 103 can be implemented as follows: when the sending end is currently in the first mode, N consecutive encoded frames are packaged into a data packet, and the sending end is controlled to send the data packet to the receiving end; when the sending end is currently in the second mode, the current encoded frame and encoded frames within a set time period before the current time are combined and packaged into a data packet, and the sending end is controlled to send the data packet to the receiving end. Where N ≥ 2, and N is a positive integer.
[0048] Specifically, when packaging encoded frames into data packets according to the current working mode of the sending end, if the sending end is currently in the first mode, it packages N consecutive encoded frames into a data packet; if the sending end is currently in the second mode, it packages the current encoded frame and the encoded frames within a set time period before the current time into a data packet. That is, if the sending end is currently in the second mode, it not only packages the current encoded frame, but also packages the encoded frames within a set time period before the current time. These encoded frames are historical encoded frames relative to the current encoded frame. In other words, if the sending end is in the second mode, it packages the current encoded frame and the historical encoded frames into a data packet.
[0049] It should be noted that the set time period is the time period before and adjacent to the current time. For example, setting the time period to 400ms is equivalent to combining the current encoded frame at the current time with the encoded frames within 400ms before the current time. The set time period can be set according to actual needs; for example, a time period of 400ms, or even 600ms. This embodiment of the application does not limit this. Furthermore, the specific value of N can be set according to actual needs. For example, N=6, which is equivalent to packaging 6 consecutive encoded frames into a data packet when the sending end is currently in the first mode; or N=13, which is equivalent to packaging 13 consecutive encoded frames into a data packet when the sending end is currently in the first mode.
[0050] Furthermore, when the sending end is currently in the second mode, when combining the current encoded frame with encoded frames within a set time period prior to the current time into a data packet, it's equivalent to updating the data with each packetization process—adding new data and deleting old data—thus creating data packets with redundant information. The resulting data packets are Real-Time Transport Protocol (RTP) packets.
[0051] For example, when the transmitting end collects audio data at 80ms intervals, with a set time period of 400ms, in the second mode, it can combine the current 80ms encoded frame with five historical encoded frames from the last 400ms. This is equivalent to combining the current 80ms encoded frame, the first 80ms-previous encoded frame, the second 80ms-previous encoded frame, the third 80ms-previous encoded frame, the fourth 80ms-previous encoded frame, and the fifth 80ms-previous encoded frame into a data packet. This packaging method updates the data with each packet, adding new data and deleting old data. Furthermore, each data packet contains redundant data.
[0052] In addition, in this embodiment of the application, when the sending end sends a data packet to the receiving end, the sending end marks the data packet with a sequence number each time it sends a data packet, and then the sending end sends the data packet carrying the sequence number and timestamp to the receiving end.
[0053] Step 104: Control the receiving end to switch to the target mode, so that the receiving end receives data packets and decodes the data packets according to the target mode in which the receiving end is located, obtains decoded data, and sends the decoded data outward. The target mode is the same as the working mode in which the sending end is currently located.
[0054] Once the sending end's working mode is switched, in order to ensure that the receiving end's working mode corresponds to the sending end's working mode and facilitate the receiving end to receive data packets of the corresponding mode, the receiving end can be controlled to switch to the same working mode as the sending end. Then, according to the receiving end's working mode, the data packet is decoded to obtain decoded data, and the decoded data is sent out.
[0055] It should be noted that the audio transmission system includes a broadcasting end, which can send decoded data to the broadcasting end, allowing the broadcasting end to play the decoded data.
[0056] In addition, in some implementations, controlling the receiving end to switch to the target mode, enabling the receiving end to receive data packets, and decoding the data packets according to the target mode in which the receiving end is located, can be implemented as follows: when the sending end is in the first mode, controlling the receiving end to switch to the first mode, enabling the receiving end to receive data packets, and directly decoding the data packets based on the first mode to obtain the decoded data; when the sending end is in the second mode, controlling the receiving end to switch to the second mode, enabling the receiving end to receive data packets, and decoding the data packets based on the second mode to obtain the decoded data.
[0057] Specifically, when the receiving end switches to the target mode to receive data packets, if the sending end is in the first mode, the receiving end is controlled to switch to the first mode to receive data packets. Based on the first mode, the received data packets are directly decoded to obtain decoded data. If the sending end is in the second mode, the receiving end is controlled to switch to the second mode to receive data packets. Based on the second mode, the received data packets are decoded to obtain decoded data. In other words, the receiving end is controlled to switch to the same working mode as the sending end to receive data, and subsequently, the received data packets can be decoded based on the receiving end's working mode to obtain decoded data.
[0058] In some implementations, the method for obtaining decoded data based on the second mode can be as follows: split the data packets according to a set duration based on the second mode to obtain multiple first split frame data; reassign sequence numbers to the multiple first split frame data and generate multiple new data packets, each new data packet corresponding to a sequence number; sort the multiple new data packets according to the sequence numbers to obtain a sorted new data packet sequence; delete duplicate new data packets in the sorted new data packet sequence; and decode the new data packet sequence after deleting duplicate new data packets to obtain decoded data.
[0059] In this embodiment, when the sending end is in the second mode, the receiving end is switched to the second mode and will split data packets according to a set duration based on the second mode. The set duration can be set according to actual needs. For example, a set duration of 40ms means that data packets are split every 40ms, resulting in multiple first split frame data packets. Alternatively, a set duration of 80ms means that data packets are split every 80ms, resulting in multiple first split frame data packets. The specific value of the set duration is not limited in this embodiment.
[0060] Once a data packet is split into multiple first-segment frames, sequence numbers are reassigned to each of these first-segment frames, resulting in multiple consecutive sequence numbers for each first-segment frame. These reassigned sequence numbers are then used to generate multiple new data packets, with each first-segment frame generating a new data packet, also with a new sequence number. These new data packets are then sorted according to their sequence numbers to obtain a sorted sequence of new data packets. Typically, when decoding a data packet, it is first buffered. This buffer will contain multiple sorted sequences of new data packets, some of which may contain duplicates. These duplicates can be removed from the sorted sequences, and the resulting sequence of new data packets is then decoded to obtain the decoded data.
[0061] This configuration effectively prevents redundant information from being transmitted in the second mode even when the network is unstable and experiences packet loss or delay. The redundant data packets are split into multiple new packets, their sequence numbers are reassigned, and then these reassigned packets are enqueued before being passed to the buffer. Duplicate packets are then removed, effectively resolving the issue of packet loss or delay causing transmission failures. This ensures uninterrupted communication even during network instability or high packet loss, maintaining voice intelligibility of over 95%.
[0062] In addition, in some implementations, the method of directly decoding the data packet based on the first mode to obtain the decoded data can be as follows: the data packet is split into multiple second split frame data based on the first mode, and the multiple second split frame data is decoded to obtain the decoded data.
[0063] When the sending end is in the first mode, the receiving end is switched to the first mode. It can then directly split the data packet into multiple second split frame data based on the first mode. Once the data packet is split into multiple second split frame data, it can be directly decoded to obtain the decoded data.
[0064] For example, when the sending end is in the first mode and the receiving end is switched to the first mode, the data packet can be directly split into multiple 40ms frames based on the first mode, and the multiple 40ms frames can be decoded to obtain the decoded data.
[0065] In this embodiment, the transmitting end is controlled to collect audio data and encode the audio data to form encoded frames; network parameters are determined, and the transmitting end is controlled to switch to the corresponding working mode according to the network parameters. The working modes include a first mode and a second mode, which are different from the second mode. The network parameters include the network speed change rate and bandwidth; the encoded frames are packaged into data packets according to the current working mode of the transmitting end, and the transmitting end is controlled to send the data packets to the receiving end; the receiving end is controlled to switch to the target mode, so that the receiving end receives the data packets, and decodes the data packets according to the target mode of the receiving end to obtain decoded data, and sends the decoded data outward. The target mode is the same as the current working mode of the transmitting end. In other words, in this embodiment, the working mode of the transmitting end can be switched according to the network parameters, and the working mode of the receiving end can be switched to the same working mode as the transmitting end. Thus, when decoding the data packets received by the receiving end to obtain decoded data, it is equivalent to associating the decoded data with the network parameters, thereby effectively avoiding the problem of poor network parameters leading to low intelligibility of transmitted audio data. That is, in this application, the intelligibility of transmitted audio data can be improved when the network parameters are poor.
[0066] This application provides an audio transmission device applied to an audio transmission system. The audio transmission system includes a transmitting end and a receiving end, and is connected to a network to transmit data, such as... Figure 2 As shown, the audio transmission device 200 includes: The first control module 201 is used to control the transmitting end to collect audio data and encode the audio data to form an encoded frame; The determining module 202 is used to determine the network parameters and control the transmitting end to switch to the corresponding working mode according to the network parameters. The working modes include a first mode and a second mode. The first mode is different from the second mode. The network parameters include the network speed change rate and bandwidth. Packaging module 203 is used to package the encoded frames into data packets according to the current working mode of the sending end, and control the sending end to send the data packets to the receiving end; The second control module 204 is used to control the receiving end to switch to the target mode, so that the receiving end receives data packets, decodes the data packets according to the target mode in which the receiving end is located, obtains decoded data, and sends the decoded data outward. The target mode is the same as the working mode in which the sending end is currently located.
[0067] Optionally, module 202 is defined, including: The first determining unit is used to determine the network speed change rate and bandwidth; The first control unit is used to control the transmitting end to switch to the first mode when the network speed change rate is less than or equal to the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold. The second control unit is used to control the transmitter to switch to the second mode when the network speed change rate is greater than the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold.
[0068] Optionally, the package module 203 includes: The first packet unit is used to package N consecutive encoded frames into a data packet when the sending end is currently in the first mode, and to control the sending end to send the data packet to the receiving end. The second packetizing unit is used to combine the current encoded frame and the encoded frames within a set time period before the current time into a data packet when the sending end is currently in the second mode, and to control the sending end to send the data packet to the receiving end.
[0069] Optionally, the second control module 204 includes: The third control unit is used to control the receiver to switch to the first mode when the transmitter is in the first mode, so that the receiver can receive data packets and directly decode the data packets based on the first mode to obtain decoded data. The fourth control unit is used to control the receiver to switch to the second mode when the transmitter is in the second mode, so that the receiver can receive data packets and decode the data packets based on the second mode to obtain decoded data.
[0070] Optionally, the fourth control unit is also used for: Based on the second mode, data packets are split according to a set duration to obtain multiple first split frame data; The sequence numbers of multiple first split frame data are reassigned, and multiple new data packets are generated from the multiple first split frame data, with each new data packet corresponding to a sequence number; Sort multiple new data packets according to their sequence numbers to obtain a sorted sequence of new data packets; Remove duplicate new data packets from the sorted new data packet sequence; The sequence of new data packets after removing duplicate data packets is decoded to obtain the decoded data.
[0071] Optionally, the third control unit is also used for: Based on the first mode, the data packet is split into multiple second split frame data, and the multiple second split frame data are decoded to obtain decoded data.
[0072] Optionally, the first control module 201 is also used for: Control the transmitter to collect audio data at a set acquisition period.
[0073] Optionally, the first control module is also used for: The audio data is encoded according to the target coding rate, which is less than or equal to a set coding rate threshold, and the coding rate threshold is less than or equal to 3kbps.
[0074] In this embodiment, the transmitting end is controlled to collect audio data and encode the audio data to form encoded frames; network parameters are determined, and the transmitting end is controlled to switch to the corresponding working mode according to the network parameters. The working modes include a first mode and a second mode, which are different from the second mode. The network parameters include the network speed change rate and bandwidth; the encoded frames are packaged into data packets according to the current working mode of the transmitting end, and the transmitting end is controlled to send the data packets to the receiving end; the receiving end is controlled to switch to the target mode, so that the receiving end receives the data packets, and decodes the data packets according to the target mode of the receiving end to obtain decoded data, and sends the decoded data outward. The target mode is the same as the current working mode of the transmitting end. In other words, in this embodiment, the working mode of the transmitting end can be switched according to the network parameters, and the working mode of the receiving end can be switched to the same working mode as the transmitting end. Thus, when decoding the data packets received by the receiving end to obtain decoded data, it is equivalent to associating the decoded data with the network parameters, thereby effectively avoiding the problem of poor network parameters leading to low intelligibility of transmitted audio data. That is, in this application, the intelligibility of transmitted audio data can be improved when the network parameters are poor.
[0075] The audio transmission device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0076] The audio transmission device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0077] The audio transmission device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0078] Optionally, such as Figure 3 As shown, this application embodiment also provides an electronic device 100, including a processor 110 and a memory 109. The memory 109 stores a program or instructions that can run on the processor 110. When the program or instructions are executed by the processor 110, they implement the various steps of the above-described audio transmission method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0079] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0080] This application also provides a storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio transmission method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0081] The processor is the processor in the electronic device described in the above embodiments. The storage medium includes computer storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0082] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio transmission method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0083] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0084] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0085] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. An audio transmission method, characterized in that, An audio transmission system, comprising a transmitter and a receiver, is applied to an audio transmission system connected to a network for transmitting data via the network. The audio transmission method includes: The transmitting end is controlled to acquire audio data and encode the audio data to form an encoded frame; The parameters of the network are determined, and the transmitting end is controlled to switch to the corresponding working mode according to the parameters of the network. The working mode includes a first mode and a second mode. The first mode is different from the second mode. The parameters of the network include the network speed change rate and bandwidth. The coded frames are packaged into data packets according to the current working mode of the sending end, and the sending end is controlled to send the data packets to the receiving end. The receiving end is controlled to switch to the target mode, so that the receiving end receives the data packet and decodes the data packet according to the target mode in which the receiving end is located, to obtain decoded data, and then sends the decoded data outward. The target mode is the same as the working mode currently in which the sending end is located.
2. The audio transmission method according to claim 1, characterized in that, The step of determining the parameters of the network and controlling the transmitting end to switch to the corresponding working mode based on the parameters of the network includes: Determine the network speed change rate and bandwidth; If the network speed change rate is less than or equal to the network speed change threshold, and the bandwidth is less than or equal to the set bandwidth threshold, the sending end is controlled to switch to the first mode. If the network speed change rate is greater than the network speed change threshold and the bandwidth is less than or equal to the set bandwidth threshold, the sending end is controlled to switch to the second mode.
3. The audio transmission method according to claim 1, characterized in that, The step of packaging the encoded frame into a data packet according to the current working mode of the sending end, and controlling the sending end to send the data packet to the receiving end, includes: When the sending end is currently in the first mode, N consecutive encoded frames are packaged into the data packet, and the sending end is controlled to send the data packet to the receiving end; When the sending end is currently in the second mode, the current encoded frame and the encoded frames within a set time period prior to the current time are combined and packaged into the data packet, and the sending end is controlled to send the data packet to the receiving end.
4. The audio transmission method according to claim 1, characterized in that, The control of the receiving end to switch to the target mode, enabling the receiving end to receive the data packet, and decoding the data packet according to the target mode of the receiving end to obtain decoded data, includes: When the sending end is in the first mode, the receiving end is controlled to switch to the first mode, so that the receiving end receives the data packet and directly decodes the data packet based on the first mode to obtain the decoded data; When the sending end is in the second mode, the receiving end is controlled to switch to the second mode, so that the receiving end receives the data packet and decodes the data packet based on the second mode to obtain the decoded data.
5. The audio transmission method according to claim 4, characterized in that, Based on the second mode, the data packet is decoded to obtain the decoded data, including: Based on the second mode, the data packets are split according to a set duration to obtain multiple first split frame data; The sequence numbers of the multiple first split frame data are reassigned, and the multiple first split frame data are used to generate multiple new data packets, each of which corresponds to a sequence number; The multiple new data packets are sorted according to their sequence numbers to obtain a sorted sequence of new data packets; Remove duplicate new data packets from the sorted sequence of new data packets; The sequence of new data packets after removing duplicates is decoded to obtain the decoded data.
6. The audio transmission method according to claim 4, characterized in that, Based on the first mode, the data packet is directly decoded to obtain the decoded data, including: Based on the first mode, the data packet is split into multiple second split frame data, and the multiple second split frame data are decoded to obtain the decoded data.
7. The audio transmission method according to any one of claims 1-6, characterized in that, Controlling the transmitting end to collect audio data includes: The transmitting end is controlled to collect audio data at a set acquisition period.
8. The audio transmission method according to any one of claims 1-6, characterized in that, Encoding the audio data includes: The audio data is encoded according to a target coding rate, wherein the target coding rate is less than or equal to a set coding rate threshold, and the coding rate threshold is less than or equal to 3kbps.
9. An audio transmission device, characterized in that, An audio transmission device is used in an audio transmission system, which includes a transmitter and a receiver, and is connected to a network to transmit data. The audio transmission device includes: The first control module is used to control the transmitting end to collect audio data and encode the audio data to form an encoded frame; The determining module is used to determine the parameters of the network and control the transmitting end to switch to the corresponding working mode according to the parameters of the network. The working mode includes a first mode and a second mode, the first mode being different from the second mode. The parameters of the network include the network speed change rate and bandwidth. The packetization module is used to package the encoded frame into a data packet according to the current working mode of the sending end, and control the sending end to send the data packet to the receiving end; The second control module is used to control the receiving end to switch to the target mode, so that the receiving end receives the data packet, decodes the data packet according to the target mode of the receiving end, obtains decoded data, and sends the decoded data outward. The target mode is the same as the working mode of the sending end.
10. A storage medium, characterized in that, The storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio transmission method as described in any one of claims 1-8.