Packet Loss Retransmission Method, System, Device, Computer Readable Storage Medium and Equipment
By acquiring the loudness of the audio data packet in audio data transmission and resending the data packet according to the loudness, the problem of long data resend time in the prior art is solved, and data transmission efficiency and network resource utilization are improved.
Patent Information
- Application Number
- CN202010601648.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-06-28
AI Technical Summary
During the real-time audio data transmission process, the existing packet loss and resend method can easily lead to a long data resend time, thereby reducing data transmission efficiency.
By obtaining the loudness of the target audio packet and resend the target audio packet according to the loudness when receiving the packet loss state, the data resend time is optimized.
Improve data resend time, improve data transmission efficiency, and reduce network resource usage.
Smart Images

Figure CN113936670B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method for retransmitting lost packets, a system for retransmitting lost packets, a device for retransmitting lost packets, a computer-readable storage medium, and an electronic device. Background Art
[0002] During the transmission of audio data, packet loss often occurs during the data transmission process due to reasons such as unstable transmission networks. In response to packet loss, data retransmission is generally used to ensure that the receiving party receives complete data. The existing packet loss retransmission methods are usually as follows: when the packet loss situation feedback by the receiving party is detected, the lost data packets included in the packet loss situation are retransmitted. However, during the transmission of real-time audio data, the following situation usually exists: the audio data in the audio data packet (for example, slight ambient sound) may not necessarily be perceptible to the human ear after being decoded and output. If the packet loss situation of such data is also fed back according to the above packet loss retransmission method, it is likely to cause the problem of a long data retransmission time, and thus it is likely to cause low data transmission efficiency.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present application is to provide a method for retransmitting lost packets, a system for retransmitting lost packets, a device for retransmitting lost packets, a computer-readable storage medium, and an electronic device, which can improve the problem of a long data retransmission time and enhance the data transmission efficiency.
[0005] Other features and advantages of the present application will become apparent through the following detailed description, or will be partially learned through the practice of the present application.
[0006] According to one aspect of the present application, a method for retransmitting lost packets is provided, including:
[0007] Obtaining the loudness corresponding to a target audio data packet;
[0008] When receiving a packet loss status indicating that the target audio data packet is lost, retransmitting the target audio data packet according to the loudness corresponding to the target audio data packet.
[0009] In an exemplary embodiment of the present application, before obtaining the loudness corresponding to the target audio data packet, the above method further includes:
[0010] Screening out target audio data packets whose audio features meet preset conditions from the received multiple audio data packets.
[0011] In an exemplary embodiment of the present application, before screening out target audio data packets whose audio features meet preset conditions from the received multiple audio data packets, the above method further includes:
[0012] Performing packet loss detection on the received multiple audio data packets;
[0013] If the packet loss detection result includes a loss status, feedback the loss status to the sender terminal so that the sender terminal retransmits data for the loss status.
[0014] In an exemplary embodiment of the present application, after the sender terminal retransmits data for the loss status and before screening out target audio data packets whose audio features meet preset conditions from the received multiple audio data packets, the above method further includes:
[0015] Updating the multiple audio data packets according to the retransmitted data packets.
[0016] In an exemplary embodiment of the present application, the audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features corresponding to the audio bitstream. The audio features corresponding to the audio bitstream include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream.
[0017] In an exemplary embodiment of the present application, the received multiple audio data packets are sent by the sender terminal;
[0018] Among them, the specific manner in which the sender terminal sends multiple audio data packets is:
[0019] The sender terminal collects an audio signal and extracts features from the audio signal to obtain audio features;
[0020] The sender terminal encodes the audio signal to obtain an audio bitstream;
[0021] The sender terminal packs the audio bitstream and the audio features into an audio data packet and sends it to the server.
[0022] In an exemplary embodiment of the present application, the preset conditions include a preset energy amplitude and / or a preset signal-to-noise ratio. Screening out target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets includes:
[0023] If at least one energy amplitude greater than the preset energy amplitude is detected in the audio features, determining the audio data packet to which the audio bitstream corresponding to the audio features belongs as the target audio data packet; and / or,
[0024] If at least one signal-to-noise ratio greater than the preset signal-to-noise ratio is detected in the audio features, determining the audio data packet to which the audio bitstream corresponding to the audio features belongs as the target audio data packet.
[0025] In an exemplary embodiment of the present application, before obtaining the loudness corresponding to the target audio data packet, the above method further includes:
[0026] The sender terminal frames the audio bitstream according to a preset duration to obtain a plurality of audio frames;
[0027] The sender terminal processes the plurality of audio frames through a preset window function respectively to obtain a plurality of reference frames;
[0028] The sender terminal calculates the power spectra respectively corresponding to the plurality of reference frames;
[0029] The sender terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum.
[0030] In an exemplary embodiment of the present application, the preset window function is a Hanning window function, a Hamming window function, a Blackman window function, a Kaiser window function, a triangular window function or a rectangular window function.
[0031] In an exemplary embodiment of the present application, the sender terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum, including:
[0032] The sender terminal calculates the frequency point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum;
[0033] The sender terminal calculates the loudness weight of each frequency point in the power spectrum according to the frequency point loudness;
[0034] The sender terminal calculates the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum;
[0035] The sender terminal determines the sum of the loudness values corresponding to the plurality of reference frames as the loudness corresponding to the target audio data packet.
[0036] In an exemplary embodiment of the present application, retransmitting the target audio data packet according to the loudness corresponding to the target audio data packet includes:
[0037] If the loudness corresponding to the target audio data packet is greater than the preset loudness, retransmit the target audio data packet to the receiver terminal so that the receiver terminal decodes and outputs the target audio data packet after and before retransmission.
[0038] In an exemplary embodiment of the present application, the manner in which the receiver terminal decodes and outputs the target audio data packet after and before retransmission is specifically:
[0039] The receiver terminal decodes the target audio data packet after and before retransmission to obtain a plurality of audio signals to be output;
[0040] The receiving - end terminal mixes multiple audio signals to be output, obtains a mixed signal and plays it.
[0041] According to one aspect of the present application, a packet - loss retransmission system is provided, including a sending - end terminal, a server, and a receiving - end terminal, where:
[0042] The sending - end terminal is used to send a target audio data packet to the server, and the target audio data packet includes the loudness corresponding to the target audio data packet;
[0043] The server is used to obtain the loudness corresponding to the target audio data packet;
[0044] The receiving - end terminal is used to send a packet - loss status indicating the loss of the target audio data packet to the server;
[0045] The server is further used to, when receiving the packet - loss status, re - send the target audio data packet to the receiving - end terminal according to the loudness corresponding to the target audio data packet;
[0046] The receiving - end terminal is further used to receive the target audio data packet.
[0047] According to one aspect of the present application, a packet - loss retransmission device is provided, including:
[0048] A loudness acquisition unit, used to acquire the loudness corresponding to the target audio data packet;
[0049] A data sending unit, used to, when receiving a packet - loss status indicating the loss of the target audio data packet, re - send the target audio data packet according to the loudness corresponding to the target audio data packet.
[0050] In an exemplary embodiment of the present application, the above - mentioned device further includes a data - packet screening unit, where:
[0051] The data - packet screening unit is used to screen out target audio data packets whose audio features meet preset conditions from multiple received audio data packets before the loudness acquisition unit acquires the loudness corresponding to the target audio data packet.
[0052] In an exemplary embodiment of the present application, the above - mentioned device further includes a packet - loss detection unit and a packet - loss status feedback unit, where:
[0053] The packet - loss detection unit is used to perform packet - loss detection on multiple received audio data packets before the data - packet screening unit screens out target audio data packets whose audio features meet preset conditions from multiple received audio data packets;
[0054] The packet - loss status feedback unit is used to, if the packet - loss detection result includes a loss status, feedback the loss status to the sending - end terminal so that the sending - end terminal performs data re - transmission for the loss status.
[0055] In an exemplary embodiment of the present application, the above device further includes a data packet updating unit, where:
[0056] The data packet updating unit is configured to update a plurality of audio data packets according to the retransmitted data packets after the sender terminal retransmits data for the lost state and before the data packet screening unit screens target audio data packets with audio features meeting preset conditions from the received plurality of audio data packets.
[0057] In an exemplary embodiment of the present application, the audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features corresponding to the audio bitstream. The audio features corresponding to the audio bitstream include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream.
[0058] In an exemplary embodiment of the present application, the received plurality of audio data packets are sent by the sender terminal;
[0059] Among them, the specific manner for the sender terminal to send a plurality of audio data packets is as follows:
[0060] The sender terminal collects an audio signal and extracts features from the audio signal to obtain audio features;
[0061] The sender terminal encodes the audio signal to obtain an audio bitstream;
[0062] The sender terminal packs the audio bitstream and the audio features into an audio data packet and sends it to the server.
[0063] In an exemplary embodiment of the present application, the preset conditions include a preset energy amplitude and / or a preset signal-to-noise ratio. The data packet screening unit screens target audio data packets with audio features meeting the preset conditions from the received plurality of audio data packets, including:
[0064] If at least one energy amplitude greater than the preset energy amplitude is detected in the audio features, the audio data packet to which the audio bitstream corresponding to the audio features belongs is determined as the target audio data packet; and / or,
[0065] If at least one signal-to-noise ratio greater than the preset signal-to-noise ratio is detected in the audio features, the audio data packet to which the audio bitstream corresponding to the audio features belongs is determined as the target audio data packet.
[0066] In an exemplary embodiment of the present application, it further includes:
[0067] Before the loudness acquisition unit of the sender terminal acquires the loudness corresponding to the target audio data packet, the sender terminal performs frame division processing on the audio bitstream according to a preset duration to obtain a plurality of audio frames;
[0068] The sender terminal processes multiple audio frames respectively through a preset window function to obtain multiple reference frames;
[0069] The sender terminal calculates the power spectra respectively corresponding to the multiple reference frames;
[0070] The sender terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum.
[0071] In an exemplary embodiment of the present application, the preset window function is a Hanning window function, a Hamming window function, a Blackman window function, a Kaiser window function, a triangular window function or a rectangular window function.
[0072] In an exemplary embodiment of the present application, the sender terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum, including:
[0073] The sender terminal calculates the frequency point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum;
[0074] The sender terminal calculates the loudness weight of each frequency point in the power spectrum according to the frequency point loudness;
[0075] The sender terminal calculates the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum;
[0076] The sender terminal determines the sum of the loudness values corresponding to the multiple reference frames as the loudness corresponding to the target audio data packet.
[0077] In an exemplary embodiment of the present application, the data sending unit retransmits the target audio data packet according to the loudness corresponding to the target audio data packet, including:
[0078] If the loudness corresponding to the target audio data packet is greater than the preset loudness, the sender terminal retransmits the target audio data packet to the receiver terminal so that the receiver terminal decodes and outputs the target audio data packet after and before retransmission.
[0079] In an exemplary embodiment of the present application, the manner in which the receiver terminal decodes and outputs the target audio data packet after and before retransmission is specifically:
[0080] The receiver terminal decodes the target audio data packet after and before retransmission to obtain multiple audio signals to be output;
[0081] The receiver terminal performs mixing processing on the multiple audio signals to be output to obtain a mixed signal and plays it.
[0082] According to one aspect of the present application, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method described in any one of the above by executing the executable instructions.
[0083] According to one aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0084] The exemplary embodiments of the present application may have some or all of the following beneficial effects:
[0085] In the packet loss retransmission method provided by an exemplary embodiment of the present application, the loudness corresponding to the target audio data packet can be obtained, and when a packet loss state indicating the loss of the target audio data packet is received, the target audio data packet is retransmitted according to the loudness corresponding to the target audio data packet. According to the above description of the solution, on the one hand, the present application can use the loudness corresponding to the audio data packet as a data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency; on the other hand, the target audio data packet can be retransmitted in a targeted manner. Compared with the prior art that retransmits all the data packets of the current transmission including the target audio data packet, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0086] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0088] Figure 1 A schematic diagram showing an exemplary system architecture of a packet loss retransmission method and a packet loss retransmission device to which the embodiments of the present application can be applied;
[0089] Figure 2 A schematic diagram showing the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application;
[0090] Figure 3 A schematic diagram showing the architecture of a packet loss retransmission method based on the embodiments of the present application;
[0091] Figure 4Schematically shows a flowchart of a packet loss retransmission method according to an embodiment of the present application;
[0092] Figure 5 Schematically shows a loudness weight curve graph according to an embodiment of the present application;
[0093] Figure 6 Schematically shows an acoustic equal-loudness curve graph according to an embodiment of the present application;
[0094] Figure 7 Schematically shows a sequence diagram of a packet loss retransmission method according to an embodiment of the present application;
[0095] Figure 8 Schematically shows a structural block diagram of a packet loss retransmission system according to an embodiment of the present application;
[0096] Figure 9 Schematically shows a structural block diagram of a packet loss retransmission system according to another embodiment of the present application;
[0097] Figure 10 Schematically shows a structural block diagram of a packet loss retransmission system according to still another embodiment of the present application;
[0098] Figure 11 Schematically shows a structural block diagram of a packet loss retransmission device according to an embodiment of the present application. Detailed implementation manners
[0099] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of this application.
[0100] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0101] Figure 1 The figure shows a schematic diagram of a system architecture of an exemplary application environment to which the packet loss retransmission method and the packet loss retransmission device according to the embodiments of the present application can be applied.
[0102] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with display screens, including but not limited to desktop computers, portable computers, smart phones, and tablet computers, etc. It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0103] are merely illustrative. Specifically, the packet loss retransmission method provided by the embodiments of the present application is generally executed by the server 105. Correspondingly, the packet loss retransmission device is generally disposed in the server 105. However, those skilled in the art can easily understand that the packet loss retransmission method provided by the embodiments of the present application can also be executed by the terminal devices 101, 102, or 103. Correspondingly, the packet loss retransmission device can also be disposed in the terminal devices 101, 102, or 103. No special limitation is made in this exemplary embodiment. For example, in an exemplary embodiment, the server 105 may obtain the loudness corresponding to the target audio data packet; if a packet loss status indicating the loss of the target audio data packet is received, the target audio data packet is retransmitted according to the loudness corresponding to the target audio data packet. According to the implementation requirements, there may be any number of terminal devices, networks, and servers.
[0104] Figure 2 The figure shows a schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.
[0105] It should be noted that Figure 2The computer system 200 of the illustrated electronic device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.
[0106] As Figure 2 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 202 or the programs loaded from the storage section 208 into the random access memory (RAM) 203. In the RAM 203, various programs and data required for system operation are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0107] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as required. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as required so that a computer program read from it can be installed into the storage section 208 as required.
[0108] Specifically, according to the embodiments of the present application, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 209, and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the methods and apparatuses of the present application are executed.
[0109] Generally speaking, it is difficult to avoid packet loss during audio transmission. The common reasons for packet loss usually include: wireless channel interference of WIFI or mobile network, router congestion during peak periods, insufficient performance of mobile devices, etc. When an audio data packet takes too long to be transmitted over the network, that is, it cannot be conveyed in time when it needs to be played, it will be determined as a lost packet even if it is received later. The existing packet loss handling methods generally are: when the receiving end receives the data transmitted by the sending end, it performs packet loss detection on it. If there is a packet loss situation, it will feedback to the sending end to make it retransmit the data until it is detected that there is no packet loss in the received data.
[0110] Please refer to Figure 3 , Figure 3 which schematically shows the architecture diagram of a packet loss retransmission method based on the embodiments of the present application. Figure 3 The shown architecture diagram includes: a sending terminal 310, a server 320, and a receiving terminal 330; wherein, the sending terminal 310 includes a feature extraction module 311, a voice encoding module 312, and a data retransmission module 313, the server 320 includes a packet loss detection module 321, an audio routing module 322, and a data retransmission module 323, and the receiving terminal 330 includes a packet loss detection module 331, a voice decoding module 332, a mixing module 333, and a playback module 334.
[0111] Specifically, the feature extraction module 311 can collect audio signals, extract features from the collected audio signals to obtain audio features, and send the collected audio signals to the voice encoding module 312. The voice encoding module 312 can encode the audio signals to obtain corresponding audio bitstreams, and send the audio bitstreams to the data retransmission module 323. The data retransmission module 323 can package the matching audio bitstreams and audio features into audio data packets, and after packing multiple audio data packets, forward them as a data group to the server 320, so that the packet loss detection module 321 in the server 320 performs packet loss detection on the received data group.
[0112] If the packet loss detection result includes audio data packets in a lost state, the packet loss detection module 321 feeds back the lost state to the data retransmission module 313, so that the data retransmission module 313 retransmits data for the lost state. Furthermore, when the packet loss detection module 321 receives the retransmitted audio data packets and detects that there are no audio data packets in a lost state in the data packet group, it updates the data packet group according to the retransmitted audio data packets and sends the updated data packet group to the audio routing module 322. Furthermore, the audio routing module 322 can screen the audio data packets in the data packet group according to the audio features corresponding to each data packet in the data packet group to obtain target audio data packets; wherein, the maximum energy amplitude corresponding to the target audio data packets is greater than that of other audio data packets in the data packet group, or the similarity between the energy spectrum distribution corresponding to the target audio data packets and the preset energy spectrum distribution is greater than that of other audio data packets in the data packet group. Furthermore, the audio routing module 322 can send the target audio data packets to the data retransmission module 323, so that the data retransmission module 323 forwards the target audio data packets to the receiving terminal 330. Furthermore, the packet loss detection module 331 in the receiving terminal 330 performs packet loss detection on the target audio data packets.
[0113] If the packet loss detection result includes target audio data packets in a lost state, the packet loss detection module 331 feeds back the lost state to the data retransmission module 323, so that the data retransmission module 323 retransmits data for the lost state. Furthermore, when the packet loss detection module 331 receives the retransmitted target audio data packets and detects that there are no target audio data packets in a lost state among all the target audio data packets, it updates all the target audio data packets according to the retransmitted target audio data packets and sends the updated all target audio data packets to the voice decoding module 332. Furthermore, the voice decoding module 332 can decode the audio code streams in each target audio data packet into audio signals and send the decoded audio signals to the mixing module 333, so that the mixing module 333 performs mixing processing on the audio signals. Furthermore, the mixing module 333 can send the mixing processing result to the playback module 334, so that the playback module 334 plays the mixing processing result.
[0114] In the above packet loss retransmission method, each data packet transmission process may include multiple data packets to be transmitted. Only when all the data packets are received completely (i.e., there is no packet loss) can the next module be triggered to perform corresponding operations. When applied to the field of multi-person meetings, the amount of audio signal acquisition is large. If the above method is used for packet loss retransmission, it is easy to cause the problem of low data transmission efficiency, which will lead to a long time difference between the moment when the receiving terminal plays the audio and the moment when the sending terminal sends the audio, affecting the real-time performance of audio output in multi-person meetings and the user experience.
[0115] In addition, for a multi-person conference, when playing multi-channel audio signals, users generally have a weaker perception of audio signals with lower loudness. Therefore, the applicant thought that audio signals with lower loudness that are not easily perceptible to the user's hearing can be used as the screening condition for audio data packets. Specifically, if the server can not only screen audio data packets according to the energy of the audio signal, but also perform secondary screening on the screened audio data packets through the loudness of the audio signal, and then transmit the audio data packets after the secondary screening to the receiving terminal, the data transmission efficiency can be improved in case of packet loss, and the time difference between the moment when the receiving terminal plays the audio and the moment when the sending terminal sends the audio can be shortened, thereby improving the real-time performance of audio output and the user experience in a multi-person conference.
[0116] Based on the above problems, the present exemplary embodiment provides a packet loss retransmission method. This packet loss retransmission method can be applied to the above-mentioned server 105, or can be applied to one or more of the above-mentioned terminal devices 101, 102, 103, and no special limitation is made in this exemplary embodiment. Refer to Figure 4 As shown, this packet loss retransmission method may include the following steps S410 to step S420:
[0117] Step S410: Obtain the loudness corresponding to the target audio data packet.
[0118] Step S420: When receiving a packet loss status indicating that the target audio data packet is lost, retransmit the target audio data packet according to the loudness corresponding to the target audio data packet.
[0119] Among them, step S410 and step S420 can be executed by the server. For example, the server can be a routing server for screening path signals.
[0120] Implement Figure 1 The method shown can use the loudness corresponding to the audio data packet as the data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency. In addition, the target audio data packet can be retransmitted in a targeted manner. Compared with the prior art of retransmitting all data packets sent in the current time including the target audio data packet, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0121] Next, the above steps of the present exemplary embodiment will be described in more detail.
[0122] In step S410, obtain the loudness corresponding to the target audio data packet.
[0123] In an embodiment of the present application, optionally, before obtaining the loudness corresponding to the target audio data packet, the above method further includes: screening out the target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets.
[0124] Among them, the audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features corresponding to the audio bitstream. The audio features corresponding to the audio bitstream may include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream. The target audio data packet may be one or more.
[0125] It can be seen that implementing this optional embodiment can screen the received audio data packets through audio features to reduce the amount of data to be forwarded, thereby reducing the consumption of network resources and improving data transmission efficiency.
[0126] In an embodiment of the present application, optionally, before screening out the target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets, the above method further includes: performing packet loss detection on the received multiple audio data packets; if the packet loss detection result includes a loss status, feedback the loss status to the sending terminal so that the sending terminal retransmits the data for the loss status.
[0127] Specifically, the number of sending terminals may be one or more, which is not limited in the embodiment of the present application. When the embodiment of the present application is applied to a multi-person conference, the conference member terminals in the multi-person conference can be used as the sending terminals and also as the receiving terminals. In addition, multiple audio data packets may belong to the same data packet. The sending terminal can send one data packet at a time, and each data packet includes one or more audio data packets. The multiple audio data packets correspond to different audio channels.
[0128] As an alternative implementation, the method for detecting packet loss of multiple received audio data packets can be specifically as follows: Determine the sequence numbers corresponding to the multiple received audio data packets within a unit time (e.g., 5 s); then, detect the continuity of the sequence numbers corresponding to the multiple audio data packets according to the packet header (e.g., Sequence Number consecutive numbers) of the data transmission protocol (e.g., TCP protocol); if the sequence numbers corresponding to the multiple audio data packets are not continuous (e.g., 1, 2, 4, 5), it is determined that a packet loss situation occurs; if the sequence numbers corresponding to the multiple audio data packets are continuous (e.g., 1, 2, 3, 4), it is determined that no packet loss situation occurs. Among them, when no packet loss situation occurs, the packet loss statuses corresponding to the multiple audio data packets are non-loss statuses; then, determine the lost data packets corresponding to the packet loss situation according to the sequence numbers corresponding to the multiple audio data packets; then, generate a packet loss detection result according to the lost data packets; where the number of lost data packets can be one or more, which is not limited in the embodiments of the present application. In addition, the packet loss detection result can include at least one of the lost data packets, the sequence numbers corresponding to the lost data packets, and the packet loss statuses corresponding to the lost data packets; among them, the packet loss status can be represented by the value 0 or 1, the value 0 can represent the lost status in the packet loss status, and the value 1 can represent the non-loss status in the packet loss status.
[0129] As an alternative implementation, the method for feeding back the lost status to the sending terminal can be specifically as follows: Feed back a Negative Acknowledgement (NACK) signal including the packet loss detection result to the sending terminal. It can be understood that in this alternative implementation, if the packet loss detection result does not include the lost status, the following steps can be included: Send a positive Acknowledgement (ACK) signal indicating normal data reception to the sending terminal.
[0130] As an alternative implementation, the method for the sending terminal to retransmit data for the lost status can be specifically as follows: The sending client determines the lost data packets corresponding to the lost status and retransmits the lost data packets; or, the sending client retransmits the data packets corresponding to the lost status, where the data packets are composed of the above-mentioned multiple audio data packets.
[0131] It can be seen that implementing this alternative embodiment can trigger the sending party to retransmit data when a packet loss status is detected, which can improve the integrity of the transmitted data.
[0132] In an embodiment of the present application, optionally, after the sending terminal retransmits data for the lost state and before screening target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets, the above method further includes: updating the multiple audio data packets according to the retransmitted data packets.
[0133] Specifically, the number of retransmitted data packets can be one or more, and the updated multiple audio data packets include the above-mentioned retransmitted data packets.
[0134] As an optional implementation manner, the manner of updating the multiple audio data packets according to the retransmitted data packets can specifically be: determining the sequence numbers corresponding to the retransmitted data packets and the sequence numbers corresponding to the multiple audio data packets respectively, and integrating the retransmitted data packets and the above-mentioned multiple audio data packets according to the sequence number order to implement the update of the multiple audio data packets.
[0135] It can be seen that implementing this optional embodiment can ensure the integrity of the data sent to the receiving terminal through data packet update.
[0136] In an embodiment of the present application, optionally, the received multiple audio data packets are sent by the sending terminal;
[0137] Among them, the manner in which the sending terminal sends multiple audio data packets is specifically: the sending terminal collects an audio signal and extracts features from the audio signal to obtain audio features; the sending terminal encodes the audio signal to obtain an audio bitstream; the sending terminal packs the audio bitstream and the audio features into audio data packets and sends them to the server.
[0138] It should be noted that the audio signal collected by the sending terminal can be an analog signal. The audio bitstream is used to represent the data traffic used by the audio signal per unit time, and the audio bitstream = sampling rate * number of bits * number of channels. For example, 44100 * 16 * 2 = 1.41 Mbit / sec.
[0139] As an optional implementation manner, the audio features corresponding to the audio signal can include at least one of the zero-crossing rate, short-time energy, short-time autocorrelation function, and short-time average magnitude difference, which is not limited in the embodiments of the present application.
[0140] When the audio features corresponding to the audio signal include the zero-crossing rate, the manner in which the sending terminal collects the audio signal and extracts features from the audio signal can specifically be: the sending terminal collects the audio signal and according to The zero-crossing rate of the audio signal is calculated as the audio feature of the audio signal, where N is the frame length of the audio signal and n is the number of frames of the audio signal. The calculated zero-crossing rate can be used to characterize the number of times each frame of the audio signal passes through the zero value. The above zero-crossing rate can be used to determine the unvoiced and voiced sounds in the audio signal, which is beneficial for the server to screen multiple audio data packets received.
[0141] When the audio feature corresponding to the audio signal includes short-time energy, the sending terminal collects the audio signal and extracts the feature of the audio signal. The specific method for obtaining the audio feature can be: the sending terminal collects the audio signal and extracts the feature according to the short-time energy. Detect the nth frame audio signal x in the collected audio signal n (m) short-time energy E n , the short-time energy corresponding to each frame is obtained as the audio feature of the audio signal; wherein N is the frame length of the audio signal, and n is a positive integer. Generally speaking, the energy of human voice is greater than the energy of noise. The calculated short-time energy corresponding to each frame can be used to distinguish between human voice and noise in the audio signal, which is beneficial for the server to filter out noise from multiple received audio data packets according to the energy corresponding to each frame, thereby improving the processing efficiency of audio data.
[0142] When the audio feature corresponding to the audio signal includes a short-time autocorrelation function, the sending terminal collects the audio signal and extracts the feature of the audio signal. The specific method for obtaining the audio feature can be: the sending terminal collects the audio signal and extracts the feature according to The short-time autocorrelation function of the audio signal is calculated as the audio feature of the audio signal; wherein N is the frame length of the audio signal, n is a positive integer, w is used to represent the window function, and w'(m) is used to represent the audio frame after windowing. The calculated short-time autocorrelation function can be used to measure the similarity of the signal's own time waveform, which is beneficial for the server to detect similar characteristics of the audio according to the short-time autocorrelation function, thereby improving the processing efficiency of the audio data.
[0143] When the audio feature corresponding to the audio signal includes the short-time average amplitude difference, the sending terminal collects the audio signal and extracts the feature of the audio signal. The specific method for obtaining the audio feature can be: the sending terminal collects the audio signal and extracts the feature according to the short-time average amplitude difference. The short-time average amplitude difference of the audio signal is calculated as the audio feature of the audio signal; wherein k=0, 1, ..., N-1, and the calculated short-time average amplitude difference can be used to measure the change of the audio amplitude.
[0144] In addition, the audio features corresponding to the audio signal may also include spectrogram, short-time power spectral density, spectral entropy, fundamental frequency, resonance peak and other features, which are not limited in the embodiments of the present application.
[0145] It can be seen that by implementing this optional embodiment, the sender terminal can extract features from the audio signal, which is conducive to the server selecting the path signal based on the result of feature extraction, thereby ensuring the real-time performance of the audio output in a multi-person conference while also ensuring the audio output effect.
[0146] In an embodiment of the present application, optionally, the preset condition includes a preset energy amplitude and / or a preset signal-to-noise ratio. Screening the target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets includes: if at least one energy amplitude greater than the preset energy amplitude is detected in the audio features, determining the audio data packet to which the audio code stream corresponding to the audio features belongs as the target audio data packet; and / or, if at least one signal-to-noise ratio greater than the preset signal-to-noise ratio is detected in the audio features, determining the audio data packet to which the audio code stream corresponding to the audio features belongs as the target audio data packet.
[0147] Specifically, each audio frame in the audio code stream corresponds to an energy amplitude, and the energy amplitude is used to represent the short-time energy corresponding to the audio frame. The signal-to-noise ratio (SNR) is the ratio of the average power of the audio signal to the average power of the noise, that is: SNR (dB) = 10 * log10(S / N) (dB).
[0148] It can be seen that by implementing this optional embodiment, the audio data packets can be screened by the energy amplitude or the signal-to-noise ratio to streamline the audio data packets to be transmitted and improve the data transmission efficiency.
[0149] In an embodiment of the present application, optionally, before obtaining the loudness corresponding to the target audio data packet, the above method further includes: the sender terminal performs frame division processing on the audio code stream according to a preset duration to obtain multiple audio frames; processes the multiple audio frames through a preset window function respectively to obtain multiple reference frames; calculates the power spectra corresponding to the multiple reference frames respectively; and calculates the loudness corresponding to the target audio data packet according to the power spectra.
[0150] Specifically, the durations of the multiple audio frames are the same (for example, 10 ms or 20 ms), and the durations of the multiple reference frames are the same as those of the corresponding audio frames. In addition, since the Fast Fourier Transform (FFT) can only transform time-domain data of a finite length, signal truncation is required for the time-domain signal. Even for a periodic signal, if the time length of the truncated periodic signal is not an integer multiple of the period, leakage is likely to occur in the intercepted signal. Therefore, this application applies a preset window function that can make the time-domain audio signal meet the periodic requirements of the Fourier transform to reduce signal leakage. The preset window function is a Hanning window function, a Hamming window function, a Blackman window function, a Kaiser window function, a triangular window function, or a rectangular window function.
[0151] As an alternative implementation, the method for obtaining multiple reference frames by separately processing multiple audio frames through a preset window function may be: determining the time-domain expressions respectively corresponding to the multiple audio frames; multiplying the time-domain expressions respectively corresponding to the multiple audio frames by the preset window function to obtain the reference frames respectively corresponding to the multiple audio frames.
[0152] As an alternative implementation, the method for calculating the power spectra respectively corresponding to the multiple reference frames may be: performing a fast Fourier transform on the multiple reference frames to determine the power spectra respectively corresponding to the multiple reference frames.
[0153] It can be seen that implementing this alternative embodiment can calculate the audio loudness of the target audio data packet, and this audio loudness can be used as a data retransmission condition. When the audio loudness is low, the server may not retransmit the lost target audio data packet to reduce the occupancy of network resources.
[0154] In an embodiment of the present application, optionally, the method for the sending terminal to calculate the loudness corresponding to the target audio data packet according to the power spectrum includes: the sending terminal calculates the frequency-point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum; calculates the loudness weight of each frequency point in the power spectrum according to the frequency-point loudness; calculates the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum; determines the sum of the loudness values corresponding to the multiple reference frames as the loudness corresponding to the target audio data packet.
[0155] Specifically, a frequency point is the number of a fixed frequency and can be used as the unique representation of the fixed frequency.
[0156] As an alternative implementation, the method for calculating the frequency-point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum may specifically be: determining the absolute value P(i, j) of the energy amplitude of each frequency point i in the power spectrum, where j = 0 to K - 1 and K is the total number of frequency points.
[0157] As an alternative implementation, the method for calculating the loudness weight of each frequency point in the power spectrum according to the frequency-point loudness may specifically be: calculating the loudness weight cof(freq) of each frequency point freq in the power spectrum based on the following formula; where
[0158]
[0159]
[0160]
[0161] The loudness loud of freq is 4.2 + afy * (dB - cfy) / (1 + bfy * (dB - cfy));
[0162]
[0163] Refer to the expression of cof(freq) Figure 5 , Figure 5 which schematically shows a loudness weight curve graph according to an embodiment of the present application. Figure 5 In the shown curve graph, the horizontal axis is used to represent frequency, and the vertical axis is used to represent loudness weight. Given the loudness at a known frequency point, the weight corresponding to this frequency point can be determined through Figure 5 the shown curve graph.
[0164] As an alternative embodiment, the way to calculate the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum can be specifically: According to the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum, as the loudness value EP(i) of the reference frame corresponding to the power spectrum; where i is the frame number and k is the frequency point number.
[0165] It can be seen that implementing this alternative embodiment can determine the loudness corresponding to the target audio data packet as the loudness for the data retransmission condition according to the loudness corresponding to each frequency point, thereby improving the data retransmission efficiency of the server.
[0166] In step S420, when receiving a packet loss status indicating the loss of the target audio data packet, retransmit the target audio data packet according to the loudness corresponding to the target audio data packet.
[0167] In the embodiments of the present application, optionally, retransmitting the target audio data packet according to the loudness corresponding to the target audio data packet includes: if the loudness corresponding to the target audio data packet is greater than the preset loudness, retransmit the target audio data packet to the receiving terminal, so that the receiving terminal decodes and outputs the target audio data packet after and before retransmission.
[0168] Specifically, loudness is a subjective perception quantity used to describe the size of sound, and the unit is sone. For a 1000Hz pure tone, the loudness is 1 sone when the sound pressure level is 40dB; the sound of 2 sones is twice the loudness of a 40-phon sound; 4 sones is 4 times the loudness of a 40-phon sound. That is to say, when the sound pressure level increases by 10dB, the loudness doubles, and the human ear's perception of loudness changes with the sound pressure level. Please refer to Figure 6 , Figure 6 which schematically shows an acoustic equal-loudness curve graph according to an embodiment of the present application. Figure 6The acoustic equal-loudness curves shown display curves of multiple loudness levels (i.e., 100 phon, 80 phon, 60 phon, 40 phon, 20 phon, threshold phon). The horizontal axis is used to represent the sound wave frequency (Frequency), and the vertical axis is used to represent the sound pressure level (Sound Pressure Level). By Figure 6 As can be seen from the acoustic equal-loudness curve diagram shown, the equal-loudness curve is a curve that describes the relationship between the sound pressure level and the sound wave frequency under equal-loudness conditions. In the mid-low frequency range (below 1 kHz), the lower the frequency, the greater the sound pressure intensity (energy) required for equal loudness. That is, the greater the sound energy, the more consistent the auditory perception of the human ear. In the mid-high frequency range (above 1 kHz), different frequency bands correspond to different acoustic auditory perception characteristics. Refer to Figure 6 It can be seen that if the audio data packets are screened by loudness, the sounds with weak human ear perception ability can be filtered out, which can improve the data transmission effect.
[0169] As an optional implementation manner, if the loudness of the target audio data packet is greater than the preset loudness, the method of retransmitting the target audio data packet to the receiving party terminal can be: detecting whether the loudness of the target audio data packet is greater than the preset loudness C; if so, triggering the enabling of the data retransmission function (i.e., triggering the enable switch ArqEnable = 1), and retransmitting the target audio data packet to the receiving party terminal through the data retransmission function; if not, closing the data retransmission function (i.e., triggering the enable switch ArqEnable = 0). That is to say, the target audio data packet can be retransmitted to the receiving party terminal according to the following expression
[0170] It can be seen that implementing this optional embodiment can screen the audio data packets according to loudness, reduce the occupancy of network resources by the sounds with weak human ear perception ability, and thus improve the data transmission efficiency. In addition, since this application can retransmit the lost data packets, the problem of low transmission efficiency caused by retransmitting all the data of the entire data packet every time in the prior art can be solved.
[0171] In the embodiment of this application, optionally, the method for the receiving party terminal to decode and output the target audio data packet after and before retransmission is specifically: the receiving party terminal decodes the target audio data packet after and before retransmission to obtain multiple audio signals to be output; the receiving party terminal performs mixing processing on the multiple audio signals to be output to obtain a mixed signal and play it.
[0172] As an alternative embodiment, before the receiving terminal mixes multiple audio signals to be output, the following steps may further be included: The receiving terminal performs format normalization on the multiple audio signals to be output, which can ensure that the formats of the multiple audio signals to be output are unified; further, the multiple audio signals to be output are converted to a preset sampling rate (such as 16 kHz, 32 kHz, 44.1 kHz, 48 kHz, etc.); further, the consistency of the bit depth (Bit-Depth) or sample format (Sample Format) of the multiple audio signals to be output is detected. If the bit depths are inconsistent or the sample formats are inconsistent, corresponding normalization processing is performed on the multiple audio signals to be output so that the number of bits carrying the audio data of each sampling point is the same; further, it is detected whether the channels of the multiple audio signals to be output (such as mono or stereo) are consistent. If they are inconsistent, a prompt message indicating inconsistent channels is output; if they are consistent, the operation of mixing the multiple audio signals to be output is performed. In addition, optionally, before the receiving terminal mixes the multiple audio signals to be output, the receiving terminal may also perform processing such as echo cancellation, noise suppression, and silence detection on the multiple audio signals to be output, which is not limited in the embodiments of the present application.
[0173] As an alternative embodiment, the way for the receiving terminal to mix multiple audio signals to be output to obtain a mixed signal may be: The receiving terminal adjusts the volume of the multiple audio signals to be output according to the energy amplitudes respectively corresponding to the multiple audio signals to be output, and mixes the results of the equalization adjustment to obtain a mixed signal.
[0174] As an alternative embodiment, after the receiving terminal mixes multiple audio signals to be output to obtain a mixed signal, the following steps may further be included: The receiving terminal performs overflow detection on the mixed signal. If there is an overflow, the overflow sampling points are processed / smoothed, and then the processing result is played.
[0175] It can be seen that by implementing this alternative embodiment, after receiving a complete plurality of audio data packets, the receiving terminal can mix and output them. The output result retains the important audio content in the meeting and discards the audio content with weak human ear perception ability, which can improve the message forwarding efficiency of a multi-person meeting, ensure the real-time nature of the output audio signal, and thus improve the user experience.
[0176] Please refer to Figure 7 , Figure 7 which schematically shows a sequence diagram of a packet loss retransmission method according to an embodiment of the present application. As Figure 7 shown, it includes step S700 to step S780, where:
[0177] In step S700, the sender terminal may perform feature extraction on the collected audio signal to obtain audio features; and, encode the audio signal to obtain an audio code stream corresponding to the audio signal; and, package the loudness, audio code stream, and audio features corresponding to the audio data packet into an audio data packet; and, send the multiple audio data packets obtained by packaging to the server.
[0178] In step S710, after receiving multiple audio data packets, the server may perform packet loss detection.
[0179] In step S720, if the packet loss detection result includes a loss status, the server feeds back the loss status to the sender terminal.
[0180] In step S730, the sender terminal retransmits data to the server for the loss status.
[0181] In step S740, the server may update multiple audio data packets according to the retransmitted data packets, screen out target audio data packets whose audio features meet the preset conditions from the multiple audio data packets, and forward the target audio data packets to the receiver terminal.
[0182] In step S750, after receiving the target audio data packet, the receiver terminal may perform packet loss detection.
[0183] In step S760, if the packet loss detection result includes a loss status, the receiver terminal feeds back the loss status of the target audio data packet to the server.
[0184] In step S770, if the loudness of the target audio data packet is greater than the preset loudness, the server retransmits the target audio data packet to the receiver terminal.
[0185] In step S780, the receiver terminal decodes and outputs the target audio data packet after and before retransmission.
[0186] Optionally, the audio data packet may not include loudness either. After the sender terminal sends the audio data packet to the server, the server may calculate the loudness corresponding to the audio signal according to the audio features in the audio data packet.
[0187] It can be seen that implementing Figure 7 the method shown can initiate the data retransmission function for audio data with strong human ear perception ability and does not initiate the data retransmission function for audio data with weak human ear perception ability to ensure the data transmission quality. Furthermore, it can ensure that the transmission quality of the main sound source in a multi-person network call is guaranteed to the greatest extent, avoid unnecessary consumption of network bandwidth resources, and at the same time reduce the poor user call experience of increased call latency caused by packet loss retransmission without screening conditions.
[0188] Further, in the present exemplary embodiment, a packet loss retransmission system 800 is also provided. Please refer to Figure 8 , which includes a sender terminal 801, a server 802, and a receiver terminal 803, where:
[0189] The sender terminal 801 is configured to send a target audio data packet to the server 802, and the target audio data packet includes the loudness corresponding to the target audio data packet;
[0190] The server 802 is configured to obtain the loudness corresponding to the target audio data packet;
[0191] The receiver terminal 803 is configured to send a packet loss status indicating the loss of the target audio data packet to the server 802;
[0192] The server 802 is further configured to, when receiving the packet loss status, retransmit the target audio data packet to the receiver terminal 803 according to the loudness corresponding to the target audio data packet;
[0193] The receiver terminal 803 is further configured to receive the target audio data packet.
[0194] It can be seen that implementing Figure 8 the system shown can use the loudness corresponding to the audio data packet as the data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency. In addition, the target audio data packet can be retransmitted in a targeted manner. Compared with the prior art that retransmits all the packets of the current transmission including the target audio data packet, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0195] Specifically, please refer to Figure 9 . Figure 9 Schematically shows a structural block diagram of a packet loss retransmission system 900 according to another embodiment of the present application. As Figure 9 shown, Figure 9 the structural block diagram shown includes: a sender terminal 910, a server 920, and a receiver terminal 930; wherein, the sender terminal 910 includes a feature extraction module 911, a voice encoding module 912, and a data retransmission module 913, the server 920 includes a packet loss detection module 921, an audio routing module 922, a data retransmission module 923, and a perception analysis module 924, and the receiver terminal 930 includes a packet loss detection module 931, a voice decoding module 932, a mixing module 933, and a playback module 934.
[0196] Specifically, the feature extraction module 911 can collect an audio signal, extract features from the collected audio signal to obtain audio features, and send the collected audio signal to the voice encoding module 912. The voice encoding module 912 can encode the audio signal to obtain a corresponding audio bitstream, and send the audio bitstream to the data retransmission module 923. The data retransmission module 923 can package the matched audio bitstream and audio features into an audio data packet. The audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features corresponding to the audio bitstream. After packing multiple audio data packets, it is forwarded as a data packet to the server 920, so that the packet loss detection module 921 in the server 920 can perform packet loss detection on the received data packet.
[0197] If the packet loss detection result includes an audio data packet in a lost state, the packet loss detection module 921 feeds back the lost state to the data retransmission module 913, so that the data retransmission module 913 can retransmit data for this lost state. Furthermore, when the packet loss detection module 921 receives the retransmitted audio data packet and detects that there is no audio data packet in a lost state in the data packet, it updates the data packet according to the retransmitted audio data packet and sends the updated data packet to the audio routing module 922. Furthermore, the audio routing module 922 can screen the audio data packets in the data packet according to the audio features corresponding to each data packet in the data packet to obtain target audio data packets. Among them, the maximum energy amplitude corresponding to the target audio data packet is greater than that of other audio data packets in the data packet, or the similarity between the energy spectrum distribution corresponding to the target audio data packet and the preset energy spectrum distribution is greater than that of other audio data packets in the data packet. Furthermore, the audio routing module 922 can send the target audio data packets to the perception analysis module 924 and the data retransmission module 923, so that the data retransmission module 923 can forward the target audio data packets to the receiving terminal 930. Furthermore, the packet loss detection module 931 in the receiving terminal 930 performs packet loss detection on the target audio data packets. Among them, the number of target audio data packets can be one or more.
[0198] If the lost packet detection result includes target audio data packets in a lost state, the lost packet detection module 931 feeds back the lost state to the data retransmission module 923. Furthermore, if the loudness of a target audio data packet is greater than a preset loudness, the data retransmission module 923 can retransmit data for this lost state; among them, the target audio data packets in a lost state can be one or more. Furthermore, when the lost packet detection module 931 receives the retransmitted target audio data packets and detects that there are no target audio data packets in a lost state among all the target audio data packets, it updates all the target audio data packets according to the retransmitted target audio data packets, and sends the updated all target audio data packets to the voice decoding module 932. Furthermore, the voice decoding module 932 can decode the audio code stream in each target audio data packet into an audio signal, and send the decoded audio signal to the mixing module 933, so that the mixing module 933 performs mixing processing on the audio signal. Furthermore, the mixing module 933 can send the mixing processing result to the playback module 934, so that the playback module 934 plays the mixing processing result.
[0199] It can be seen that implementing Figure 9 the system shown can use the loudness corresponding to the audio data packet as the data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency. In addition, the target audio data packets can be retransmitted in a targeted manner. Compared with the prior art of retransmitting all the data packets in the current transmission including the target audio data packets, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0200] Furthermore, in this embodiment of the present example, optionally, the number of the above-mentioned sender terminals and receiver terminals can both be multiple, and the above-mentioned server can be a server cluster. Please refer to Figure 10 , Figure 10 which schematically shows a structural block diagram of a lost packet retransmission system 1000 according to another embodiment of the present application. As Figure 10 shown, it includes sender terminals 1011, sender terminals 1012, ……, sender terminals 101n, a server cluster 1030, receiver terminals 1021, receiver terminals 1022, ……, receiver terminals 102n; where n is a positive integer and greater than or equal to 3.
[0201] See Figure 10It can be known that the server cluster 1030 in the present application can receive data packets sent from at least one of the sender terminals 1011, 1012, ……, 101n. Each data packet may include one or more audio data packets. If the audio data packet meets the preset conditions, the audio data packet can be determined by the server cluster 1030 as a target audio data packet and forwarded to the receiver terminals 1021, 1022, ……, 102n. When receiving a packet loss status indicating the loss of the target audio data packet sent from any one of the receiver terminals 1021, 1022, ……, 102n, the target audio data packet is retransmitted according to the loudness corresponding to the target audio data packet.
[0202] It can be seen that implementing Figure 10 the system shown can use the loudness corresponding to the audio data packet as the data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency. In addition, the target audio data packet can be retransmitted in a targeted manner. Compared with the prior art of retransmitting all data packets of the current transmission including the target audio data packet, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0203] Furthermore, in the present exemplary embodiment, a packet loss retransmission device is also provided. Referring to Figure 11 the figure shown, the packet loss retransmission device 1100 may include:
[0204] A loudness acquisition unit 1101, configured to acquire the loudness corresponding to the target audio data packet;
[0205] A data sending unit 1102, configured to retransmit the target audio data packet according to the loudness corresponding to the target audio data packet when receiving a packet loss status indicating the loss of the target audio data packet.
[0206] It can be seen that implementing Figure 11 the device shown can use the loudness corresponding to the audio data packet as the data retransmission condition, thereby improving the problem of long data retransmission time and enhancing the data transmission efficiency. In addition, the target audio data packet can be retransmitted in a targeted manner. Compared with the prior art of retransmitting all data packets of the current transmission including the target audio data packet, the amount of retransmitted data can be reduced, and the occupation of network resources can be reduced.
[0207] In an exemplary embodiment of the present application, the above device further includes a data packet screening unit (not shown), where:
[0208] The data packet screening unit is configured to screen out target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets before the loudness acquisition unit 1101 acquires the loudness corresponding to the target audio data packet.
[0209] It can be seen that implementing this optional embodiment can screen the received audio data packets through audio features to reduce the amount of data forwarded, thereby reducing the consumption of network resources and improving the data transmission efficiency.
[0210] In an exemplary embodiment of the present application, the above device further includes a packet loss detection unit (not shown) and a packet loss status feedback unit (not shown), where:
[0211] The packet loss detection unit is configured to perform packet loss detection on the received multiple audio data packets before the data packet screening unit screens the target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets;
[0212] The packet loss status feedback unit is configured to, if the packet loss detection result includes a loss status, feedback the loss status to the sending terminal so that the sending terminal retransmits the data for the loss status.
[0213] Among them, the audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features corresponding to the audio bitstream. The audio features corresponding to the audio bitstream include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream.
[0214] It can be seen that implementing this optional embodiment can trigger the sender to retransmit the data when a packet loss status is detected, which can improve the integrity of the transmitted data.
[0215] In an exemplary embodiment of the present application, the above device further includes a data packet update unit (not shown), where:
[0216] The data packet update unit is configured to update the multiple audio data packets according to the retransmitted data packets after the sending terminal retransmits the data for the loss status and before the data packet screening unit screens the target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets.
[0217] It can be seen that implementing this optional embodiment can ensure the integrity of the data sent to the receiving terminal through data packet update.
[0218] In an exemplary embodiment of the present application, the received multiple audio data packets are sent by the sending terminal;
[0219] Among them, the specific manner in which the sending terminal sends multiple audio data packets is as follows:
[0220] The sending terminal collects an audio signal and extracts features from the audio signal to obtain audio features;
[0221] The sending terminal encodes the audio signal to obtain an audio bitstream;
[0222] The sending terminal packs the audio bitstream and the audio features into an audio data packet and sends it to the server.
[0223] It can be seen that by implementing this optional embodiment, the sending terminal can extract features from the audio signal, which is beneficial for the server to select the path signal according to the result of feature extraction, so as to ensure the real-time performance of the audio output in a multi-person conference while also ensuring the audio output effect.
[0224] In an exemplary embodiment of the present application, the preset condition includes a preset energy amplitude and / or a preset signal-to-noise ratio. The data packet screening unit screens target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets, including:
[0225] If at least one energy amplitude greater than the preset energy amplitude is detected in the audio features, the audio data packet to which the audio bitstream corresponding to the audio features belongs is determined as the target audio data packet; and / or,
[0226] If at least one signal-to-noise ratio greater than the preset signal-to-noise ratio is detected in the audio features, the audio data packet to which the audio bitstream corresponding to the audio features belongs is determined as the target audio data packet.
[0227] It can be seen that by implementing this optional embodiment, the audio data packets can be screened by the energy amplitude or the signal-to-noise ratio to streamline the audio data packets to be transmitted and improve the data transmission efficiency.
[0228] In an exemplary embodiment of the present application, it further includes:
[0229] Before the loudness acquisition unit of the sending terminal acquires the loudness corresponding to the target audio data packet, the sending terminal performs frame division processing on the audio bitstream according to a preset duration to obtain multiple audio frames;
[0230] The sending terminal processes the multiple audio frames through a preset window function respectively to obtain multiple reference frames;
[0231] The sending terminal calculates the power spectrum corresponding to each of the multiple reference frames;
[0232] The sending terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum.
[0233] Wherein, the preset window function is a Hanning window function, a Hamming window function, a Blackman window function, a Kaiser window function, a triangular window function or a rectangular window function.
[0234] It can be seen that implementing this optional embodiment can calculate the audio loudness of the target audio data packet, and this audio loudness can be used as a data retransmission condition. When the audio loudness is low, the server can refrain from retransmitting the lost target audio data packet to reduce the occupancy of network resources.
[0235] In an exemplary embodiment of the present application, the sender terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum, including:
[0236] The sender terminal calculates the frequency point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum;
[0237] The sender terminal calculates the loudness weight of each frequency point in the power spectrum according to the frequency point loudness;
[0238] The sender terminal calculates the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum;
[0239] The sender terminal determines the sum of the loudness values corresponding to multiple reference frames as the loudness corresponding to the target audio data packet.
[0240] It can be seen that implementing this optional embodiment can determine the loudness corresponding to the target audio data packet as a data retransmission condition according to the loudness corresponding to each frequency point, thereby improving the data retransmission efficiency of the server.
[0241] In an exemplary embodiment of the present application, the data sending unit 1102 retransmits the target audio data packet according to the loudness corresponding to the target audio data packet, including:
[0242] If the loudness corresponding to the target audio data packet is greater than the preset loudness, the target audio data packet is retransmitted to the receiving terminal so that the receiving terminal decodes and outputs the target audio data packet after and before retransmission.
[0243] It can be seen that implementing this optional embodiment can screen the audio data packets according to the loudness, reduce the occupancy of network resources by sounds with weak human ear perception ability, and further improve the data transmission efficiency. In addition, since the present application can retransmit the lost data packets, it can solve the problem of low transmission efficiency caused by retransmitting all the data of the entire data packet every time in the prior art.
[0244] In an exemplary embodiment of the present application, the manner in which the receiving terminal decodes and outputs the target audio data packet after and before retransmission is specifically:
[0245] The receiving terminal decodes the target audio data packet after and before retransmission to obtain multiple audio signals to be output;
[0246] The receiving terminal mixes multiple audio signals to be output, obtains a mixed signal and plays it.
[0247] It can be seen that by implementing this optional embodiment, the receiving terminal can mix and output after receiving a complete set of multiple audio data packets. The output result retains the important audio content in the meeting and discards the audio content with weak human ear perception ability, which can improve the message forwarding efficiency of multi-person meetings, ensure the real-time nature of the output audio signal, and thus improve the user experience.
[0248] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0249] Since each functional module of the packet loss retransmission device in the exemplary embodiment of the present application corresponds to the steps of the exemplary embodiment of the above packet loss retransmission method, for details not disclosed in the device embodiment of the present application, please refer to the above embodiments of the packet loss retransmission method of the present application.
[0250] On the other hand, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.
[0251] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0252] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0253] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0254] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.
[0255] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A packet loss retransmission method, characterized in that, Including: Screening target audio data packets whose audio features meet preset conditions from the received multiple audio data packets to reduce the amount of data to be forwarded, where the audio features include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream; Obtaining the loudness corresponding to the target audio data packet; When a packet loss status indicating the loss of the target audio data packet is received, if the loudness corresponding to the target audio data packet is greater than or equal to a preset loudness, the target audio data packet is retransmitted, and if the loudness corresponding to the target audio data packet is less than the preset loudness, retransmission of the target audio data packet is abandoned, where the preset loudness corresponds to the perceptibility of the human ear.
2. The method according to claim 1, wherein Before screening target audio data packets whose audio features meet preset conditions from the received multiple audio data packets, the method further includes: Performing packet loss detection on the received multiple audio data packets; If the packet loss detection result includes a loss status, the loss status is fed back to the sending terminal so that the sending terminal retransmits data for the loss status.
3. The method according to claim 2, wherein After the sending terminal retransmits data for the loss status and before screening target audio data packets whose audio features meet preset conditions from the received multiple audio data packets, the method further includes: Updating the multiple audio data packets according to the retransmitted data packets.
4. The method according to claim 1, wherein The audio data packet includes the loudness corresponding to the audio data packet, the audio bitstream, and the audio features.
5. The method according to claim 4, wherein The received multiple audio data packets are sent by a sending terminal; Among them, the specific manner in which the sending terminal sends the multiple audio data packets is: The sending terminal collects an audio signal and extracts features from the audio signal to obtain the audio features; The sending terminal encodes the audio signal to obtain the audio bitstream; The sending terminal packs the audio bitstream and the audio features into the audio data packet and sends it to the server.
6. The method according to claim 4, wherein The preset conditions include a preset energy amplitude and / or a preset signal-to-noise ratio. Screening target audio data packets whose audio features meet the preset conditions from the received multiple audio data packets includes: If at least one energy amplitude greater than the preset energy amplitude is detected in the audio features, determining the audio data packet to which the audio bitstream corresponding to the audio features belongs as the target audio data packet; and / or, If at least one signal-to-noise ratio greater than the preset signal-to-noise ratio is detected in the audio features, determining the audio data packet to which the audio bitstream corresponding to the audio features belongs as the target audio data packet.
7. The method according to claim 1, wherein Before obtaining the loudness corresponding to the target audio data packet, the method further includes: The sending terminal performs frame division on the audio bitstream according to a preset duration to obtain multiple audio frames; The sending terminal processes the multiple audio frames through a preset window function respectively to obtain multiple reference frames; The sending terminal calculates the power spectrum corresponding to each of the multiple reference frames; The sending terminal calculates the loudness corresponding to the target audio data packet according to the power spectrum.
8. The method according to claim 7, characterized in that The preset window function is a Hanning window function, a Hamming window function, a Blackman window function, a Kaiser window function, a triangular window function, or a rectangular window function.
9. The method according to claim 7, wherein The sender terminal calculates the loudness corresponding to the target audio packet according to the power spectrum, including: The sender terminal calculates the frequency point loudness of each frequency point in the power spectrum according to the energy amplitude of each frequency point in the power spectrum; The sender terminal calculates the loudness weight of each frequency point in the power spectrum according to the frequency point loudness; The sender terminal calculates the weighted sum between the energy amplitude of each frequency point in the power spectrum and the loudness weight of each frequency point in the power spectrum as the loudness value of the reference frame corresponding to the power spectrum; The sender terminal determines the sum of the loudness values corresponding to the multiple reference frames as the loudness corresponding to the target audio packet.
10. The method according to claim 1, characterized in that, Retransmitting the target audio packet according to the loudness corresponding to the target audio packet includes: If the loudness corresponding to the target audio packet is greater than the preset loudness, retransmit the target audio packet to the receiver terminal so that the receiver terminal decodes and outputs the target audio packet after retransmission and before retransmission.
11. The method according to claim 10, wherein The manner in which the receiver terminal decodes and outputs the target audio packet after retransmission and before retransmission is specifically: The receiver terminal decodes the target audio packet after retransmission and before retransmission to obtain a plurality of audio signals to be output; The receiver terminal performs a mixing process on the plurality of audio signals to be output to obtain a mixed signal and play it.
12. A packet loss retransmission system, characterized in that, Including a sender terminal, a server, and a receiver terminal, where: The sender terminal is configured to screen out a target audio packet whose audio features meet preset conditions from the received multiple audio packets, and send the target audio packet to the server to reduce the amount of data to be forwarded. The audio features include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream. The target audio packet includes the loudness corresponding to the target audio packet; The server is configured to obtain the loudness corresponding to the target audio packet; The receiver terminal is configured to send a packet loss status indicating that the target audio packet is lost to the server; The server is further configured to, when receiving the packet loss status, if the loudness corresponding to the target audio packet is greater than or equal to the preset loudness, retransmit the target audio packet, and if the loudness corresponding to the target audio packet is less than the preset loudness, abandon retransmitting the target audio packet. The preset loudness corresponds to the perceptibility of the human ear; The receiver terminal is further configured to receive the target audio packet.
13. A packet loss retransmission device, characterized in that, Including: A loudness acquisition unit, configured to screen out a target audio packet whose audio features meet preset conditions from the received multiple audio packets to reduce the amount of data to be forwarded, and acquire the loudness corresponding to the target audio packet. The audio features include the energy distribution corresponding to the audio bitstream and the energy amplitude corresponding to each frequency point in the audio bitstream; A data sending unit, configured to, when receiving a packet loss status indicating the loss of the target audio packet, retransmit the target audio packet if the loudness corresponding to the target audio packet is greater than or equal to a preset loudness, and abandon retransmitting the target audio packet if the loudness corresponding to the target audio packet is less than the preset loudness, where the preset loudness corresponds to the perceptibility of the human ear.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-11 is implemented.
15. An electronic device, characterized in that, Comprising: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1-11 by executing the executable instructions.
16. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of a computer device reads and executes the computer program, so that the computer device executes the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for allocating internet protocol (IP) network resources
CN102137438A
Loudness detection method and system
CN104105045A
Voice signal processing system of electronic communication device
CN110600049A