Terminal-side voice service end-to-end time delay reduction method, system and device and storage medium

By optimizing the voice packet transmission process, ensuring that the scheduling request SR is synchronized with the audio data sending cycle, and sending the SR in advance to reduce waiting time, the resource request delay problem when the voice packet reaches the MAC layer is solved, and low-latency and efficient voice data transmission is achieved.

CN120676465APending Publication Date: 2025-09-19GUANGZHOU HAIGE COMMUNICATION GROUP INCORPORATED COMPANY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510823884.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the prior art, the time when the voice packet arrives at the MAC layer is inconsistent with the sending time of the scheduling request SR, which increases the resource request time of the voice packet and further increases the end-to-end delay.

Method used

By optimizing the audio data collection, encoding, packaging, transmission, and resource allocation processes, we ensure that the scheduling request (SR) sending cycle is synchronized with the audio data sending cycle. SR is sent in advance to request resources just when the voice packet arrives, reducing waiting time.

Benefits of technology

It significantly reduces the total transmission time of voice data from the sender to the receiver, ensures the real-time performance and transmission efficiency of voice data, and improves the user experience, especially in low-latency communication scenarios such as online game voice chat and remote medical consultation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676465A_ABST
    Figure CN120676465A_ABST
Patent Text Reader

Abstract

The invention discloses a terminal-side voice service end-to-end time delay reduction method, system and device and a storage medium, under the condition that a scheduling request SR sending period TSR is T0 times of a voice packet sending period, the whole voice period is ensured, the sending time sequence of all voice packets is started from the initial collection time by controlling the initial voice collection time of an audio channel, and the time delay of the whole voice period is reduced. The method comprises the following steps of: sending voice packets which are sent at intervals of 20ms regularly, simultaneously, predicting the arrival time of each voice packet by a protocol stack MAC (Media Access Control) layer, triggering a scheduling request SR before the voice packets arrive at a physical layer, and ensuring that when the voice packets arrive, transmission resources applied by the scheduling request SR can just directly send out the voice packets which arrive at the physical layer. Therefore, the transmission time delay of the voice packet does not contain the SR sending waiting time, the authorization indication waiting time and the sending waiting time any more, and then the end-to-end time delay of the voice packet is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular to a method, system, device, and storage medium for reducing end-to-end delay of terminal-side voice services. Background Art

[0002] After the uplink audio frame is prepared, it is transmitted to the MAC layer. The MAC layer starts to apply for network resources for the transmission of the voice packet and initiates a scheduling request (SR). After the network receives the SR request from the terminal, it allocates transmission resources and informs the terminal. After receiving the network's resource indication, the terminal sends the voice packet on the allocated time-frequency resources. Therefore, after the voice packet reaches the MAC, it is not sent directly to the physical layer. Instead, there is a waiting time during the actual transmission. The waiting time includes the waiting time for the SR to be sent, the waiting time for resource authorization, and the waiting time for sending.

[0003] Therefore, when the time when the voice packet arrives at the MAC is inconsistent with the SR sending time, an SR sending waiting time will be generated. In particular, when the time when the voice packet arrives at the MAC lags behind the SR sending position, the time the voice packet waits for SR sending is close to the SR sending period. When the SR sending period is long, the resource request time of the voice packet will increase, thereby increasing the end-to-end delay. Summary of the Invention

[0004] The present invention aims to overcome at least one of the above-mentioned shortcomings of the prior art and provides a method, system, device and storage medium for reducing the end-to-end delay of terminal-side voice services. The method is used to solve the technical problem in the prior art that when the time when the voice packet arrives at the MAC is inconsistent with the SR sending time, an SR sending waiting time is generated, the resource request time of the voice packet is increased, and the end-to-end delay is increased.

[0005] The present invention provides a method for reducing the end-to-end delay of terminal-side voice services, which is used to optimize the transmission process of voice data from the sending end to the receiving end, reduce the delay of voice services during transmission, and thus improve the user experience, especially in real-time voice application scenarios requiring low-latency communication. The specific steps of this method include:

[0006] S01, the sending end collects audio data and encodes it to form voice frames;

[0007] S02, grouping the voice frames into voice packets;

[0008] S03, transmitting the voice packet to the MAC layer;

[0009] S04: The MAC layer initiates a scheduling request SR;

[0010] S05. After receiving the scheduling request SR, the network side allocates time-frequency resources for the voice packet and sends a resource indication to the terminal;

[0011] S06. After receiving the resource indication from the network side, the terminal sends the voice packet on the allocated time-frequency resource;

[0012] S07: The voice packet is sent.

[0013] By optimizing the audio data acquisition, encoding, packetization, transmission and resource allocation processes, the total time required for voice data to be received and decoded by the receiving end is significantly reduced, that is, the end-to-end latency is reduced, the real-time nature of voice data is ensured, and users can communicate with each other almost without delay. After initiating a scheduling request SR at the MAC layer and waiting for the network side to allocate video resources, the terminal can accurately send voice packets on the allocated resources, avoiding unnecessary waiting and resource waste, improving transmission efficiency, and being able to adapt to various network environments.

[0014] The present invention optimizes the transmission process of voice packets based on the regularity and periodicity of audio data transmission, especially the relationship between the sending timing of the scheduling request SR and the audio data sending period, so as to further reduce the end-to-end delay of the voice service. Since the characteristic of audio data transmission is regularity, once a voice call is established, audio data will be continuously collected and played. In order to ensure that the sending period of the scheduling request SR is coordinated with the sending period of the audio data to reduce unnecessary waiting time and resource waste, a frame of audio data will be sent every 20ms in the uplink sending direction, and a frame of playback data needs to be provided every 20ms in the downlink receiving direction. Therefore, the sending of voice packets is periodic, and the sending period of audio data in step S01 is T0 = 20ms.

[0015] After the uplink voice frame is prepared, it is packaged into a voice packet through audio and transmitted to the MAC layer. The MAC layer applies for network resources for the transmission of the voice packet and initiates a scheduling request SR. The time interval between the SR and the authorization indication received from the network is fixed. The period of the scheduling request SR is T SR It is a multiple of the audio data sending period T0:

[0016] T SR =nT0

[0017] Where n = 1, 2, 3, ..., represents T SR It is an integer multiple of T0, which makes the scheduling request and the sending of audio data more synchronized, reduces the additional delay caused by waiting for SR to be sent, and improves transmission efficiency.

[0018] By properly setting the SR sending cycle, the network side can more effectively allocate video resources for voice packets. At the same time, the waiting time of voice packets during transmission is reduced, including the time waiting for SR sending, waiting for resource authorization, and waiting for sending. This reduces end-to-end latency and improves user experience. This is especially true in voice communication scenarios that require fast response and high synchronization, such as online game voice chat and remote medical consultation. Users can experience a smoother and more natural communication experience.

[0019] When the SR period T SR When the audio data transmission period T0 is not an integer multiple, the transmission is performed according to the protocol specification. That is, the SR is triggered only after the voice packet reaches the MAC layer. At this time, the waiting time for the actual voice packet transmission includes the waiting time for SR transmission, the waiting time for resource authorization, and the waiting time for transmission.

[0020] After a voice packet arrives at the MAC layer, there is a waiting time before it is actually sent to the physical layer. It is necessary to wait for the scheduling request SR to apply for time and frequency resources for the voice packet. This waiting time includes the waiting time for the SR to be sent, the waiting time for resource authorization, and the waiting time for sending. To ensure that the scheduling request SR is triggered before the voice packet arrives at the MAC layer and that the requested transmission resources are available when the voice packet arrives, so as to achieve timely transmission of the voice packet, the voice packet transmission to the MAC layer in step S03 further includes:

[0021] The time t when the voice packet arrives at the MAC layer m satisfy:

[0022] t m =t s +Δt1+Δt2

[0023] Among them, t s Indicates the time when SR is sent, Δt1 indicates the time interval from when SR is sent to when the authorization indication from the network is received, and Δt2 indicates the time interval from when the authorization indication from the network is received to when the actual data is sent.

[0024] The above formula indicates that the Scheduling Request (SR) is triggered before the voice packet reaches the physical layer, ensuring that the transmission resources requested by the Scheduling Request (SR) are just sufficient to transmit the voice packet upon its arrival. This eliminates the need for the SR to be sent, the time it takes for the authorization indication to be indicated, and the time it takes for the voice packet to be transmitted, thereby reducing the end-to-end latency of the voice packet. By precisely controlling the timing of SR triggering and the resource allocation process, network resources can be more efficiently utilized, avoiding overallocation or waste of resources, thereby optimizing overall network performance.

[0025] Since voice data is sent every 20ms, in order to ensure the continuity of audio data, it is necessary to ensure that after each sending is completed, the next frame of audio data can be processed and sent immediately to maintain the continuity of the audio. Therefore, the completion of sending the voice packet in step S07 also includes: returning to step S01 according to the sending period T0 of the audio data to continuously process subsequent audio data.

[0026] By looping and processing subsequent audio data, real-time transmission of voice data is ensured, meeting the requirements of real-time voice communication for low latency and high continuity.

[0027] The present invention also provides a terminal-side voice service end-to-end delay reduction system, which is characterized by comprising:

[0028] Acquisition module: used to collect audio data and encode it into voice frames;

[0029] Processing module: used for grouping the voice frames into voice packets;

[0030] Transmission module: used for transmitting the voice packet to the MAC layer;

[0031] Request module: MAC layer initiates scheduling request SR;

[0032] Allocation module: After receiving the scheduling request SR, the network layer allocates time and frequency resources for the voice packet and sends a resource indication to the terminal;

[0033] Sending module: after receiving the resource indication from the network side, the terminal sends the voice packet on the allocated time-frequency resource;

[0034] End module: Voice packet sending is completed.

[0035] The present invention also provides a device, characterized in that it includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the method for reducing the end-to-end delay of terminal-side voice services.

[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is processed and executed, the steps of the method for reducing the end-to-end delay of the terminal-side voice service are implemented.

[0037] The present invention provides a method, system, device and storage medium for reducing the end-to-end delay of terminal-side voice services. When the SR sending period is a multiple of the voice packet sending period, the timing of the initial voice collection of the audio path is controlled to ensure that the sending timing of all voice packets in the entire voice cycle starts from the initial collection time and is sent regularly at intervals of 20ms. This precise timing control not only ensures the continuity and smoothness of the audio data, but also provides a basis for subsequent resource allocation and transmission optimization. At the same time, the MAC layer of the protocol stack needs to predict the arrival time of each voice packet and send the SR in advance before the voice packet arrives, thereby actively applying for the required transmission resources. The early prediction and resource reservation mechanism avoids the increase in delay caused by insufficient resources or allocation delays, and ensures that the transmission resource sending time requested by the SR can just directly send the incoming voice packet, realizing seamless transmission of the voice packet. This can effectively reduce the end-to-end delay of the voice service, thereby directly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 It is a schematic diagram of the steps of a method for reducing the end-to-end delay of terminal-side voice services.

[0040] Figure 2 This is a timing diagram when SR is not sent in advance in Example 2.

[0041] Figure 3 This is a timing diagram of the early transmission of SR in Example 2.

[0042] Figure 4 This is a timing diagram when SR is not sent in advance in Example 3.

[0043] Figure 5 This is a timing diagram of SR early transmission in Example 3. DETAILED DESCRIPTION

[0044] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings:

[0045] An embodiment of the present invention provides an embodiment of a method for reducing the end-to-end delay of terminal-side voice services. It should be noted that although the logical order is shown in the flowchart, under certain data, the steps shown or described can be completed in an order different from that here.

[0046] Example 1: Reference Figure 1

[0047] The present invention provides a method for reducing the end-to-end delay of a terminal-side voice service, which is characterized by comprising:

[0048] S01, the sending end collects audio data and encodes it to form voice frames;

[0049] S02, grouping the voice frames into voice packets;

[0050] S03, transmitting the voice packet to the MAC layer;

[0051] S04: The MAC layer initiates a scheduling request SR;

[0052] S05. After receiving the scheduling request SR, the network side allocates time-frequency resources for the voice packet and sends a resource indication to the terminal;

[0053] S06. After receiving the resource indication from the network side, the terminal sends the voice packet on the allocated time-frequency resource;

[0054] S07: The voice packet is sent.

[0055] By optimizing the audio data acquisition, encoding, packetization, transmission and resource allocation processes, the total time required for voice data to be received and decoded by the receiving end is significantly reduced, that is, the end-to-end latency is reduced, the real-time nature of voice data is ensured, and users can communicate with each other almost without delay. After initiating a scheduling request SR at the MAC layer and waiting for the network side to allocate video resources, the terminal can accurately send voice packets on the allocated resources, avoiding unnecessary waiting and resource waste, improving transmission efficiency, and being able to adapt to various network environments.

[0056] The present invention optimizes the transmission process of voice packets based on the regularity and periodicity of audio data transmission, especially the relationship between the sending timing of the scheduling request SR and the audio data sending period, so as to further reduce the end-to-end delay of the voice service. Since the characteristic of audio data transmission is regularity, once a voice call is established, the audio data will be continuously collected and played. In order to ensure that the sending period of the scheduling request SR is coordinated with the sending period of the audio data to reduce unnecessary waiting time and resource waste, the characteristic of audio data transmission is regularity. Once a voice call is established, the audio data will be continuously collected and played. A frame of audio data will be sent every 20ms in the uplink sending direction, and a frame of playback data needs to be provided every 20ms in the downlink receiving direction. Therefore, the sending of voice packets is periodic, and the sending period T0 of the audio data in step S01 is 20ms.

[0057] After the uplink voice frame is prepared, it is packaged into a voice packet through audio and transmitted to the MAC layer. The MAC layer applies for network resources for the transmission of the voice packet and initiates a scheduling request SR. The time interval between the SR and the authorization indication received from the network is fixed. The period of the scheduling request SR is T SR It is a multiple of the audio data sending period T0:

[0058] T SR =nT0

[0059] Where n = 1, 2, 3, ..., represents T SR It is an integer multiple of T0, which makes the scheduling request and the sending of audio data more synchronized, reduces the additional delay caused by waiting for SR to be sent, and improves transmission efficiency.

[0060] By properly setting the SR sending cycle, the network side can more effectively allocate video resources for voice packets. At the same time, the waiting time of voice packets during transmission is reduced, including the time waiting for SR sending, waiting for resource authorization, and waiting for sending. This reduces end-to-end latency and improves user experience. This is especially true in voice communication scenarios that require fast response and high synchronization, such as online game voice chat and remote medical consultation. Users can experience a smoother and more natural communication experience.

[0061] When the SR period T SR When the audio data transmission period T0 is not an integer multiple, the transmission is performed according to the protocol specification. That is, the SR is triggered only after the voice packet reaches the MAC layer. At this time, the waiting time for the actual voice packet transmission includes the waiting time for SR transmission, the waiting time for resource authorization, and the waiting time for transmission.

[0062] After a voice packet arrives at the MAC layer, there is a waiting time before it is actually sent to the physical layer. It is necessary to wait for the scheduling request SR to apply for time and frequency resources for the voice packet. This waiting time includes the waiting time for the SR to be sent, the waiting time for resource authorization, and the waiting time for sending. To ensure that the scheduling request SR is triggered before the voice packet arrives at the MAC layer and that the requested transmission resources are available when the voice packet arrives, so as to achieve timely transmission of the voice packet, the voice packet transmission to the MAC layer in step S03 further includes:

[0063] The time t when the voice packet arrives at the MAC layer m satisfy:

[0064] t m =t s +Δt1+Δt2

[0065] Among them, t sIndicates the time when SR is sent, Δt1 indicates the time interval from when SR is sent to when the authorization indication from the network is received, and Δt2 indicates the time interval from when the authorization indication from the network is received to when the actual data is sent.

[0066] The above formula indicates that the Scheduling Request (SR) is triggered before the voice packet reaches the physical layer, ensuring that the transmission resources requested by the Scheduling Request (SR) are just sufficient to transmit the voice packet upon its arrival. This eliminates the need for the SR to be sent, the time it takes for the authorization indication to be indicated, and the time it takes for the voice packet to be transmitted, thereby reducing the end-to-end latency of the voice packet. By precisely controlling the timing of SR triggering and the resource allocation process, network resources can be more efficiently utilized, avoiding overallocation or waste of resources, thereby optimizing overall network performance.

[0067] Since voice data is sent every 20ms, in order to ensure the continuity of audio data, it is necessary to ensure that after each sending is completed, the next frame of audio data can be processed and sent immediately to maintain the continuity of the audio. Therefore, the completion of sending the voice packet in step S07 also includes: returning to step S01 according to the sending period T0 of the audio data to continuously process subsequent audio data.

[0068] By looping and processing subsequent audio data, real-time transmission of voice data is ensured, meeting the requirements of real-time voice communication for low latency and high continuity.

[0069] In the case where the SR sending period is a multiple of the voice packet sending period, the present invention controls the moment of initial voice collection in the audio path to ensure that the sending timing of all voice packets in the entire voice period starts from the initial collection moment and is sent regularly at intervals of 20ms. This precise timing control not only ensures the continuity and fluency of the audio data, but also provides a basis for subsequent resource allocation and transmission optimization. At the same time, the MAC layer of the protocol stack needs to predict the arrival time of each voice packet and send SR in advance before the voice packet arrives, so as to actively apply for the required transmission resources. The advance prediction and resource reservation mechanism avoids the increase in delay caused by insufficient resources or allocation delays, and ensures that the transmission resource sending moment requested by SR can just directly send out the incoming voice packet, thereby realizing seamless transmission of voice packets, which can effectively reduce the end-to-end delay of voice services, thereby directly improving user experience.

[0070] Example 2: Reference Figure 2 and Figure 3

[0071] Sending period T with scheduling request SR SR = 20ms as an example, analyze the solution to reduce the time, from Figure 2 and Figure 3By comparing the results, we can see that by triggering SR sending in advance, the voice packet can be sent directly when it reaches the MAC. In this way, the waiting time for the voice packet to be sent is zero. For the entire voice lifecycle, the end-to-end delay of the voice packet does not include the time waiting for resource sending, thereby reducing the end-to-end delay of the audio path.

[0072] Example 3: Reference Figure 4 and Figure 5

[0073] When the scheduling request SR is sent periodically T SR =40ms, compared Figure 4 and Figure 5 , we can see that if the coordinated SR is sent in advance and the time when the 2ith (i=1,2,3,4…) voice packet arrives at the MAC is consistent with the time when the data is actually sent, the sending waiting time of two consecutive voice packets is 20ms, as shown in Figure 5 As shown in the figure, assuming that voice packet 4 can be sent out just when it arrives, the waiting time is 0, and the waiting time of voice packet 3 is 20ms. In this way, the total waiting time for voice packets 3 and 4 is 20ms. If the SR is not sent in advance, the waiting time for sending two consecutive voice packets will exceed 20ms. Figure 4 As shown, the SR period is 40ms, and SR is sent later than voice packet 2 and voice packet 3. By the time the data can actually be sent, the sending waiting time generated by voice packet 2 exceeds 20ms, and the sending waiting time generated by voice packet 3 is less than 20ms. In this way, the total sending waiting time of voice packet 2 and voice packet 3 exceeds 20ms.

[0074] From the above analysis, we can see that when the SR sending period is twice the voice sending period, controlling the SR sending in advance to ensure that the sending waiting time of the second data packet of two consecutive data packets is 0 can still minimize the overall sending waiting time, thereby achieving the goal of reducing end-to-end transmission delay.

[0075] The above two sets of analysis show that in voice services, when the SR sending period is a multiple of the voice sending period, sending the SR in advance ensures that the voice data is ready and can be sent directly when the resources requested by the SR are sent. This solution can effectively reduce the end-to-end latency of voice services and improve the voice call experience.

Claims

1. A method for reducing the end-to-end delay of terminal-side voice services, characterized in that: include: S01, the sending end collects audio data and encodes it to form voice frames; S02, grouping the voice frames into voice packets; S03, transmitting the voice packet to the MAC layer; S04: The MAC layer initiates a scheduling request SR; S05. After receiving the scheduling request SR, the network side allocates time-frequency resources for the voice packet and sends a resource indication to the terminal; S06. After receiving the resource indication from the network side, the terminal sends the voice packet on the allocated time-frequency resource; S07: The voice packet is sent.

2. A method for reducing end-to-end delay of terminal-side voice service according to claim 1, characterized in that: In step S01 , the audio data transmission period T0 is 20 ms.

3. The method for reducing the end-to-end delay of terminal-side voice service according to claim 2, characterized in that: The period T of the scheduling request SR in step S04 SR It is a multiple of the audio data sending period T0: T SR =nT0 Where n = 1, 2, 3, ..., represents T SR It is an integer multiple of T0.

4. The method for reducing end-to-end delay of voice service according to claim 1, characterized in that: The transmission of the voice packet to the MAC layer in step S03 further includes: The time t when the voice packet arrives at the MAC layer m satisfy: t m =t s +Δt1+Δt2 Among them, t s Indicates the time when SR is sent, Δt1 indicates the time interval from when SR is sent to when the authorization indication from the network is received, and Δt2 indicates the time interval from when the authorization indication from the network is received to when the actual data is sent.

5. A method for reducing end-to-end delay of terminal-side voice service according to claim 4, characterized in that: Step S05 also includes triggering the scheduling request SR before the voice packet reaches the physical layer.

6. A method for reducing end-to-end delay of terminal-side voice service according to claim 5, characterized in that: The sending time when the scheduling request SR applies for transmission resources is consistent with the time when the voice packet arrives at the physical layer, so that the voice packet is sent out.

7. The method for reducing the end-to-end delay of terminal-side voice service according to claim 2, characterized in that: The completion of sending the voice packet in step S07 further includes: returning to step S01 according to the sending period T0 of the audio data to continuously process subsequent audio data.

8. A terminal-side voice service end-to-end delay reduction system, characterized in that: include: Acquisition module: used to collect audio data and encode it into voice frames; Processing module: used for grouping the voice frames into voice packets; Transmission module: used for transmitting the voice packet to the MAC layer; Request module: MAC layer initiates scheduling request SR; Allocation module: After receiving the scheduling request SR, the network layer allocates time and frequency resources for the voice packet and sends a resource indication to the terminal; Sending module: after receiving the resource indication from the network side, the terminal sends the voice packet on the allocated time-frequency resource; End module: Voice packet sending is completed.

9. A device, characterized in that The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for reducing the end-to-end delay of the terminal-side voice service according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is processed and executed, the steps of the method for reducing the end-to-end delay of the terminal-side voice service are implemented as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Voice service processing method and device

    CN109152000A

  • Uplink resource pre-application method and related equipment

    CN114765831A

  • Time matching method, terminal and network side equipment

    CN115604732A

  • Pre-scheduling method, electronic equipment and system

    CN118250803A