Resource scheduling method and apparatus

By dynamically scheduling the transmission resources of encoded frames and determining the remaining latency budget based on the actual time, the problem of unstable video playback caused by network jitter is solved, improving user experience and resource utilization.

CN116569597BActive Publication Date: 2026-05-22HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-01-29
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

When there is network jitter, the terminal device cannot play the video at the preset frame rate, resulting in stuttering, lag, frame skipping and other phenomena, causing poor user experience such as screen misalignment and jumping, and degrading the user interaction experience.

Method used

The first network element determines the remaining delay budget based on the actual arrival time of the encoded frame, dynamically schedules transmission resources, prioritizes scheduling delayed encoded frames, delays the transmission of early-arriving encoded frames, and takes into account other service needs to achieve flexible resource allocation.

Benefits of technology

It alleviates stuttering, lag, and frame skipping caused by network jitter, reduces screen misalignment and jumping, and improves resource utilization and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116569597B_ABST
    Figure CN116569597B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a resource scheduling method and device. The method comprises: receiving a first encoded frame; determining a residual delay budget of the first encoded frame according to an actual time at which the first encoded frame arrives at a first network element; determining a scheduling priority according to the residual delay budget; and scheduling a transmission resource for transmitting the first encoded frame for the first encoded frame based on the scheduling priority. In a case where the scheduling priority is less than or equal to a preset priority threshold, the transmission resource is preferentially scheduled for the first encoded frame, and in a case where the scheduling priority is greater than the priority threshold, the delay of the first encoded frame is increased, so that resources can be flexibly scheduled for encoded frames that arrive in advance and lag at the first network element, and the phenomenon of lag, lag, frame skipping and the like caused by network jitter in the process of encoded stream transmission is alleviated, and adverse experiences such as picture misalignment and skipping are reduced, which is conducive to improving user interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more specifically, to a resource scheduling method and apparatus. Background Technology

[0002] With the development of communication technology, real-time multimedia services such as cloud extended reality (Cloud XR) and cloud gaming have also developed rapidly. Multimedia service data, such as video, after being encoded by a cloud server, can reach the terminal device for decoding and playback via multi-hop transmission.

[0003] However, network jitter is common in actual transmission, and in severe cases, it can exceed 20 milliseconds (ms). In the presence of jitter, terminal devices may be unable to play videos at the preset frame rate, and playback may experience phenomena such as stuttering, lag, and frame skipping. This can result in poor user experience, such as screen misalignment and skipping, and a decline in user interaction. Summary of the Invention

[0004] This application provides a resource scheduling method and apparatus to alleviate stuttering, lag, and frame skipping caused by network jitter during encoded stream transmission, and to reduce poor user experience such as screen misalignment and skipping.

[0005] Firstly, a resource scheduling method is provided, which can be executed by a first network element or by a component (such as a chip, chip system, etc.) configured in the first network element. This application does not limit this approach.

[0006] For example, the method includes: a first network element receiving a first coded frame; the first network element determining a remaining delay budget for the first coded frame based on the actual time the first coded frame arrives at the first network element; and the first network element scheduling transmission resources for the first coded frame based on the remaining delay budget.

[0007] The remaining delay budget for the first coded frame in the first network element is determined based on the actual time it arrives. This means the remaining delay budget is dynamically updated in real-time according to the actual transmission status of the current frame, and resources are then allocated to the current frame based on this budget. For example, resources can be prioritized when the remaining delay budget is small, and delayed when it is large. In this way, the first network element can flexibly allocate transmission resources for the first coded frame based on the dynamic remaining delay budget. This allows for prioritizing resources for delayed coded frames, mitigating stuttering, delays, and frame skipping caused by network jitter, and reducing poor user experience such as image misalignment and jumps. It also allows for balancing the transmission needs of other services, prioritizing resources for more urgent services. Overall, resources are flexibly allocated, thus improving resource utilization.

[0008] In conjunction with the first aspect, in some possible implementations of the first aspect, scheduling transmission resources for the first coded frame based on the remaining delay budget includes: the first network element determining a scheduling priority based on the remaining delay budget; the first network element scheduling the transmission resources for the first coded frame based on the scheduling priority; wherein, the smaller the remaining delay budget, the higher the scheduling priority of the transmission resources. In one possible design, the scheduling priority can be represented by a priority value, with a smaller priority value indicating a higher scheduling priority. Optionally, the priority value is inversely proportional to the remaining delay budget.

[0009] Therefore, multiple scheduling priorities can be allocated for multiple different remaining delay budgets, and transmission resources can be scheduled based on multiple different scheduling priorities, thereby enabling more flexible resource allocation and improving resource utilization.

[0010] This application provides two possible implementations for determining the remaining delay budget of the first coded frame. These two implementations will be described below.

[0011] In a first possible implementation, the first network element can determine the remaining delay budget of the first coded frame based on the time interval between the actual time and the expected time of the first coded frame arriving at the first network element, and the maximum delay budget allocated to the first network element. That is, the remaining delay budget is the time interval between the actual time and the expected time of the first coded frame arriving at the first network element.

[0012] The expected time is: the latest time that the first network element is expected to send the first encoded frame, or the latest time that the first encoded frame is expected to be transmitted to the second network element, or the latest time that the first encoded frame is expected to be decoded by the second network element.

[0013] Here, the desired latest time is also the acceptable latest time.

[0014] The expected time is the latest time that the first network element is expected to send the first coded frame, which means the latest acceptable time for the first network element to send the first coded frame. In other words, the first coded frame should not be sent later than this expected time.

[0015] The expected time is the latest time at which the first coded frame is expected to be transmitted from the first network element to the second network element. In other words, the first coded frame should arrive at the second network element no later than this expected time.

[0016] The expected time is the latest time that the first encoded frame is expected to be decoded by the second network element. In other words, the first encoded frame should not be decoded later than this expected time. For example, for multimedia services, one possible manifestation of the first encoded frame being decoded is that the first encoded frame can be played. Therefore, the expected time could refer to the latest time that the first encoded frame of the multimedia service can be played on the second network element.

[0017] In this embodiment of the application, the remaining delay budget is determined based on the actual time and expected time of the arrival of the first coded frame at the first network element. Resources can be scheduled for the first coded frame in the remaining delay budget so that the actual time of the first coded frame being sent from the first network element is no later than the latest time of the desired transmission, or the actual time of the first coded frame being transmitted to the second network element is no later than the latest time of the desired arrival, or the actual time of the first coded frame being decoded in the second network element is no later than the latest time of the desired decoding completion.

[0018] In conjunction with the first aspect, in some possible implementations of the first aspect, the first network element determines the expected time based on the actual arrival times of multiple coded frames arriving at the first network element within a first time period, the ideal time for the first coded frame to arrive at the first network element, the maximum delay budget allocated to the first network element, and a predefined frame drop rate threshold; wherein, the end time of the first time period is the actual time for the first coded frame to arrive at the first network element, and the duration of the first time period is a predefined value; the ideal time for the first coded frame to arrive at the first network element is obtained based on learning the network jitter of the transmission link between the encoding end of the first coded frame and the first network element; the maximum delay budget is determined based on the end-to-end round trip time (RTT) of the first coded frame.

[0019] Since network jitter is unstable, by learning network jitter, the ideal time for the first coded frame to arrive at the first network element can be predicted relatively accurately. Furthermore, as time progresses, subsequent coded frames can be processed sequentially as the first coded frame to determine the ideal time of arrival at the first network element; that is, the ideal time for the coded frame to arrive at the first network element can also be updated in real time.

[0020] Furthermore, it should be noted that the end-to-end RTT of the encoded stream is related to the end-to-end RTT of each encoded frame in the encoded stream. During transmission, the end-to-end RTT of multiple encoded frames in the encoded stream may be equal to the end-to-end RTT of the encoded stream.

[0021] It should be understood that the actual time when the first coded frame arrives at the first network element may or may not be included in the first time period. In other words, the multiple coded frames arriving at the first network element within the first time period may or may not include the first coded frame. This application does not limit this.

[0022] It should also be understood that the multiple coded frames within the first time period (excluding the first coded frame) and the first coded frame may belong to the same coded stream or to different related coded streams (such as the primary stream and the enhancement stream), and this application embodiment does not limit this.

[0023] Further, the first network element determines the expected time based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period, the ideal time for the first coded frame to arrive at the first network element, the maximum latency budget allocated to the first network element, and a predefined frame drop rate threshold. This includes: the first network element determining a candidate expected time for the first coded frame to arrive at the first network element based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period and the frame drop rate threshold; and the first network element determining the expected time based on the candidate expected time, the ideal time for the first coded frame to arrive at the first network element, and the maximum latency budget.

[0024] Furthermore, the first network element determines a candidate expected time for the first coded frame to arrive at the first network element based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period and the frame drop rate threshold. This includes: the first network element learning the network jitter of the transmission link between the encoding end and the first network element based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period; the first network element predicting the ideal time for the first coded frame to arrive at the first network element based on the network jitter; the first network element determining a delay distribution, which indicates the offset of the actual time of each coded frame arriving at the first network element relative to the ideal time, and the number of frames corresponding to different offsets; and the first network element determining a candidate expected time for the first coded frame to arrive at the first network element based on the delay distribution and the frame drop rate threshold, wherein the proportion of the number of frames corresponding to the offset of the candidate expected time relative to the ideal time of the first coded frame arriving at the first network element in the delay distribution is less than 1-φ, where φ is the frame drop rate threshold, 0 < φ < 1.

[0025] Specifically, this percentage can refer to the proportion of the number of frames corresponding to the offset of the candidate expected time from the ideal time of the first coded frame arriving at the first network element in the delay distribution to the total number of frames in the delay distribution.

[0026] To ensure the normal playback of the encoded stream, determining the expected time ensures that the transmission of encoded frames meets a predefined frame drop rate threshold, such as being less than (or less than or equal to) the frame drop rate threshold, thus enabling the encoded stream to play normally.

[0027] Since the determination of this expected time also needs to take into account the maximum delay budget, for ease of distinction, the expected time determined based on the actual time of the arrival of the multiple coded frames in the first time period to the first network element and the frame drop rate threshold is recorded as the candidate expected time.

[0028] Furthermore, the first network element determines the expected time based on the candidate expected time, the ideal time of the first coded frame arriving at the first network element, and the maximum delay budget, including: when the time interval between the ideal time of the first coded frame arriving at the first network element and the candidate expected time is greater than or equal to the maximum delay budget, the first network element determines the expected time of the first coded frame based on the maximum delay budget, such that the time interval between the expected time of the first coded frame and the ideal time of the first coded frame arriving at the first network element is the maximum delay budget; or, when the time interval between the ideal time of the first coded frame arriving at the first network element and the candidate expected time is less than the maximum delay budget, the first network element determines the candidate expected time as the expected time.

[0029] To avoid excessive end-to-end RTT of the encoded stream, the time interval between the above-mentioned expected time and the ideal time for the first encoded frame to arrive at the first network element should not be less than the maximum delay budget allocated to the first network element.

[0030] Therefore, when determining the expected time, both the candidate expected time obtained from network jitter factors and the maximum remaining delay budget are considered. Based on both, smooth transmission of coded frames is achieved, improving the user experience.

[0031] In conjunction with the first aspect, in some possible implementations of the first aspect, the network jitter of the transmission link between the first network elements is learned by linear regression or Kalman filtering based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period.

[0032] Because network jitter is unstable, the first network element can learn from the actual arrival times of multiple coded frames that arrived before the first coded frame to simulate real network jitter, and thus can more accurately predict the ideal time for the first coded frame to arrive at the first network element.

[0033] It should be understood that the linear regression and Kalman filtering methods are merely examples and should not be construed as limiting this application. Learning network jitter can also be achieved using other algorithms, including but not limited to those described in this application.

[0034] In conjunction with the first aspect, in some possible implementations of the first aspect, the first network element determines the expected time of the second encoded frame based on the expected time of the first encoded frame and the frame rate of the encoded stream, wherein the second encoded frame is the next frame after the first encoded frame.

[0035] For a given encoded stream, each frame can be used as the first encoded frame, and the expected time can be determined according to the method provided in the embodiments of this application, thereby determining the remaining latency budget. Alternatively, some frames can be used as the first encoded frames, and the expected time of the remaining frames can be determined according to the frame rate of the encoded stream, thereby determining the remaining latency budget. In a second possible implementation, the remaining latency budget is determined by the end-to-end RTT of the first encoded frame, the instruction processing time of the application layer device, the processing time of the encoding end, the actual time when the first encoded frame is sent from the encoding end, the actual time when the first encoded frame arrives at the first network element, and the processing time of the decoding end.

[0036] This implementation requires the participation of other network elements used to transmit the encoded stream. For example, the encoding end can carry the time of sending each encoded frame in the data packet so that the first network element can determine the expected delay budget for each encoded frame.

[0037] In conjunction with the first aspect, in some possible implementations of the first aspect, scheduling transmission resources for the first coded frame based on the scheduling priority includes: when the scheduling priority is less than or equal to (or less than) a preset priority threshold, the first network element preferentially schedules transmission resources for the first coded frame, and the number of physical resource blocks (PRBs) in the scheduled transmission resources is greater than or equal to the number of PRBs required for the transmission of the first coded frame.

[0038] In conjunction with the first aspect, in some possible implementations of the first aspect, scheduling transmission resources for the first encoded frame based on the scheduling priority includes: when the scheduling priority is greater than (or greater than or equal to) a preset priority threshold, the first network element increases the latency of the first encoded frame.

[0039] Therefore, if the first coded frame arrives late, the corresponding remaining delay budget is small, and the scheduling priority of transmission resources is high. Sufficient physical resources can then be allocated to ensure the first coded frame is transmitted within that remaining delay budget. Conversely, if the first coded frame arrives early, the corresponding remaining delay budget is large, and the scheduling priority of transmission resources is low. In this case, the delay can be increased, and resources don't need to be allocated hastily for the first coded frame. Resources can then be used for other more urgent services. This allows for flexible resource allocation, which is beneficial for improving resource utilization.

[0040] It should be understood that the above description of the relationship between scheduling priority and priority threshold is based on the following assumption: the smaller the priority value, the higher the priority; the larger the priority value, the lower the priority. However, this should not constitute any limitation on this application. Those skilled in the art will understand that if the larger the priority value, the higher the priority; and the smaller the priority value, the lower the priority, then the above relationship between scheduling priority and priority threshold can be changed accordingly, and these changes should also fall within the protection scope of this application.

[0041] Optionally, the transmission resources are air interface resources, the first network element is an access network device, and the second network element is a terminal device.

[0042] Optionally, the transmission resource is a routing resource, the first network element is a core network device, and the second network element is an access network device.

[0043] In a second aspect, a resource scheduling apparatus is provided, the apparatus being used to perform the methods of the first aspect and any possible implementation thereof.

[0044] Thirdly, a resource scheduling apparatus is provided, the apparatus including a processor. The processor is coupled to a memory and can be used to execute a computer program in the memory to implement the methods in any of the possible implementations of the above aspects. Optionally, the apparatus further includes a memory. Optionally, the apparatus further includes a communication interface, to which the processor is coupled.

[0045] This resource scheduling device can correspond to the first network element in the first aspect above.

[0046] If the first network element is an access network device, in one implementation, the device is an access network device, and the communication interface can be a transceiver or an input / output interface. In another implementation, the device is a chip configured in the access network device, and the communication interface can be an input / output interface.

[0047] If the first network element is a core network device, in one implementation, the device is a core network device; in another implementation, the device is a chip configured within the core network device. The communication interface can be an input / output interface.

[0048] Optionally, the transceiver described above can be a transceiver circuit. Optionally, the input / output interface described above can be an input / output circuit.

[0049] Fourthly, a computer-readable storage medium is provided for storing a computer program, which, when executed by a computer or processor, is used to implement the methods of the first aspect and any possible implementation thereof.

[0050] Fifthly, a computer program product is provided, the computer program product including instructions that, when executed, cause a computer to perform the methods of the first aspect and any possible implementation thereof.

[0051] In a sixth aspect, a system-on-a-chip (SoC) or system-on-a-chip (SoC) is provided, which can be applied to an electronic device. The SoC or SoC includes: at least one communication interface, at least one processor, and at least one memory. The communication interface, memory, and processor are interconnected via a bus. The processor executes instructions stored in the memory, enabling the terminal device to perform the methods described in the first aspect and any possible implementation thereof.

[0052] In a seventh aspect, a communication system is provided. The communication system includes the aforementioned first network element and second network element.

[0053] It should be understood that the second to sixth aspects of the embodiments of this application correspond to the technical solutions of the first aspect of the embodiments of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be described again. Attached Figure Description

[0054] Figure 1 This is an example diagram of a network architecture provided in one embodiment of this application;

[0055] Figure 2 A schematic diagram of the transmission of an encoded stream under ideal conditions, provided for one embodiment of this application;

[0056] Figure 3 This is a schematic diagram illustrating the transmission of an encoded stream in the presence of network jitter, as provided in one embodiment of this application.

[0057] Figure 4 A flowchart illustrating a resource scheduling method provided in one embodiment of this application;

[0058] Figure 5 A schematic diagram of a first time period provided for one embodiment of this application;

[0059] Figure 6 This is a schematic diagram of a linear function simulating the actual arrival time of multiple coded frames to the first network element, provided as an embodiment of this application.

[0060] Figure 7 and Figure 8 A schematic diagram illustrating the process of determining the expected time of a first coded frame according to an embodiment of this application;

[0061] Figure 9 This is a schematic diagram illustrating the expected time of each encoded frame obtained by performing a resource scheduling method once for each encoded frame to determine the remaining delay budget, according to one embodiment of this application.

[0062] Figure 10 A schematic diagram illustrating the expected timing of multiple encoded frames following a first encoded frame determined based on frame rate f, provided in one embodiment of this application;

[0063] Figure 11 and Figure 12 A schematic block diagram of a resource scheduling device provided in an embodiment of this application. Detailed Implementation

[0064] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0065] The resource scheduling method provided in this application can be applied to various communication systems, such as: Long Term Evolution (LTE) systems, LTE frequency division duplex (FDD) systems, LTE time division duplex (TDD) systems, Universal Mobile Telecommunication System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX) systems, future 5th Generation (5G) mobile communication systems, or new radio access technology (NR). Among these, 5G mobile communication systems can include non-standalone (NSA) and / or standalone (SA) networks.

[0066] In this embodiment of the application, the access network device can be any device with wireless transceiver functionality. Access network equipment includes, but is not limited to: evolved Node B (eNB), radio network controller (RNC), Node B (NB), base station controller (BSC), base transceiver station (BTS), home base station (e.g., home evolved Node B, or home Node B, HNB), baseband unit (BBU), access point (AP), wireless relay node, wireless backhaul node, transmission point (TP), or transmission and reception point (TRP) in a wireless fidelity (WiFi) system. It can also be a gNB in ​​a 5G system, such as NR, or a transmission point (TRP or TP), one or a group of antenna panels (including multiple antenna panels) of a base station in a 5G system, or a network node constituting a gNB or transmission point, such as a baseband unit (BBU) or a distributed unit (DU).

[0067] Terminal equipment can also be called user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device.

[0068] Terminal devices can be devices that provide voice / data connectivity to users, such as handheld devices with wireless connectivity, in-vehicle devices, etc. Currently, examples of such terminals include: mobile phones, tablets, computers with wireless transceiver capabilities (such as laptops and PDAs), mobile internet devices (MIDs), virtual reality (VR) devices, augmented reality (AR) devices, cloud virtual reality (Cloud VR) devices, cloud augmented reality (Cloud AR) devices, cloud XR devices, cloud gaming devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, and personal digital assistants (PDAs). PDA (Power Assistant), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, terminal devices in 5G networks or terminal devices in future public land mobile networks (PLMNs), etc.

[0069] Core network equipment may include, but is not limited to, access and mobility management function (AMF), session management function (SMF), and user plane function (UPF).

[0070] The AMF (Active Location Function) is primarily used for mobility management and access management, such as user location updates, user registration with the network, and user handover. The AMF can also be used to implement other functions of the mobility management entity (MME) besides session management, such as lawful interception or access authorization (or authentication).

[0071] The Service Provider Function (SMF) is primarily used for session management, UE Internet Protocol (IP) address allocation and management, selection of manageable User Plane Functions (MPF) endpoints, policy control, or charging function interfaces, and downlink data notification. In this embodiment, the SMF is mainly responsible for session management in the mobile network, such as session establishment, modification, and release. Specific functions may include, for example, allocating IP addresses to terminal devices and selecting a User Plane Function (UPF) that provides packet forwarding capabilities.

[0072] UPF stands for Data Plane Gateway. It can be used for routing and forwarding, or for handling Quality of Service (QoS) of user plane data. User data can access the data network (DN) through this network element. In the embodiments of this application, it can be used to forward instructions received from the access network device to the cloud server, and can also be used to forward encoded streams received from the cloud server to the access network device.

[0073] In addition, for ease of understanding, a brief explanation of the terms used below will be provided.

[0074] 1. Video: Video is a series of images played in a specific time sequence. In short, video can include image sequences.

[0075] 2. Video Encoding and Video Decoding: Video encoding generally refers to processing the sequence of images that form a video. Video encoding is performed on the source side; for example, in this embodiment, video encoding is performed on a cloud server. Video encoding typically involves processing (e.g., compressing) the raw video images to reduce the amount of data required to represent the video images, thereby enabling more efficient storage and / or transmission. Video encoding can be simply referred to as encoding.

[0076] Video decoding is performed on the destination side; for example, in this embodiment, video decoding is performed on the terminal device. Video decoding typically involves performing inverse processing relative to video encoding to reconstruct video images.

[0077] 3. Coded Stream: Encoding a sequence of images yields a coded stream. In the field of video encoding and decoding,

[0078] 4. Encoded Frame: A frame is the basic element of a stream. In the field of video encoding and decoding, the terms "picture," "frame," or "image" can be used synonymously.

[0079] A encoded frame is the basic element of an encoded stream. For example, an encoded stream can include multiple encoded frames. Multiple encoded frames can be carried in Internet Protocol (IP) packets and can be transmitted in the form of IP packets.

[0080] 5. Frame Rate: For video, frame rate refers to the number of frames played per second. For example, a frame rate of 60 frames per second (fps) means that 60 frames are played per second.

[0081] To facilitate understanding of the embodiments of this application, the network architecture applicable to the methods provided in the embodiments of this application will be briefly described first. Figure 1 This is a schematic diagram of the network architecture applicable to the resource scheduling method provided in the embodiments of this application. For example... Figure 1 As shown, the network architecture 100 may include: application layer device 110, terminal device 120, access network device 130, core network device 140 and cloud server 150.

[0082] The application layer device 110 can be, for example, an application layer device such as a cloud-based VR device, cloud-based AR device, cloud-based XR device, or cloud gaming device, like a headset or controller. The application layer device 110 can collect user operations, such as controller operation or voice control, and generate action commands based on these operations. These action commands can be transmitted over the network to the cloud server 150. The cloud server 150 can provide services such as logical operations, rendering, and encoding. In this embodiment, the cloud server 150 can be used to process and render action commands, and can also be used for encoding to generate an encoded stream. The encoded stream can be transmitted over the network to the application layer device 110.

[0083] For example, the application layer device 110 can be connected to the terminal device 120, such as a mobile phone or tablet. The terminal device 120 can connect to the access network device 130 via an air interface, and then the core network device 140 accesses the Internet, ultimately communicating with the cloud server 150. In other words, the encoded stream can sequentially pass through the core network device 140, the access network device 130, and the terminal device 120 before finally reaching the application layer device 110. Optionally, the terminal device 120 can also be used to decode the encoded stream and present the decoded video to the user.

[0084] It should be understood that Figure 1 The architecture shown is merely an example and should not be construed as limiting this application. For example, in some other possible designs, the application layer device 110 and the terminal device 120 may be integrated, and this application does not limit this.

[0085] To achieve a better user experience, real-time multimedia services require low latency, high reliability, and high bandwidth. For example, XR requires the latency between human movement and screen refresh, i.e., the MTP (motion to photon latency), to be no more than 20ms to avoid causing dizziness.

[0086] based on Figure 1 The network architecture shown can be divided into several stages: client processing, network transmission, and cloud server processing. A latency budget can be allocated to each stage. If each stage can be kept within its allocated latency budget, low-latency ideal transmission—"encode one frame, transmit one frame, play one frame"—can be achieved, resulting in ideal and stable transmission.

[0087] Figure 2 A schematic diagram of the transmission of the encoded stream under ideal conditions is shown. Figure 2 The encoded stream shown achieves ideal, smooth transmission by "encoding one frame, transmitting one frame, and playing one frame." For example... Figure 2 As shown, each small square in the diagram represents an encoded frame, or in other words, each small square represents an IP data packet used to carry an encoded frame.

[0088] Figure 2 The diagram shows the timing sequence of the six encoded frames 0 to 5 as they are encoded, transmitted over a fixed network, transmitted over an air interface, and then decoded and played back.

[0089] The encoding can be processed on a cloud server. The resulting six encoded frames have a stable frame interval T1, which can be inversely proportional to the frame rate of the encoded stream. For example, for an encoded stream with a frame rate of 60fps, the frame interval T1 can be 16.67ms.

[0090] The encoded frames are transmitted to the access network device via a fixed network. In ideal fixed network transmission, without considering factors such as network jitter, the arrival time of these 6 encoded frames at the access network device also has a stable frame interval, T1. The interval between the time each frame is sent from the cloud server and the time it arrives at the access network device is the fixed network transmission delay T2. It can be seen that in ideal fixed network transmission, the fixed network transmission delay of these 6 encoded frames is also the same.

[0091] The encoded frames are transmitted from the access network device to the terminal device. In ideal air interface transmission, without considering factors such as network jitter, the arrival time of these 6 encoded frames at the access network device also has a stable frame interval, T1. The interval between the time each frame is sent from the access network device and the time it arrives at the terminal device is the air interface transmission delay T3. It can be seen that in ideal air interface transmission, the air interface transmission delay of these 6 encoded frames is also the same.

[0092] Subsequently, the terminal device decodes and plays the received encoded frames. Since the transmission delay of the six encoded frames is stable and the frame interval also remains stable, the six decoded frames can be smoothly decoded and played back on the terminal device.

[0093] In summary, in real-time multimedia services, the end-to-end round-trip time (RTT) can include: the time spent on action commands, the processing time on the cloud server, the fixed-line transmission time, the air interface transmission time, and the decoding time. For example, the time spent acquiring action commands and uploading them to the cloud can be denoted as T. act The processing time for cloud server rendering, encoding, etc., is denoted as T. clo The fixed-line transmission time is denoted as T. fn The air interface transmission time is denoted as T. nr The processing time for terminal decoding and playback is denoted as T. ter Therefore, RTT can be expressed by the formula: RTT = T act +T clo +T fn +T nr +T ter .

[0094] It should be understood that Figure 2 For ease of understanding only, a timing diagram is shown illustrating the ideal "encode one frame, transmit one frame, play one frame" process for six encoded frames. However, this should not constitute any limitation on this application. This application does not limit the number of encoded frames, bitrate, or other parameters contained in the encoded stream.

[0095] However, network jitter is common in actual transmission. When jitter is present, terminal devices may be unable to play video at a pre-set frame rate. Figure 3 This diagram illustrates the transmission of the encoded stream in the presence of network jitter. For example... Figure 3 As shown, due to potential time delay jitter in actual fixed-line network transmission, some encoded frames may arrive early, while others may arrive late. For ease of understanding, Figure 3 The dashed squares represent the arrival time of each frame under ideal fixed-line transmission, while solid squares represent the arrival time of each frame in actual fixed-line transmission. As can be seen, encoded frames numbered 2, 3, and 4 arrive late, even after the ideal playback time. This can lead to stuttering, lag, and frame skipping. This can result in poor user experience, such as misaligned or jerky visuals, and a degraded user interaction experience.

[0096] It should be understood that Figure 3 This example merely illustrates a scenario where network jitter during fixed-line transmission causes encoded frames to arrive late. In reality, network jitter can be prevalent during the transmission of encoded streams, leading to unpleasant experiences such as misaligned or skipped playback, and a decline in user interaction.

[0097] In view of this, this application provides a resource scheduling method that determines the remaining delay budget of a coded frame in the first network element based on the actual time the coded frame arrives at the first network element, and schedules resources for the coded frame based on the remaining delay budget. Therefore, the remaining delay budget can be updated in real time according to the actual transmission status of each coded frame, and transmission resources can be flexibly scheduled based on the remaining delay budget. This allows for flexible resource scheduling for coded frames that arrive early or late in the first network element, alleviating stuttering, delays, and frame skipping caused by network jitter during coded stream transmission, reducing unpleasant experiences such as image misalignment and jumps, and improving the user interaction experience.

[0098] The resource scheduling method provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0099] It should be understood that, for ease of understanding and explanation, the following description uses the first network element as the execution subject, but this should not constitute any limitation on this application. Any entity that can execute the method provided in this application can serve as the execution subject, as long as it can run the code or program that records the method provided in this application. For example, the first network element can also be replaced by a component configured in the first network element, such as a chip, a chip system, or other functional modules capable of calling and executing programs.

[0100] It should also be understood that the first network element can be any network element the encoded stream passes through during its transmission from the cloud server to the terminal device. The encoded stream can be transmitted from the first network element to the second network element. For example, in... Figure 1 In the network architecture shown, the first network element may be, for example, an access network device, and the second network element may be, for example, a terminal device; or, the first network element may be, for example, a core network device, and the second network element may be, for example, an access network device.

[0101] Figure 4 This is a flowchart illustrating a resource scheduling method provided in this application. Figure 4 As shown, the method 400 may include steps 410 to 440. The details are as follows. Figure 4 Each step in the process.

[0102] In step 410, the first network element receives the first encoded frame.

[0103] From the above text Figure 1 and Figure 2 As described above, cloud servers can generate encoded streams by rendering and encoding video streams. An encoded stream can include multiple encoded frames, each of which can be transmitted via the internet or carrier networks (including fixed-line and air interface transmission) to the terminal device, where it is then decoded and played.

[0104] As an example, the first network element can be an access network device, which can receive the first encoded frame from a core network device, such as a UPF, and can send the received first encoded frame to the terminal device through air interface resources.

[0105] In another example, the first network element can be a core network device, which can receive the first coded frame from the Internet and send the received first coded frame to the access network device through routing resources.

[0106] In step 420, the first network element determines the remaining delay budget of the first coded frame based on the actual time when the first coded frame arrives at the first network element.

[0107] Since the remaining delay budget of the first coded frame is determined based on the actual time it arrives at the first network element, the remaining delay budget can be determined based on the actual transmission status of the first coded frame. In other words, for a given coded stream, the remaining delay budget of the multiple coded frames contained therein may be dynamically changing; therefore, this remaining delay budget can also be called a dynamic delay budget.

[0108] In this embodiment of the application, the first network element can determine the latency budget of the first coded frame based on the following two possible implementation methods.

[0109] In a first possible implementation, the first network element can determine the remaining delay budget of the first encoded frame based on the time interval between the actual and expected arrival time of the first encoded frame, and the maximum delay budget allocated to the first network element. In other words, in this implementation, the first network element can determine the remaining delay budget of the first encoded frame without the cooperation of other network elements. In a second possible implementation, the first network element can determine the remaining delay budget based on the end-to-end RTT of the encoded stream, the instruction processing time of the application layer device, the processing time of the encoding end, the actual time the first encoded frame is sent from the encoding end, the actual time the first encoded frame arrives at the first network element, and the processing time of the decoding end. In other words, in this implementation, the first network element can determine the remaining time budget of the first encoded frame through the cooperation of other network elements.

[0110] The two possible implementation methods described above are explained in detail below.

[0111] In the first possible implementation, the expected time of the first encoded frame can specifically refer to: the latest time that the first network element is expected to send the first encoded frame; or, the latest time that the first encoded frame is expected to be transmitted to the second network element; or, the latest time that the first encoded frame is expected to be decoded by the second network element.

[0112] The maximum delay budget allocated to the first network element is determined based on the end-to-end RTT of the first coded frame. Depending on the definition of expected time, the maximum delay budget allocated to the first network element can also be defined differently. As mentioned earlier, RTT can be expressed by the formula: RTT = T act +T clo +T fn +T nr +T ter .

[0113] For example, taking the first network element as an access network device, the maximum latency budget allocated to the first network element can specifically refer to the maximum latency budget allocated to the access network device. Assume the time when the first network element receives the first coded frame is denoted as t. rec The time when the first encoded frame was sent is denoted as t. sen The air interface transmission time for the first encoded frame from the first network element to the second network element is t. tra Then the air interface transmission time T nr It should satisfy: T nr =t tra +t sen -t rec .

[0114] If the expected time of the first coded frame is the latest time that the first network element is expected to send the first coded frame, the maximum delay budget (e.g., T) is... maxdThe value of ) can be the time t when the first network element sends the first encoded frame. sen The time t when the first encoded frame is received rec The maximum value of the difference. Assume the time interval between the transmission of the first coded frame from the first network element to the second network element is t. tra This maximum value can be determined by the expected RTT (i.e., the predefined RTT; for ease of distinction from the actual RTT, the expected RTT is denoted as RTT). exp Subtract the time T spent on command acquisition and uploading to the cloud. act The processing time T of the cloud server clo Fixed-line transmission time T fn air interface transmission time t tra Processing time T of terminal devices ter We obtain, that is, T maxd This can be expressed by the following formula:

[0115] T maxd =RTT exp -T act -T clo -T fn -T ter -t tra .

[0116] If the expected time for the first coded frame is the latest time at which the first coded frame is expected to be transmitted to the second network element, the maximum delay budget T maxd The value can be the maximum difference between the time when the second network element receives the first coded frame and the time when the first network element receives the first coded frame, that is, the aforementioned T. nr The maximum value. This maximum value can be calculated by subtracting the time T spent on command acquisition and uploading to the cloud from the RTT. act The processing time T of the cloud server clo Fixed-line transmission time T fn Processing time T of terminal devices ter We obtain, that is, T maxd This can be expressed by the following formula:

[0117] T maxd =RTT exp -T act -T clo -T fn -T ter .

[0118] If the expected time of the first coded frame is the latest time at which the first coded frame is expected to be decoded in the second network element, the maximum delay budget T maxd The value can be obtained by subtracting the time T for command acquisition and uploading to the cloud from the RTT. act The processing time T of the cloud server cloTransmission time T between fixed network and fixed network fn We obtain, that is, T maxd This can be expressed by the following formula:

[0119] T maxd =RTT exp -T act -T clo -T fn .

[0120] It should be understood that the above example, using an access network device as the first network element, is for ease of understanding only, illustrating the maximum latency budget allocated to the first network element. Based on the same concept, the maximum latency budget allocated to the first network element when other devices are used as the first network element can also be determined. For the sake of brevity, examples are not provided here.

[0121] The following details the specific process for determining the remaining delay budget for the first coded frame.

[0122] For example, step 420 may specifically include:

[0123] Step 4201: The first network element learns the network jitter of the transmission link between the encoding end and the first network element based on the arrival times of multiple encoded frames arriving at the first network element within the first time period.

[0124] Step 4202: The first network element predicts the ideal time for the first coded frame to arrive at the first network element based on network jitter;

[0125] Step 4203: The first network element determines the candidate expected time of the first coded frame based on the actual arrival time of multiple coded frames to the first network element and the frame drop rate threshold.

[0126] Step 4204: The first network element determines the expected time of the first coded frame based on the ideal time of arrival of the first coded frame, the candidate expected time of the first coded frame, and the maximum delay budget allocated to the first network element; and

[0127] Step 4205: The first network element determines the remaining delay budget of the first coded frame based on the time interval between the actual time and the expected time when the first coded frame arrives at the first network element.

[0128] Specifically, the expected time of the first coded frame can be determined based on the ideal time of the first coded frame arriving at the first network element. The ideal time of the first coded frame arriving at the first network element can be determined by the actual arrival times of multiple coded frames arriving at the first network element within a certain period before the first coded frame arrives at the first network element.

[0129] For ease of explanation, a period of time before the first coded frame arrives at the first network element is referred to as the first time period. Specifically, the first time period can refer to a period of time obtained by counting backwards from the actual time the first coded frame arrives at the first network element. In other words, the end time of the first time period can be the actual time the first coded frame arrives at the first network element. The duration of the first time period can be a predefined value; for example, the first time period can be 1 second (s) or 5 seconds. This application does not limit the specific duration of the first time period.

[0130] For ease of understanding, Figure 5 An example from the first time period is shown. Figure 5 The diagram shows n coded frames numbered 0, 1, 2, 3 up to n-1. Assuming frame n is the first coded frame, the time preceding the actual arrival time of frame n at the first network element can be called the first time period. During this first time period, multiple coded frames arrive at the first network element, such as... Figure 5 As shown, the n encoded frames numbered 0, 1, 2, 3 up to n-1 all arrive at the first network element within the first time period.

[0131] It should be understood that Figure 5 For ease of understanding only, a first time period and multiple encoded frames within the first time period are illustrated exemplarily. This application does not limit the specific duration of the first time period or the number of encoded frames arriving at the first network element within the first time period. Furthermore, the first time period may also include the first encoded frame, i.e. Figure 5 The actual time when the encoded frame numbered n arrives at the first network element. In this case, the multiple encoded frames arriving at the first network element within this first time period may also include the first encoded frame.

[0132] It should also be understood that the multiple coded frames within the first time period (excluding the first coded frame) and the first coded frame may belong to the same coded stream or to different related coded streams (such as the primary stream and the enhancement stream), and this application embodiment does not limit this.

[0133] In step 4201, the first network element can learn the network jitter of the transmission link between the encoding end and the first network element based on the actual arrival times of multiple encoded frames arriving at the first network element within the first time period.

[0134] Specifically, the first network element can predict the ideal time for the first coded frame to arrive at the first network element based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period, using linear regression or Kalman filtering.

[0135] Taking linear regression as an example, the first network element can simulate a linear function based on the actual arrival times of multiple coded frames within the first time period. This linear function can be used to represent the linear relationship between the coded frame and its arrival time. Based on this linear function, the arrival time of the next coded frame (such as the first coded frame mentioned above) can be predicted.

[0136] The process of simulating the linear function based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period is essentially the process of learning the network jitter of the transmission link from the encoding end to the first network element during the fixed network transmission process. Based on the learned network jitter, step 4202 can be executed to predict the ideal time for the first coded frame to arrive at the first network element. It should be understood that the ideal time for the first coded frame to arrive at the first network element is the time predicted under the condition of learning the actual network jitter. For ease of distinction, the time for the first coded frame to arrive at the first network element predicted based on the learned network jitter is called the ideal time for the first coded frame to arrive at the first network element.

[0137] Figure 6 The curve of a simulated linear function based on the actual arrival times of multiple coded frames in the first network element is shown. For example... Figure 6 As shown, Figure 6 The horizontal axis represents the number of encoded frames arriving at the first network element, and the unit can be frames; the vertical axis represents the actual time of arrival of the encoded frame at the first network element, and the unit can be milliseconds (ms). Each discrete point can represent the actual time of arrival of an encoded frame at the first network element. For example, the arrival time of the first encoded frame arriving at the first network element can be used as a reference, corresponding to the origin of the coordinate system. From there, the arrival times of the second, third, and so on, up to the last encoded frame among these multiple encoded frames, can be recorded.

[0138] The linear function curve shown in the figure can be simulated using linear regression. It can be seen that the points corresponding to multiple coded frames are distributed around this curve.

[0139] Based on this curve, the first network element can predict the ideal time for the first coded frame to arrive at the first network element.

[0140] It should be understood that the above text, in combination with... Figure 6 The process of predicting the ideal time for the first coded frame to arrive at the first network element is described using linear regression as an example, but this should not be construed as limiting this application. For example, the ideal time for the first coded frame to arrive at the first network element can also be predicted based on algorithms such as Kalman filtering. For the sake of brevity, this will not be elaborated here.

[0141] It should also be understood that any coded frame arriving at the first network element before the first coded frame can learn network jitter based on the method described above, and predict its ideal arrival time at the first network element based on the network jitter. For the sake of brevity, this will not be repeated here.

[0142] In steps 4203 to 4204, the first network element determines the expected time of the first coded frame.

[0143] On the one hand, determining the expected time ensures that the transmission of the encoded frame meets predefined thresholds, such as the frame drop rate threshold mentioned above, so that the encoded stream can be played normally; on the other hand, the time interval between the expected time and the ideal time for the first encoded frame to arrive at the first network element should not be lower than the maximum delay budget allocated to the first network element.

[0144] To ensure proper playback of the encoded stream, a frame drop rate threshold can be predefined, for example, 1% or 0.1%. It can be understood that the lower the frame drop rate, the higher the proportion of successfully transmitted encoded frames in the stream, thus guaranteeing proper playback. Alternatively, a lower frame drop rate threshold also means a higher requirement for the proportion of successfully transmitted encoded frames, potentially resulting in a larger time interval between the expected arrival time and the ideal arrival time of the first encoded frame at the first network element.

[0145] It should be understood that the actual time when the aforementioned multiple encoded frames arrive at the first network element is also the actual arrival time of the multiple encoded frames that arrive at the first network element within the first time period mentioned above.

[0146] The following text combines Figure 7 and Figure 8 The process of determining the expected time for the first coded frame is described in detail.

[0147] First, based on the delay distribution of the actual arrival times of multiple coded frames arriving at the first network element within the first time period, and a predefined frame drop rate threshold, a candidate expected time for the first coded frame is determined. It is called a candidate expected time because the time interval between the expected time of the first coded frame determined based on the aforementioned delay distribution and frame drop rate threshold and the ideal time is not necessarily lower than the maximum delay budget allocated to the first network element. To distinguish it from the finally determined expected time of the first coded frame, the expected time of the first coded frame determined based on the delay distribution of the actual arrival times of multiple coded frames arriving at the first network element within the first time period, and the predefined frame drop rate threshold, is referred to as the candidate expected time.

[0148] The time delay distribution of the actual arrival times of multiple coded frames arriving at the first network element within the first time period can be used to indicate the offset between the ideal time and the actual time of each coded frame arriving at the first network element within the first time period, as well as the number of coded frames corresponding to different offsets.

[0149] The ideal time for each coded frame to arrive at the first network element can be determined based on the actual arrival times of multiple coded frames within a certain time period prior to its arrival. The specific process can be found in the above text. Figure 6 The specific process for determining the ideal time for the first coded frame to arrive at the first network element is described. For the sake of brevity, it will not be elaborated here.

[0150] Figure 7 This shows the offset between the ideal and actual arrival times of multiple coded frames arriving at the first network element within the first time period. For example... Figure 7 As shown in the figure, the ideal time for n coded frames to arrive at the first network element within the first time period is indicated by the dashed box, and the actual time for n coded frames to arrive at the first network element within the first time period is indicated by the solid box. By taking the offset Δt of a certain boundary (such as the right boundary) of two boxes with the same number, the offset between the ideal time and the actual time for each coded frame to arrive at the first network element can be obtained, as shown in t0, t1, t2, t3 up to t in the figure. n-1 As shown, these represent the offsets between the ideal and actual times of arrival of n encoded frames numbered 0, 1, 2, 3 up to n-1 to the first network element.

[0151] By counting the number of coded frames at each offset, the delay distribution of the actual arrival times of multiple coded frames arriving at the first network element within the first time period can be obtained, as shown below. Figure 8 As shown. Figure 8 The horizontal axis represents the offset Δt, and the vertical axis represents the number of coded frames corresponding to different offset values ​​Δt.

[0152] After determining the delay distribution of the actual arrival times of multiple coded frames arriving at the first network element within the first time period, the candidate expected time of the first coded frame can be determined according to the predefined frame drop rate threshold.

[0153] exist Figure 8 In the delay distribution diagram shown, with an offset of 0 as a reference, an offset (e.g., denoted as τ) is found such that the proportion of encoded frames with an offset greater than τ is less than a preset frame drop rate threshold. In other words, the proportion of encoded frames located to the right of this offset τ in the total number of encoded frames in the delay distribution diagram is less than the preset frame drop rate threshold. If the frame drop rate threshold is denoted as φ, where 0 < φ < 1, then the proportion of frames corresponding to offsets τ and 0 in the aforementioned delay distribution is less than 1 - φ. In other words, the proportion of frames corresponding to the offset of the candidate expected time of the first encoded frame arriving at the first network element relative to the ideal time of the first encoded frame arriving at the first network element in the delay distribution is less than 1 - φ.

[0154] Therefore, by adding the offset τ to the ideal time of the first coded frame arriving at the first network element, the candidate expected time of the first coded frame can be obtained. It can be understood that the time interval between the candidate expected time of the first coded frame and the actual time of its arrival at the first network element is τ.

[0155] If the time interval τ is less than or equal to the maximum delay budget allocated to the first network element, the aforementioned candidate expected time can be determined as the expected time of the first coded frame. If the time interval τ is greater than the maximum delay budget allocated to the first network element, the expected time of the first coded frame can be determined based on the maximum delay budget. For example, the sum of the ideal time for the first coded frame to arrive at the first network element and the maximum delay budget allocated to the first network element can be determined as the expected time of the first coded frame; or, the time interval between the expected time of the first coded frame and the ideal time for the first coded frame to arrive at the first network element is the maximum delay budget.

[0156] For ease of understanding, let's assume the maximum delay budget allocated to the first network element is denoted as T. maxd The ideal time for the first coded frame to arrive at the first network element is denoted as T. iarr Then the expected time T of the first coded frame exp It can be represented as: T exp =T iarr +min(τ,T maxd It should be understood that the above text, in conjunction with... Figure 7 and Figure 8 The descriptions provided are for illustrative purposes only and should not be construed as limiting the scope of this application. This application does not limit the offset corresponding to each coded frame, the number of coded frames corresponding to different offset values, or the relationship between the offset τ and the maximum delay budget allocated to the first network element.

[0157] In step 4205, the first network element determines the remaining delay budget based on the actual time and expected time of the arrival of the first coded frame at the first network element.

[0158] The remaining delay budget can be obtained by calculating the time interval between the actual time and the expected time of the arrival of the first coded frame at the first network element. Let T be the actual time of the arrival of the first coded frame at the first network element. farr The first network element can obtain the remaining delay budget of the first coded frame, denoted as T. db Then the remaining delay budget T of the first coded frame db T can be calculated using the following formula: db =T exp -T farr .

[0159] It should be noted that the first coded frame in the embodiments of this application can be one frame from multiple coded frames in the same coded stream, or it can be one frame from multiple coded frames in related different coded streams. The related different coded streams can include a primary stream and an enhancement stream. The primary stream and the enhancement stream can be transmitted in parallel, or they can be transmitted sequentially according to time sequence. This application does not limit this.

[0160] The first network element can perform step 420 of the above method to determine the remaining delay budget for multiple consecutive coded frames. In this case, each of the multiple consecutive coded frames can be referred to as the first coded frame. In other words, the arrival time of the first coded frame at the first network element is a process of real-time updating based on the actual transmission situation, and the expected time of the first coded frame is also a process of real-time updating based on the actual transmission situation. The remaining delay budget of the first coded frame determined in this way is more consistent with the actual network conditions of transmission. Figure 9 The figure illustrates the expected time for each coded frame obtained by performing steps 4201 to 4205 of the above steps individually for each coded frame to determine the remaining delay budget. The figure shows the expected times of multiple consecutive coded frames numbered n, n+1, n+2, and n+3, and the remaining delay budget determined by the expected time and the actual time of arrival at the first network element, corresponding to T in the figure. db,n T db,n+1 T db,n+2 and T db,n+3 As can be seen, the expected time interval between any two adjacent coded frames is different.

[0161] Alternatively, the first network element can designate one of the coded frames as the first coded frame, and the expected time of the coded frames following the first coded frame can be calculated according to the frame rate f. The interval between the expected times of two adjacent coded frames is 1 / f. For example, the expected time interval between the expected time of the next frame after the first coded frame (e.g., called the second coded frame) and the expected time of the first coded frame is 1 / f, the expected time interval between the expected time of the next frame after the second coded frame and the expected time of the second coded frame is also 1 / f, and so on, not listed here. Figure 10 The expected times for multiple coded frames following the first coded frame, determined based on the frame rate f, are shown. It can be seen that the interval between the expected times of any two adjacent coded frames is 1 / f.

[0162] Alternatively, the first network element can use N consecutive coded frames as a period, and take the first coded frame in each period as the first coded frame to determine the remaining delay budget. The expected time of other coded frames in the same period can be calculated according to the frame rate.

[0163] In the second possible implementation, the remaining latency budget can be determined by the end-to-end RTT of the first coded frame, the instruction processing time of the application layer device, the processing time of the encoding end, the actual time when the first coded frame is sent from the encoding end, the actual time when the first coded frame arrives at the first network element, and the processing time of the decoding end. The end-to-end RTT of the first coded frame can be the end-to-end RTT of the coded stream to which the first coded frame belongs.

[0164] The actual time when the first encoded frame was sent from the encoding end is carried in the data packet used to carry the first encoded frame. For example, each time the first encoded frame passes through a network element, that network element can carry the time when it sent the encoded frame in the data packet, for example, in the form of a timestamp in the packet header. In this way, the first network element can obtain the actual time when the first encoded frame was sent from the encoding end when it receives the first encoded frame, and thus determine the remaining delay budget.

[0165] From the relevant calculation formulas for end-to-end RTT mentioned above, we know that the remaining delay budget T of the first network element can be calculated using the following formula. db :T db =RTT-T act -T clo -T ter -(T send -T rec ), where T send T represents the time when the first encoded frame was sent from the cloud server. rec This indicates the time when the first coded frame arrives at the first network element, or in other words, the time when the first network element receives the first coded frame. It should be understood that the receiving or transmitting times involved in this implementation are all actual times.

[0166] In step 430, the first network element determines the scheduling priority based on the remaining delay budget.

[0167] The scheduling priority refers to the priority used to schedule the transmission resources for the first coded frame. When the first network element is an access network device, the transmission resources can be air interface resources. Air interface resources can specifically include time-frequency resources, spatial resources, etc.

[0168] When the first network element is a core network device, the transmission resource can be a routing resource. Routing resources can specifically include routers, forwarding paths, etc.

[0169] The relationship between the remaining delay budget and the scheduling priority of transmission resources can be summarized as follows: the smaller the remaining delay budget, the higher the scheduling priority of the transmission resource. For example, if scheduling priority is represented by a priority value, and a smaller priority value corresponds to a higher scheduling priority, then the priority value is inversely proportional to the remaining delay budget. For instance, the relationship between the priority value L and the remaining delay budget D can be expressed as: L = α / D, where α is a coefficient, which can be a pre-configured or protocol-defined value, and α is not zero.

[0170] In step 440, the first network element schedules transmission resources for the first encoded frame based on scheduling priority.

[0171] The first network element can prioritize scheduling transmission resources for the first encoded frame when the scheduling priority is high, so that the first encoded frame can be sent in a shorter time; when the scheduling priority is low, the latency of the first encoded frame can be increased, and the resources can be used for more urgent services.

[0172] For example, scheduling priority can be determined based on preset conditions. For instance, the preset condition could be that the priority value is less than or equal to a preset priority threshold. Using the priority value example mentioned above, the smaller the priority value, the higher the scheduling priority. If the priority value is less than or equal to the preset priority threshold, it indicates a higher scheduling priority; if the priority value is greater than the preset priority threshold, it indicates a lower scheduling priority.

[0173] As previously mentioned, the first network element can be an access network device, and the transmission resources can be air interface resources. Under the condition of high scheduling priority, the first network element, as an access network device, can prioritize scheduling PRBs for the first coded frame, so that the number of PRBs scheduled for the first coded frame is greater than or equal to the number of PRBs required for the transmission of the first coded frame.

[0174] The first network element can also be a core network device, and the transmission resources can be routing resources. Under high scheduling priority, the first network element, acting as a core network device, can select a forwarding path with fewer hops for the first coded frame, and / or select a router with a shorter queue length to forward the first coded frame. Here, the router queue length can be understood as the number of data packets waiting to be transmitted by that router. A longer queue length indicates a larger number of data packets waiting to be transmitted, thus requiring a longer waiting time; a shorter queue length indicates a smaller number of data packets waiting to be transmitted, thus requiring a relatively shorter waiting time.

[0175] Conversely, if the first network element is an access network device, under higher scheduling priority, the access network device, as the first network element, can increase the latency of the first coded frame within the allowable range of the remaining latency budget, and allocate more resources to the transmission of other urgent services. This frees up time domain resources, accommodates the transmission needs of other services, allows for flexible allocation of time domain resources, and improves resource utilization.

[0176] If the first network element is a core network device, under low scheduling priority, the core network device, as the first network element, can, within the remaining delay budget, select a forwarding path with a relatively long transmission hop count and / or a router with a large queue length to forward the first coded frame. This allows forwarding paths with fewer transmission hops and / or routers with smaller queue lengths to be used for other urgent services. This frees up routing resources, accommodates the transmission needs of other services, enables flexible allocation of routing resources, and improves resource utilization.

[0177] It should be understood that the enumeration of preset conditions and priority thresholds herein is merely illustrative and should not constitute any limitation on this application. For example, scheduling priorities can be divided into more different levels based on priority values, and transmission resources can be scheduled for the first coded frame based on different levels.

[0178] It should also be understood that the above examples of scheduling transmission resources for the first coded frame based on scheduling priority are merely illustrative for ease of understanding and should not constitute any limitation on this application. For example, the first network element can also schedule transmission resources based on a proportional fairness scheduling (PFS) algorithm, such that the scheduling priority is directly proportional to the ratio of the current instantaneous rate to the historical average rate and inversely proportional to the remaining delay budget. The first network element can schedule transmission resources for the first coded frame based on existing scheduling priority algorithms; for the sake of simplicity, these will not be illustrated in detail here.

[0179] In another implementation, the first network element can also directly schedule transmission resources for the first coded frame based on the remaining delay budget, without needing to calculate the scheduling priority of the transmission resources. Based on the same concept described above, the first network element can prioritize scheduling transmission resources for the first coded frame when the remaining delay budget is greater than or equal to a certain preset threshold; and relax the delay of the first coded frame when the remaining delay budget is less than a certain preset threshold. This allows for flexible scheduling of transmission resources for the first coded frame based on a dynamic remaining delay budget.

[0180] Based on the above scheme, the first network element determines the remaining delay budget for the first coded frame based on its actual arrival time, and can schedule transmission resources for the first coded frame according to the dynamically calculated remaining delay budget. Therefore, in the event of network jitter, the remaining delay budget can be updated in real time based on the actual transmission status of each coded frame, and resources can be flexibly scheduled accordingly. For delayed coded frames, this helps alleviate stuttering, lag, and frame skipping caused by network jitter, reducing poor user experience such as image misalignment and jumps, and improving user interaction. For early coded frames, the delay can be relaxed to accommodate the transmission needs of other services, prioritizing resources for more urgent services. Overall, the first network element can flexibly configure resources, ensuring that resource allocation perfectly matches transmission requirements, thereby improving resource utilization.

[0181] On the other hand, based on the above solution, the first network element does not need to introduce additional buffering to achieve smooth playback, thus avoiding an increase in the latency of the entire encoded stream. In other words, it does not lead to a significant increase in end-to-end RTT, thereby avoiding issues such as black bars and screen tearing during viewpoint transitions and improving the user experience.

[0182] It should be noted that the access network equipment and core network equipment listed above as the two examples of first network elements can implement the above scheme individually or simultaneously. When the access network equipment and core network equipment implement the above scheme simultaneously, a two-stage dejitter effect can be achieved. This can significantly reduce stuttering, lag, and frame skipping caused by network jitter, thereby greatly reducing unpleasant user experiences such as image misalignment and jumps.

[0183] The methods provided by the embodiments of this application have been described in detail above with reference to several accompanying drawings. However, it should be understood that these drawings and their corresponding descriptions are merely illustrative for ease of understanding and should not constitute any limitation on this application. Not every step in each flowchart is necessarily required to be performed; for example, some steps can be skipped. Furthermore, the execution order of each step is not fixed and is not limited to what is shown in the figures. The execution order of each step should be determined by its function and internal logic.

[0184] To achieve the functions of the methods provided in the embodiments of this application, the first network element may include a hardware structure and / or a software module, implementing the above functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether a particular function is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.

[0185] The following will combine Figures 11 to 12The apparatus provided in the embodiments of this application will be described in detail.

[0186] Figure 11 This is a schematic block diagram of a resource scheduling device 1100 provided in an embodiment of this application. It should be understood that the resource scheduling device 1100 may correspond to the first network element in the above method embodiment, and may be used to execute the various steps and / or processes executed by the first network element in the above method embodiment.

[0187] like Figure 11 As shown, the resource scheduling device 1100 may include a transceiver module 1110 and a processing module 1120. Specifically, when the device 1100 is used to perform... Figure 4 When the first network element executes the method, the transceiver module 1110 can be used to execute step 410 in the method 400 above. The processing module 1120 can be used to execute some or all of the steps in steps 420, 430 and 440 in the method 400.

[0188] It should be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0189] Figure 12 This is another schematic block diagram of the resource scheduling device provided in the embodiments of this application. For example... Figure 12 As shown, Figure 12 As shown, the resource scheduling device 1200 includes at least one processor 1210, which is used to implement the function of the first network element in the method provided in the embodiments of this application.

[0190] For example, if the resource scheduling device 1200 corresponds to the first network element in the above method embodiment, the processor 1210 can be used to determine the remaining delay budget of the first coded frame; and to schedule transmission resources for the first coded frame according to the remaining delay budget of the first coded frame. See the detailed description in the method example for details, which will not be repeated here.

[0191] The resource scheduling device 1200 may further include at least one memory 1220 for storing program instructions and / or data. The memory 1220 is coupled to the processor 1210. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, and may be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. The processor 1210 may operate in conjunction with the memory 1220. The processor 1210 may execute program instructions stored in the memory 1220. At least one of the at least one memory may be included in the processor.

[0192] The resource scheduling device 1200 may further include a communication interface 1230. This communication interface 1230 may be a transceiver, interface, bus, circuit, or device capable of transmitting and receiving functions. The communication interface 1230 is used to communicate with other devices via a transmission medium, thereby enabling communication between the devices in the resource scheduling device 1200 and other devices. For example, if the resource scheduling device 1200 corresponds to the first network element in the above method embodiment, the other devices may be the second network element or the upstream network element of the first network element. The processor 1210 uses the communication interface 1230 to transmit and receive data and is used to implement… Figure 4 The method executed by the first network element in the corresponding embodiment.

[0193] This application embodiment does not limit the specific connection medium between the processor 1210, memory 1220, and communication interface 1230. This application embodiment... Figure 12 The memory 1220, processor 1210, and communication interface 830 are connected via a bus 1240. Figure 12 The connections between other components are shown in bold and are for illustrative purposes only, not as limiting information. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0194] This application also provides a processing apparatus, including at least one processor, which is configured to execute a computer program stored in a memory, such that the processing apparatus performs the method executed by the access network device or the core network device in the above method embodiments.

[0195] In the embodiments of this application, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0196] In the embodiments of this application, the memory can be non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as random-access memory (RAM). Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited thereto. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing storage functions, used to store program instructions and / or data.

[0197] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute... Figure 4 The method executed by the first network element in the illustrated embodiment.

[0198] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when executed on a computer, causes the computer to perform... Figure 4 The method executed by the first network element in the illustrated embodiment.

[0199] According to the method provided in the embodiments of this application, this application also provides a system, including the aforementioned cloud server, core network equipment, access network equipment, and terminal equipment. The core network equipment and / or access network equipment can be used to execute the method executed by the first network element in the above method embodiments.

[0200] The technical solutions provided in this application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a terminal device, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media, etc.

[0201] In the embodiments of this application, provided there is no logical contradiction, the embodiments may reference each other. For example, the methods and / or terms between method embodiments may reference each other, the functions and / or terms between device embodiments may reference each other, and the functions and / or terms between device embodiments and method embodiments may reference each other.

[0202] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A resource scheduling method, characterized in that, include: The first network element receives the first encoded frame; The first network element determines the remaining delay budget of the first encoded frame based on the actual time the first encoded frame arrives at the first network element; the remaining delay budget is the time interval between the actual time the first encoded frame arrives at the first network element and the expected time; the expected time is the latest time that the first network element is expected to send the first encoded frame, or the latest time that the first encoded frame is expected to be transmitted to the second network element, or the latest time that the first encoded frame is expected to be decoded by the second network element. The first network element schedules transmission resources for the first coded frame based on the remaining delay budget; The method further includes: The first network element determines the expected time based on the actual arrival times of multiple coded frames arriving at the first network element within a first time period, the ideal time for the first coded frame to arrive at the first network element, the maximum latency budget allocated to the first network element, and a predefined frame drop rate threshold; wherein, the end time of the first time period is the actual time for the first coded frame to arrive at the first network element, and the duration of the first time period is a predefined value; the ideal time for the first coded frame to arrive at the first network element is obtained based on learning the network jitter of the transmission link between the encoding end of the first coded frame and the first network element; the maximum latency budget is determined based on the end-to-end round-trip time (RTT) of the first coded frame.

2. The method as described in claim 1, characterized in that, The step of scheduling transmission resources for the first coded frame based on the remaining delay budget includes: The first network element determines the scheduling priority based on the remaining delay budget; The first network element schedules the transmission resources for the first coded frame based on the scheduling priority; The smaller the remaining delay budget, the higher the scheduling priority of the transmission resources.

3. The method as described in claim 1, characterized in that, The first network element determines the expected time based on the actual arrival times of multiple coded frames arriving at the first network element within the first time period, the ideal arrival time of the first coded frame at the first network element, the maximum latency budget allocated to the first network element, and a predefined frame drop rate threshold, including: The first network element determines the candidate expected time for the first coded frame to arrive at the first network element based on the actual arrival time of multiple coded frames arriving at the first network element within the first time period and the frame drop rate threshold. The first network element determines the expected time based on the candidate expected time, the ideal time for the first coded frame to arrive at the first network element, and the maximum delay budget.

4. The method as described in claim 3, characterized in that, The first network element determines a candidate expected time for the arrival of the first coded frame to the first network element based on the actual arrival time of multiple coded frames arriving at the first network element within the first time period and the frame drop rate threshold, including: The first network element learns the network jitter of the transmission link between the encoding end and the first network element based on the actual arrival times of multiple encoded frames arriving at the first network element within the first time period; The first network element predicts the ideal time for the first coded frame to arrive at the first network element based on the network jitter; The first network element determines the delay distribution, which is used to indicate the offset of the actual time of each of the plurality of coded frames arriving at the first network element relative to the ideal time, and the number of frames corresponding to different offsets; The first network element determines the candidate expected time for the first coded frame to arrive at the first network element based on the delay distribution and the frame drop rate threshold. The offset of the candidate expected time from the ideal time for the first coded frame to arrive at the first network element is less than 1-φ in the proportion of the number of frames corresponding to the delay distribution, where φ is the frame drop rate threshold, 0 < φ < 1.

5. The method as described in claim 3 or 4, characterized in that, The first network element determines the expected time based on the candidate expected time, the ideal time for the first coded frame to arrive at the first network element, and the maximum delay budget, including: If the time interval between the ideal time of the first coded frame arriving at the first network element and the candidate expected time is greater than or equal to the maximum delay budget, the first network element determines the expected time of the first coded frame based on the maximum delay budget, such that the time interval between the expected time and the ideal time of the first coded frame arriving at the first network element is the maximum delay budget; or If the time interval between the ideal time of the arrival of the first coded frame at the first network element and the candidate expected time is less than the maximum delay budget, the first network element determines the candidate expected time as the expected time.

6. The method according to any one of claims 1-4, characterized in that, The network jitter of the transmission link between the encoding end and the first network element is learned by linear regression or Kalman filtering based on the actual arrival times of multiple encoded frames that arrive at the first network element within the first time period.

7. The method according to any one of claims 1-4, characterized in that, The method further includes: The first network element determines the expected time of the second encoded frame based on the expected time of the first encoded frame and the frame rate of the encoded stream. The second encoded frame is the next frame after the first encoded frame.

8. The method as described in claim 1 or 2, characterized in that, The remaining latency budget is determined by the end-to-end RTT of the first encoded frame, the instruction processing time of the application layer device, the processing time of the encoding end, the actual time when the first encoded frame is sent from the encoding end, the actual time when the first encoded frame arrives at the first network element, and the processing time of the decoding end.

9. The method as described in claim 8, characterized in that, The actual time when the first encoded frame is sent from the encoding end is carried in the data packet used to carry the first encoded frame.

10. The method as described in claim 2, characterized in that, The step of scheduling the transmission resources for the first encoded frame based on the scheduling priority includes: When the scheduling priority is less than or equal to a preset priority threshold, the first network element prioritizes scheduling transmission resources for the first coded frame, and the number of Physical Resource Blocks (PRBs) in the scheduled transmission resources is greater than or equal to the number of PRBs required for the transmission of the first coded frame.

11. The method as described in claim 2, characterized in that, The step of scheduling the transmission resources for the first encoded frame based on the scheduling priority includes: When the scheduling priority is greater than a preset priority threshold, the first network element increases the latency of the first encoded frame.

12. The method according to any one of claims 1-4 and 9, characterized in that, The transmission resources are air interface resources, the first network element is an access network device, and the second network element is a terminal device.

13. The method according to any one of claims 1-4 and 9, characterized in that, The transmission resources are routing resources, the first network element is a core network device, and the second network element is an access network device.

14. A resource scheduling device, characterized in that, include: The transceiver module is used to receive the first encoded frame; The processing module is configured to determine the remaining delay budget of the first encoded frame based on the actual time the first encoded frame arrives at the device; the remaining delay budget is the time interval between the actual time the first encoded frame arrives at the device and the expected time, the expected time being the latest time the device is expected to send the first encoded frame, or the latest time the first encoded frame is expected to be transmitted to the second network element, or the latest time the first encoded frame is expected to be decoded by the second network element. Used to schedule transmission resources for the first coded frame based on the remaining delay budget; The expected time is determined based on the actual arrival times of multiple encoded frames arriving at the device within a first time period, the ideal time for the first encoded frame to arrive at the device, the maximum latency budget allocated to the device, and a predefined frame drop rate threshold; wherein, the end time of the first time period is the actual time for the first encoded frame to arrive at the device, and the duration of the first time period is a predefined value; the ideal time for the first encoded frame to arrive at the device is obtained by learning the network jitter of the transmission link between the encoding end of the first encoded frame and the device; the maximum latency budget is determined based on the end-to-end round-trip time (RTT) of the first encoded frame.

15. The apparatus according to claim 14, characterized in that, The processing module is specifically used for: Based on the remaining delay budget, determine the scheduling priority; Based on the scheduling priority, schedule the transmission resources for the first encoded frame; The smaller the remaining delay budget, the higher the scheduling priority of the transmission resources.

16. The apparatus according to claim 15, characterized in that, The processing module is specifically used for: Based on the actual arrival times of multiple encoded frames arriving at the device within the first time period and the frame drop rate threshold, a candidate expected time for the first encoded frame to arrive at the device is determined. The device determines the expected time based on the candidate expected time, the ideal time for the first coded frame to arrive at the device, and the maximum delay budget.

17. The apparatus according to claim 16, characterized in that, The processing module is specifically used for: Based on the actual arrival times of multiple encoded frames arriving at the device within the first time period, the network jitter of the transmission link between the encoding end and the device is learned; Based on the network jitter, predict the ideal time for the first encoded frame to arrive at the device; Determine the time delay distribution, which is used to indicate the offset of the actual time of arrival of each of the plurality of coded frames to the device relative to the ideal time, and the number of frames corresponding to different offsets; Based on the delay distribution and the frame drop rate threshold, a candidate expected time for the first coded frame to arrive at the device is determined. The offset of the candidate expected time from the ideal time for the first coded frame to arrive at the device is less than 1-φ, where φ is the frame drop rate threshold, and 0 < φ < 1.

18. The apparatus according to claim 16 or 17, characterized in that, The processing module is also used for: If the time interval between the ideal time of the first coded frame arriving at the device and the candidate expected time is greater than or equal to the maximum delay budget, the expected time of the first coded frame is determined based on the maximum delay budget, such that the time interval between the expected time of the first coded frame and the ideal time of the first coded frame arriving at the device is the maximum delay budget. or If the time interval between the ideal time of the arrival of the first coded frame at the device and the candidate expected time is less than the maximum delay budget, the candidate expected time is determined as the expected time.

19. The apparatus as claimed in any one of claims 14-17, characterized in that, The network jitter of the transmission link between the encoding end and the device is learned by linear regression or Kalman filtering based on the actual arrival times of multiple encoded frames that arrive at the device within the first time period.

20. The apparatus according to any one of claims 14-17, characterized in that, The processing module is also used for: The expected time of the second encoded frame is determined based on the expected time of the first encoded frame and the frame rate of the encoded stream. The second encoded frame is the next frame after the first encoded frame.

21. The apparatus according to claim 14 or 15, characterized in that, The remaining latency budget is determined by the end-to-end RTT of the first encoded frame, the instruction processing time of the application layer device, the processing time of the encoding end, the actual time when the first encoded frame is sent from the encoding end, the actual time when the first encoded frame arrives at the device, and the processing time of the decoding end.

22. The apparatus according to claim 21, characterized in that, The actual time when the first encoded frame is sent from the encoding end is carried in the data packet used to carry the first encoded frame.

23. The apparatus according to claim 15, characterized in that, The processing module is also used for: When the scheduling priority is less than or equal to a preset priority threshold, transmission resources are scheduled for the first coded frame first, and the number of Physical Resource Blocks (PRBs) in the scheduled transmission resources is greater than or equal to the number of PRBs required for the transmission of the first coded frame.

24. The apparatus according to claim 15, characterized in that, The processing module is also used for: If the scheduling priority is greater than a preset priority threshold, the delay of the first encoded frame is increased.

25. The apparatus according to any one of claims 14-17, 22, characterized in that, The transmission resources are air interface resources, the device is an access network device, and the second network element is a terminal device.

26. The apparatus according to any one of claims 14-17, 22, characterized in that, The transmission resource is a routing resource, the device is a core network device, and the second network element is an access network device.

27. A resource scheduling device, characterized in that, Includes processor and memory; among which, The memory is used to store program code; The processor is configured to invoke the level code stored in the memory to execute the method according to any one of claims 1 to 13.

28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 13.

29. A computer program product, characterized in that, Includes program code that, when a computer runs the computer program, performs the method as described in any one of claims 1 to 13.