Cloud side-end collaborative service method, device and equipment
By establishing cloud-edge latency models and edge-end latency models, and combining reinforcement learning and the Grey Wolf optimization algorithm, cloud-edge routing and edge resource allocation are optimized, solving the problem of excessively long XR data backhaul latency, achieving low-latency content delivery, and meeting users' real-time needs.
Patent Information
- Application Number
- CN202511402732.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-02-17
AI Technical Summary
In existing technologies, the data transmission latency of XR applications is difficult to control within 20ms, which affects the user interaction experience, especially in the process of cloud-edge transmission and edge-device transmission, where there are problems of high latency and low resource allocation efficiency.
By establishing cloud-edge latency models and edge-end latency models, and combining reinforcement learning models and the Grey Wolf optimization algorithm, we optimize cloud-edge routing information and edge-side resource allocation, and adopt a short time slot resource allocation strategy to achieve low-latency content delivery.
It effectively reduces the dwell time of XR data packets in the cloud-edge transmission network, improves the edge-end transmission rate, and enables flexible, low-latency content delivery to meet users' real-time needs.
Smart Images

Figure CN121547864A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of extended reality (XR) technology, and more particularly to a cloud-edge-device collaborative service method, apparatus, and device. Background Technology
[0002] With the rapid development and widespread application of extended reality (XR) technology, especially in fields such as industry, entertainment, and healthcare, the immersive experiences brought by its series of technologies, such as augmented reality (AR) and virtual reality (VR), are increasingly becoming one of the emerging ways for users to interact.
[0003] In related technologies, the data generated by XR applications, such as high-resolution video, rendering data, human-computer interaction and operation, have increasingly stringent requirements for network latency. Especially during the evolution of XR technology, the demand for data traffic and bandwidth has increased dramatically. The latency of XR stream return must be controlled within 20ms, otherwise it will directly affect the user's interactive experience. Summary of the Invention
[0004] This invention provides a cloud-edge-device collaborative service method, apparatus, and device. By establishing cloud-edge latency models and edge-device latency models, cloud-edge routing information and edge-side resource allocation results are determined, effectively optimizing the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network. Based on the needs of XR services, wireless frequency band resources are rationally allocated to accelerate their transmission rate in the edge-device transmission network, enabling XR data packets from the cloud or edge to be backhauled to end users with low latency, achieving flexible and low-latency content delivery.
[0005] This invention provides a cloud-edge-device collaborative service method, comprising the following steps.
[0006] Establish a cloud-edge latency model and an edge-end latency model; the cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-end latency model is used to minimize the transmission latency between the edge and the terminal. Based on the cloud-edge latency model and the edge latency model, determine the cloud-edge routing information and the edge-side resource allocation results.
[0007] According to the present invention, a cloud-edge-device collaborative service method is provided, wherein the cloud-edge latency model includes: in, Indicates the number of users within the domain; Indicates the length of the frame; Indicates whether the service requested by the user is cached on the edge. This indicates the transmission latency between the cloud and the edge. This represents the average transmission delay of data packets within a frame; This represents the average queuing delay of data packets within a frame; This represents the average transmission delay of data packets within a frame; The edge delay model includes: in, This indicates the transmission delay between the edge and the terminal; This indicates the average transmission rate of data packets within a frame; Indicates that the data packet is in the frame Average queuing delay within; Indicates user The average transmission delay of data packets.
[0008] According to the cloud-edge-device collaborative service method provided by the present invention, the step of determining cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge-device latency model includes: Based on the reinforcement learning model and the cloud-edge latency model, determine the cloud-edge routing information; Based on the Grey Wolf Optimization Algorithm (PGWO) and the edge latency model, the resource allocation results on the edge side are determined.
[0009] According to a cloud-edge-device collaborative service method provided by the present invention, the method further includes: Meta parameters are obtained using the following method: ; in, This represents the meta-parameters of the model in the previous round. Indicates the first Training parameters of the Actor network in round-element training. Indicates the first Training parameters of the Critic network in round-element training; This represents the learning rate of the outer loop Actor network and Critic network.
[0010] According to the cloud-edge-device collaborative service method provided by the present invention, the step of determining the edge-side resource allocation result based on the Grey Wolf Optimization Algorithm (PGWO) and the edge-device latency model includes: The constrained problem of determining the resource allocation results at the edge is transformed into an unconstrained problem by using a penalty function; Based on the Gray Wolf Optimization Algorithm (PGWO) and the transformed unconstrained problem, the resource allocation results on the edge side are determined.
[0011] According to a cloud-edge-device collaborative service method provided by the present invention, the method further includes: Edge-side resource allocation is performed based on time slots.
[0012] According to the cloud-edge-device collaborative service method provided by the present invention, the state space of the reinforcement learning model includes the congestion level of each node in the network. ;in, Represents a node congestion rate , and This represents the predicted node throughput, downlink inflow, and downlink outflow; t and Indicates time information.
[0013] The present invention also provides a cloud-edge-device collaborative service device, comprising the following modules: The module establishes a cloud-edge latency model and an edge-end latency model. The cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge. The edge-end latency model is used to minimize the transmission latency between the edge and the terminal. The determination module is used to determine cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge latency model.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the cloud-edge-device collaborative service method as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the cloud-edge-device collaborative service method as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the cloud-edge-device collaborative service method as described above.
[0017] The cloud-edge-device collaborative service method, apparatus, and device provided by this invention, by establishing cloud-edge latency models and edge-device latency models, determines cloud-edge routing information and edge-side resource allocation results, effectively optimizing the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network, and rationally allocating wireless frequency band resources according to the needs of XR services to accelerate their transmission rate in the edge-device transmission network, enabling XR data packets from the cloud or edge to be backhauled to end users with low latency, achieving flexible and low-latency content delivery. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the cloud-edge-device collaborative service method provided by the present invention.
[0020] Figure 2 This is a schematic diagram of the two-stage architecture of cloud-edge collaborative XR low-latency service provided by the present invention.
[0021] Figure 3 This is a schematic diagram of edge-side resource allocation based on time slot dimension provided by the present invention.
[0022] Figure 4 This is a schematic diagram of the cloud-edge routing determination method provided by the present invention.
[0023] Figure 5 This is a schematic diagram of the resource allocation method provided by the present invention.
[0024] Figure 6 This is a schematic diagram of the cloud-edge-device collaborative service device provided by the present invention.
[0025] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] The following is combined with Figures 1-7 The present invention describes a cloud-edge-device collaborative service method, apparatus, and device.
[0028] To facilitate a clearer understanding of the technical solutions of the various embodiments of this application, some technical content related to the various embodiments of this application will be introduced first.
[0029] With the rapid development and widespread application of extended reality (XR) technology, especially in industries, entertainment, and healthcare, the immersive experiences offered by its technologies, such as augmented reality (AR) and virtual reality (VR), are increasingly becoming a new way for users to interact. However, the data generated by XR applications, such as high-resolution video, rendering data, human-computer interaction, and operations, places increasingly stringent demands on network latency. Especially during the evolution of XR technology, data traffic and bandwidth requirements have increased dramatically. The latency of XR stream return must be controlled within 20ms; otherwise, it will directly affect the user's interactive experience. Therefore, network optimization has become crucial to ensuring a good user experience.
[0030] To address this challenge, XR applications typically employ a cloud-edge-device execution model. Tasks demanding high computing and storage resources are processed in the cloud, while other tasks are cached at the edge as needed. End users need to send requests to the edge server to check for a corresponding cached response; if not, they send a request to the cloud. In both cases, it's crucial to ensure that XR content is accurately and promptly delivered back to the user's device. Cloud-edge collaboration in XR content delivery can improve network bandwidth utilization and overall service quality, providing users with a smoother and more responsive XR experience. However, the current network architecture still faces numerous challenges, particularly at the cloud-edge transmission and edge-device transmission levels.
[0031] For the cloud-edge transmission segment, with the core network becoming increasingly congested, the latency of data transmission from the cloud center to the edge is too high. Optimizing the cloud-edge transmission path based on real-time network conditions to ensure efficient and low-latency data transmission has become a critical issue that urgently needs to be addressed. Relational Analysis (RL) has proven to be a promising solution. However, traditional RL algorithms struggle to meet the millisecond-level routing decision requirements of XR services due to issues such as inaccurate state space perception leading to local optima and the large action space dimension resulting in slow policy convergence.
[0032] For the edge-to-end transmission segment, the main challenge lies in how to reduce data transmission latency by rationally and effectively scheduling the resources of edge nodes. Existing research mainly focuses on wireless channel resource allocation and scheduling. Typical methods include game theory-based edge node load balancing mechanisms and time slot allocation algorithms that incorporate channel state information. However, these schemes tend to suffer from high computational complexity and low solution efficiency when solving mixed-integer programming problems involving multiple users and multiple services.
[0033] Figure 1 This is one of the flowcharts illustrating the cloud-edge-device collaborative service method provided by the present invention, such as... Figure 1 As shown, the method includes the following: Step 101: Establish cloud-edge latency model and edge-end latency model; the cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-end latency model is used to minimize the transmission latency between the edge and the terminal.
[0034] Specifically, such as Figure 2 As shown in the embodiments of this application, a two-stage architecture for cloud-edge collaborative XR low-latency services is proposed, which is mainly divided into three layers: End Layer, Edge Layer, and Cloud Layer. Furthermore, a cloud-edge routing optimization module and an edge-end resource allocation module are designed for cloud-edge transmission mode and edge-end transmission mode. Each module will be completed through four parts: data collection, modeling, environmental awareness, and decision-making, ensuring a low-latency experience for XR users.
[0035] 1) End Layer: Composed of various XR infrastructure devices, the edge XR service needs to provide efficient mobile bandwidth and low latency to end users.
[0036] 2) Edge Layer: Composed of mobile communication equipment such as base stations. We have considered factors such as the geographical location of edge nodes and divided the edge side into multiple regions, which are managed by RIC units near base stations in each region. The Central Controller communicates with the RIC units in each region to issue policies and realize the overall management of the entire edge network.
[0037] 3) Cloud Layer: This consists of servers deployed in the cloud center for each XR application. After an end user sends a request to the cloud server, the control node on the cloud side needs to find the server of the corresponding XR application based on the user's request, process the user's request, and then transmit the corresponding XR content back to the corresponding XR device through the network.
[0038] Optionally, for the two-stage architecture of cloud-edge collaborative XR low-latency services, this embodiment establishes a cloud-edge latency model and an edge-end latency model. The cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-end latency model is used to minimize the transmission latency between the edge and the terminal. In other words, this embodiment employs a two-stage network management method that combines routing optimization and resource allocation. For XR caches deployed in the cloud or edge, it optimizes the latency of the cloud-edge transmission mode and the edge-end transmission mode, thereby enabling low-latency transmission of XR data packets from the cloud or edge to the end user.
[0039] Step 102: Determine the cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge latency model.
[0040] Specifically, after establishing the cloud-edge latency model and the edge-end latency model, the cloud-edge routing information and edge-side resource allocation results can be determined through the cloud-edge latency model and the edge-end latency model in this embodiment of the application. This effectively realizes the management of the cloud-edge-end collaborative network, optimizes the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network, and reasonably allocates wireless frequency band resources according to the needs of XR services to accelerate their transmission rate in the edge-end transmission network, thereby achieving flexible and low-latency content delivery.
[0041] The method described above establishes cloud-edge latency models and edge-end latency models to determine cloud-edge routing information and edge-side resource allocation results. This effectively optimizes the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network. Based on the needs of XR services, it rationally allocates wireless frequency band resources to accelerate their transmission rate in the edge-end transmission network. This enables XR data packets from the cloud or edge to be backhauled to end users with low latency, achieving flexible and low-latency content delivery.
[0042] In some embodiments, the cloud-edge latency model includes: in, Indicates the number of users within the domain; Indicates the length of the frame; Indicates whether the service requested by the user is cached on the edge. This indicates the transmission latency between the cloud and the edge. This represents the average transmission delay of data packets within a frame; This represents the average queuing delay of data packets within a frame; This represents the average transmission delay of data packets within a frame; Edge latency models include: in, This indicates the transmission delay between the edge and the terminal; This indicates the average transmission rate of data packets within a frame; Indicates that the data packet is in the frame Average queuing delay within; Indicates user The average transmission delay of data packets.
[0043] Specifically, this application embodiment solves the problem of low-latency content delivery for XR data packet backhaul through a cloud-edge collaborative two-stage XR low-latency service provision architecture. The architecture includes a cloud-edge transport network. Edge transmission network , among which, according to The geographical location and coverage of each base station will Divided into Domain Optionally, the system is configured in a [specific context]. XR services are provided within a time frame, with each frame lasting for a duration of [number] frames. Considering scenarios where multiple users utilize random XR applications, pay attention to... Each user The data packet return of an XR service is within a set of frames. This is completed in the middle. The user first needs to request the edge server of their nearest base station to check if there is a corresponding cache; if not, a request is sent to the cloud server. Therefore, setting response variables... Representation domain Are there users on the edge server? Request cache, where Representation domain The user set within, This indicates that there is a cache, no need to send a request to the cloud, and no XR data is being sent back from the cloud; This indicates that there is no cache, and a request needs to be sent to the cloud, taking into account the latency of the XR data packets sent back from the cloud.
[0044] Generally, an XR business model is built upon the data packets that the XR server (cloud server / edge server with XR application caching) needs to transmit to the user. This involves transmitting the data packets to the end user within frames. The set of XR business models obtained within is denoted as .in, Representation domain End users within the frame The obtained XR business model set, denoted by domain End users within In frame The XR service to be transmitted back in the random XR application is a 5-tuple. ,in, ,when hour, This indicates that the XR application is cached in the cloud, when hour, This indicates that the XR application's cache is located in the edge domain. On the server. This indicates that the user requesting this service is... In other words, the XR return data packets must ultimately be transmitted to the user. . This indicates the maximum end-to-end latency allowed for the XR service requested by the user. Indicates in frame The following needs to be sent back to the user. Data packets, specifying data packets Follow arrival rate for Poisson distribution, average data packet size for The exponential distribution. This represents the number of CPU cycles (cycles / bit) required to process a unit of data, and sets the data packet propagation rate on the network. ,when At this time, the propagation speed is Divided into two sections, respectively by Routing algorithms and The resource allocation of each RRH within the organization determines the outcome; when At this time, the propagation speed is Only by The resource allocation of each RRH within the organization determines this. When At that time, the user Requested XR services data packets Placed queues of each node The queue is processed according to a first-in, first-out (FIFO) strategy. Similarly, the XR stream is processed according to... The relationship reaches the base station Later, or when The corresponding edge server After a data packet is generated, it must be passed to the queue under its respective RRH. The data packets in the queue are queued and waited for. At the end of each frame, every effort is made to ensure that the data packets in the queue are sent to the designated user.
[0045] (1) For example, the cloud-edge transmission model in this application embodiment is as follows: For a person with One forwarding node, The edges form Using undirected multi-weighted graphs Representing a frame Down The network topology model, in which Represents a forwarding node in the transmission network. This represents the links that connect the various nodes in the network. The value represents the bandwidth between nodes; a larger value indicates a larger link bandwidth. (Using...) express The congestion level of each node, among which Represents a node congestion rate , , and This represents the predicted throughput of the node, including downlink inflow and downlink outflow. (Using...) Represents nodes in the cloud-edge transmission network exist The feature matrix is used to reflect the network load. Among them, Represents a node In frame The four characteristics are packet loss rate, failure rate, throughput, and CPU frequency.
[0046] For each node There is a forwarding queue. This is used to store data packets to be forwarded. The length of the queue... .when At that time, according to the routing algorithm, from to base station The backhaul path model is ,in , For data packets The number of hops returned. For a link in the path. ,node Sending rate for: in The proportional coefficient for allocating available CPU frequency to nodes for XR services. Indicates the CPU frequency of the node .link The transmission rate is the current link bandwidth. .
[0047] (2) For example, the edge transmission model in this application embodiment is as follows: entire From an edge control center, One base station and It consists of XR terminal users, and each base station contains a BBU and Transmitting antenna RRH, and BBU The connected RRH set is , Given and users The association, that is, for ,have .
[0048] The reserved frequency band resources for XR services are updated at the beginning of each frame. Assuming shared spectrum mode, the same... All RRHs within the domain share the frequency band resources within this domain, so at the base station In the diagram, the PRBs reserved for XR services are denoted as... , and frame Relevant. (Note) The CPU frequency of the internal RRH connection server is This represents the server's processing power.
[0049] exist In this system, all users' power and spectrum resources are managed by the base station, and power resources are managed by their respective... The available bandwidth resources are allocated and set to be shared by the entire base station, and are divided into... PRBs are assigned to individual UEs. We set , , , Representing domains The PRBs to be allocated within the domain, the total transmit power of each RRH, and the channel gain matrix of each RRH within the domain.
[0050] To meet the real-time requirements of low-latency XR scenarios, the physical network is accurately modeled, and short time slots are designed as the smallest unit for wireless resource allocation. Therefore, the set of time slots is... ,in For in frame The first in There are 1 time slot. The length of each time slot is 1. ,Pick .
[0051] Set up a binary decision variable To indicate whether to use PRB In short time slots Internal allocation to users ( (Indicates true). It stipulates that within the RRH coverage area, each PRB can be assigned to at most one XR terminal user: Set continuous decision variables Indicates in time slot The amount of transmit power allocated to the PRB between the RRH and its end users is only determined when... Power will only be allocated when the value is 1. The overall allocation is as follows: Figure 3 .
[0052] Given the constraints on power allocation, the power allocated to each PRB under the same RRH is... The total power cannot exceed the maximum allocatable power of the RRH. : RRH Using PRB Serving XR end users The channel interference ratio (SINR) is: in, PRB Up users With RRH Channel gain. It is Gaussian white noise.
[0053] Since XR users belong to URLLC service slices, the coded block length of the traffic is limited, and the channel transmission rate cannot be directly obtained using Shannon's formula. Therefore, the short packet transmission mode is used to approximate the channel throughput (Mbps): in, For users and RRH The channel between them in time slots The total bandwidth of the allocated PRB, This represents the bandwidth size of each PRB block. Indicates channel dispersion, This represents the inverse of the Gaussian Q-function. This indicates the probability of a transmission error.
[0054] To ensure that a user's packets are successfully transmitted within a short time slot, we need to set constraints between channel throughput and the size of successfully transmitted packets within the short time slot: Optionally, after establishing the cloud-edge transmission model and the edge-end transmission model, the cloud-edge latency and the edge-end latency can also be determined.
[0055] (3) For example, the cloud-edge latency in this application embodiment is as follows: Based on the sending rate ,get A data packet in a frame The average transmission delay within is: The formula for calculating the current queue length is: ,So A data packet in a frame The average queuing delay within is: Based on link transmission rate ,get A data packet in a frame The average transmission delay within is: In summary, domain medium frame On average, each XR service data packet has a throughput of [number] packets per day. The communication delay is: (4) For example, the edge delay in this application embodiment is as follows: Similar to cloud-edge transmission, when After arriving at the processing server corresponding to RRH h, a data packet in a frame The average transmission rate within is: So A data packet in a frame The average transmission delay within is: Similarly, in the transmission queue Before the first data packet in the queue, its queuing delay is first checked to see if it meets the requirements. Considering the current queue length is... ,So A data packet in a frame The average queuing delay within is: Based on the channel resource allocation method, the transmission delay of intra-frame data packets depends on how many short time slots corresponding to PRB resources this XR service can occupy within that frame. Therefore, the transmission delay of a user's intra-frame data packets depends on how many short time slots corresponding to PRB resources this XR service can occupy within that frame. The average data packet transmission delay is: In summary, domain medium frame On average, each XR service data packet has a throughput of [number] packets per day. The communication delay is: The method described in the above embodiments, by establishing cloud-edge latency models and edge-end latency models, determines cloud-edge routing information and edge-side resource allocation results. This effectively optimizes the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network. Based on the needs of XR services, it rationally allocates wireless frequency band resources to accelerate their transmission rate in the edge-end transmission network, enabling XR data packets from the cloud or edge to be backhauled to end users with low latency, thus achieving flexible and low-latency content delivery.
[0056] In some embodiments, determining cloud-edge routing information and edge-side resource allocation results based on cloud-edge latency models and edge-end latency models includes: Based on the reinforcement learning model and the cloud-edge latency model, determine the cloud-edge routing information; Based on the Grey Wolf Optimization Algorithm (PGWO) and the edge latency model, the resource allocation results on the edge side are determined.
[0057] Specifically, when the cache is on an edge server, the data packets are in... The transmission delay of is the total transmission delay. When At this time, the data packet transmission latency must be included in the cache in the cloud. Transmission delay and in The transmission delay, then the frame The average latency for internal user data packet transmission is: As mentioned above, the objective of this application is to minimize the total latency of XR traffic backhaul obtained by each XR user, while ensuring that the communication latency of data packets should not exceed their maximum tolerable end-to-end latency. Therefore, in large-scale cloud-edge collaborative networks, this problem can be formulated as an NP-hard constrained optimization problem: Since the cloud-edge side and the edge-end side are managed by different controllers, this optimization problem is divided into two sub-problems: cloud-edge latency optimization and edge-end latency optimization, which are solved separately.
[0058] Optionally, in this embodiment of the application, cloud-edge routing information is determined based on a reinforcement learning model and a cloud-edge latency model.
[0059] Optionally, the three basic elements of reinforcement learning are defined as follows: ①State Space The agent in the frame The following observations are made using an environmental model, and the observed states include the network-side states. and business side status Two parts. Forming a set of observed states: Among them, network side status Including the congestion level of each node in the network Resource status information and topology information Business-side status XR service collection .
[0060] ② Action Space Agent in frame Based on the observed state, a corresponding strategy selection is generated, forming a set of action decisions: in express Regarding the backhaul path, this patent employs a multi-step reinforcement learning method with hop-based route selection. Each step corresponds to the next hop for an XR service in the current step, and the action generated by each state is derived from a series of hops in the nodes. We assume that the XR service model is indivisible, meaning it can only be transmitted along a single path from the source node to the destination node. For the current hop... The choice of its next hop Follow these basic principles: The selected node is adjacent to the previous hop; the distance between the selected node and the target node must be shorter than the distance between the previous node and the target node; the selected node does not form a cycle with the edges of previous nodes; the resources of the selected node must meet the flow requirements; the congestion rate of the selected node. The lower the better.
[0061] ③Reward Modeling We use Indicates a frame The evaluation of the selected action. The reward settings are related to the optimization goals and node selection requirements of this application, and are set as follows: It consists of three parts: the success rate of XR data packet return, the latency reward for return, and the penalty for return failure. If the XR service return is successful, then... ,set up , , among which, if If the next jump action is closer to the target node, then give... Anyway, give If the XR service is abandoned, then take We assign different penalty values based on the circumstances of the discard. For cases where the generated next jump action creates a loop with the previous jump, we assign... The penalty is imposed if the generated next-hop action is not connected to the previous hop node. The punishment.
[0062] Optionally, in this embodiment of the application, the resource allocation result on the edge side is determined based on the Grey Wolf Optimization Algorithm (PGWO) and the transformed unconstrained problem.
[0063] This application introduces a penalty function-assisted Grey Wolf Optimization Algorithm (PGWO) to solve the edge-end resource allocation problem. It simulates the allocation of selected PRBs and power among individual grey wolves in the wolf pack. Wolf Wolf, wolves and The wolves approach to train the algorithm, causing the wolf pack to converge, i.e., finding the optimal solution. For have A wolf pack of 10 wolves, with an initial allocation scheme for each wolf. Calculate the corresponding fitness and select the candidate with the smallest, second smallest, and third smallest fitness. Wolf, wolves and Wolf.
[0064] pass calculate Wolves (for each possible PRB and power allocation scenario) and Wolf, wolves and The distance to the wolves, and readjustment based on the distance. The wolf's position (in the solution scenario) is then used to reselect wolves based on their fitness in a new round. Wolf, wolves and Wolves, until they reach iteration.
[0065] in, For the updated round of allocation decisions, and It is a control factor used to control the behavior of the wolf pack. Decisions for Alpha Wolves, Beta Wolves, or Delta Wolves.
[0066] The fitness function is the objective function for optimization, as shown in the following equation: The objective function has two variables: one discrete and one continuous. For ease of processing, the discrete variable is... Scaling up to For continuous variables, the constraint is... Can be converted into constraints : To ensure that the fitness function is The following still applies, and is introduced As a new constraint: Transformation using the Big M method for and ,in and These are a maximum and a minimum number, respectively.
[0067] After this processing, the original optimization problem can be viewed as a better-expressed constraint problem: However, the GWO algorithm is suitable for solving unconstrained optimization problems, but the current optimization problem OP-2 has constraints. Therefore, the penalty function method is used to transform the fitness function of GWO into an unconstrained optimization problem, thereby determining the solution. Constraint settings penalty items Therefore, the optimization problem can be further transformed into: in, It is a function of uniform magnitude for the penalty term. For constraints... In the At that time, set penalty items. : Since the optimization objective is to minimize latency, therefore add The penalty term increases the fitness value, which means a penalty is imposed; the same applies to subsequent constraints. Set penalty items and : Similarly, for constraints When it exceeds its own constraint, a penalty of the same amount as the excess portion is imposed. Then through the function Simply change the magnitude.
[0068] The method in the above embodiments determines cloud-edge routing information based on a reinforcement learning model and a cloud-edge latency model; and determines edge-side resource allocation results based on the Grey Wolf Optimization Algorithm (PGWO) and the edge latency model, thereby achieving accurate allocation of cloud-edge routing information and edge-side resources, which can also guarantee the low-latency experience for XR users.
[0069] In some embodiments, meta-parameters are obtained in the following manner: ; in, This represents the meta-parameters of the model in the previous round. Indicates the first Training parameters of the Actor network in round-element training. Indicates the first Training parameters of the Critic network in round-element training; This represents the learning rate of the outer loop Actor network and Critic network.
[0070] Specifically, in this embodiment, the PPO model is used as the reinforcement learning model, and MPNN+DNN is used as the baseline Actor network and the baseline Critic network. The agent generates decisions from the Actor network according to the environment of the cloud-edge transmission network and the needs of XR services. Then, the Actor network at different times is trained through the Reptile method to obtain generalizable meta-parameters for deployment in actual networks.
[0071] The Actor network will state The inner loop of meta-training is used as input to obtain the action. and rewards At the same time, state It will be updated to And store the experience vector into the experience buffer. In the buffer, the Critic network will... The sampling experience in the buffer is used for training.
[0072] For an Actor network, the expected cumulative reward is defined as follows: Our goal is to find Therefore, it is necessary to follow Directional Iterative Update Strategy Parameters .
[0073] To maximize the expected value, during each iteration's update process, we consider how to utilize the current parameters. Find a better parameter , making Here, we use the concepts of advantage function and importance sampling to set the objective function. ,in The function is represented as follows: in, The strategy generated in the previous iteration, The strategy generated for the new round of iterations This represents the strategy ratio. The advantage function in the strategy selection process is expressed by the following formula: in, For state Select Action Action value function, For state The estimated value function is obtained through the Critic network.
[0074] The Actor network is optimized using the PPO-clip approach, and its policy gradient is as follows. For Internal strategies Direct optimization using policy gradients, and vice versa. First, it needs to be truncated.
[0075] For the Critic network, the TD-Error algorithm is used for learning and updating. The following formula is the loss function of the Critic network. The goal is to minimize the loss, so gradient descent is used to update the Critic network.
[0076] Each inner loop iteration updates the meta-parameters based on the previous iteration, and the Reptile algorithm is used to generate the meta-parameters in the outer loop: , in, This represents the meta-parameters of the model in the previous round. Indicates the first Training parameters of the Actor network in round-element training. Indicates the first Training parameters of the Critic network during round training. This represents the learning rate of the outer loop Actor network and Critic network.
[0077] The method described in the above embodiments, by deploying a highly generalizable meta-parameter in the system, can quickly adapt to the current environment and tasks with the assistance of a small amount of data. This allows it to serve as a deployment parameter for actual cloud-edge transmission, enabling better iteration and ensuring a low-latency experience for XR users.
[0078] In some embodiments, edge-side resource allocation is performed in terms of time slots.
[0079] Specifically, in order to meet the real-time requirements of XR low-latency scenarios, this application embodiment accurately models the physical network and designs short time slots as the smallest unit for wireless resource allocation. Therefore, the time slot set is... ,in For in frame The first in There are 1 time slot. The length of each time slot is 1. ,Pick This allows for more precise resource allocation and effectively reduces transmission latency at the edge.
[0080] The method described in the above embodiments uses short time slots as the smallest unit for wireless resource allocation, which enables more refined resource allocation and effectively reduces edge transmission latency.
[0081] In some embodiments, the state space of the reinforcement learning model includes the congestion level of each node in the network. ;in, Represents a node congestion rate , and This represents the predicted node throughput, downlink inflow, and downlink outflow; t and Indicates time information.
[0082] Specifically, in the process of determining cloud-edge routing information, this application uses express The congestion level of each node, among which Represents a node congestion rate , , and This represents the predicted throughput of the nodes, including downlink inflow and downlink outflow. In other words, network fault conditions are fully considered during the process of determining cloud-edge routing information, which makes the determined cloud-edge routes more reasonable and accurate, and effectively reduces edge transmission latency.
[0083] The method described in the above embodiments fully considers network failure scenarios during the process of determining cloud-edge routing information, thereby making the determined cloud-edge routes more reasonable and accurate, and effectively reducing edge transmission latency.
[0084] For example, this application provides a cloud-edge-device collaborative service method, the specific process of which is as follows: (1) Data acquisition mode To achieve low-latency backhaul of XR service data packets, the controller needs to collect network data in real time to allocate backhaul schemes for XR data packets based on real-time network conditions. We use one frame as a data collection cycle. The controller and the RICs of each domain on the edge side are responsible for real-time data acquisition within their respective domains. At the beginning of each frame, the network updates the XR application cache of the edge server according to the cache update rules and updates the network resource information according to the network slicing situation. Therefore, we collect data once at the beginning of each frame and update it to the database of the corresponding control unit. The collected data is shown in Table 1.
[0085] Table 1
[0086] (2) The execution mode of the cloud-edge routing optimization algorithm based on IRPPO is as follows: Figure 4 As shown, in the cloud-edge transmission module, a cloud-edge routing optimization algorithm based on the improved Reptile-PPO (IRPPO) is proposed. Meta-reinforcement learning is used to train a policy network that can adapt to multiple time periods, where the baseline network for reinforcement learning adopts the form of MPNN+DNN.
[0087] (3) The execution mode of the PGWO-driven short-slot edge-end resource allocation algorithm is as follows: Figure 5 As shown, in the side-end transmission module, short time slots are used as the smallest unit of resource allocation, which improves the utilization rate of wireless channel resources, and the PGWO heuristic algorithm is used to make low-latency decisions for the allocation of wireless resources in short time slots.
[0088] The method in this application embodiment optimizes the backhaul path of XR data packets to reduce their dwell time in the cloud-edge transmission network, and reasonably allocates wireless frequency band resources for XR services according to their needs to accelerate their transmission rate in the edge-end transmission network, thereby achieving flexible and low-latency content delivery.
[0089] The cloud-edge-device collaborative service device provided by the present invention is described below. The cloud-edge-device collaborative service device described below can be referred to in correspondence with the cloud-edge-device collaborative service method described above. The cloud-edge-device collaborative service device of the embodiments of this application is as follows: Figure 6 As shown, it includes: Establish module 610; used to establish cloud-edge latency model and edge-end latency model; cloud-edge latency model is used to minimize transmission latency between the cloud and the edge; edge-end latency model is used to minimize transmission latency between the edge and the terminal; The determination module 620 is used to determine the cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge latency model.
[0090] Figure 7 A schematic diagram of the physical structure of an electronic device is provided. This electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can invoke logical instructions in the memory 730 to execute a cloud-edge-device collaborative service method. This method includes: establishing a cloud-edge latency model and an edge-device latency model; using the cloud-edge latency model to minimize transmission latency between the cloud and the edge; using the edge-device latency model to minimize transmission latency between the edge and the terminal; and determining cloud-edge routing information and edge-device resource allocation results based on the cloud-edge latency model and the edge-device latency model.
[0091] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0092] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cloud-edge-device collaborative service method provided by the above methods. The method includes: establishing a cloud-edge latency model and an edge-device latency model; using the cloud-edge latency model to minimize the transmission latency between the cloud and the edge; using the edge-device latency model to minimize the transmission latency between the edge and the terminal; and determining cloud-edge routing information and edge-device resource allocation results based on the cloud-edge latency model and the edge-device latency model.
[0093] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the cloud-edge-device collaborative service method provided by the above methods. The method includes: establishing a cloud-edge latency model and an edge-device latency model; the cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-device latency model is used to minimize the transmission latency between the edge and the terminal; and determining cloud-edge routing information and edge-device resource allocation results based on the cloud-edge latency model and the edge-device latency model.
[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cloud-edge-device collaborative service method, characterized in that, include: Establish a cloud-edge latency model and an edge-end latency model; the cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-end latency model is used to minimize the transmission latency between the edge and the terminal. Based on the cloud-edge latency model and the edge latency model, determine the cloud-edge routing information and the edge-side resource allocation results.
2. The cloud-edge-device collaborative service method according to claim 1, characterized in that, The cloud-edge latency model includes: ; in, Indicates the number of users within the domain; Indicates the length of the frame; Indicates whether the service requested by the user is cached on the edge. This indicates the transmission latency between the cloud and the edge. This represents the average transmission delay of data packets within a frame; This represents the average queuing delay of data packets within a frame; This represents the average transmission delay of data packets within a frame; The edge delay model includes: ; in, This indicates the transmission delay between the edge and the terminal; This indicates the average transmission rate of data packets within a frame; Indicates that the data packet is in the frame Average queuing delay within; Indicates user The average transmission delay of data packets.
3. The cloud-edge-device collaborative service method according to claim 1, characterized in that, The step of determining cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge-end latency model includes: Based on the reinforcement learning model and the cloud-edge latency model, determine the cloud-edge routing information; Based on the improved Grey Wolf Optimization Algorithm (PGWO) and the edge latency model, the resource allocation results on the edge side are determined.
4. The cloud-edge-device collaborative service method according to claim 3, characterized in that, The method further includes: Meta parameters are obtained using the following method: ; ; in, This represents the meta-parameters of the model in the previous round. Indicates the first Training parameters of the Actor network in round-element training. Indicates the first Training parameters of the Critic network in round-element training; This represents the learning rate of the outer loop Actor network and Critic network.
5. The cloud-edge-device collaborative service method according to claim 3, characterized in that, The determination of edge-side resource allocation results based on the improved Grey Wolf Optimization Algorithm (PGWO) and the edge-side latency model includes: The constrained problem of determining the resource allocation results at the edge is transformed into an unconstrained problem by using a penalty function; Based on the Gray Wolf Optimization Algorithm (PGWO) and the transformed unconstrained problem, the resource allocation results on the edge side are determined.
6. The cloud-edge-device collaborative service method according to claim 1, characterized in that, The method further includes: Edge-side resource allocation is performed based on time slots; where the time slot set is... , For the first frame Each time slot.
7. The cloud-edge-device collaborative service method according to claim 3, characterized in that, The state space of the reinforcement learning model includes the congestion level of each node in the network. ;in, Represents a node congestion rate , and This represents the predicted node throughput, downlink inflow, and downlink outflow; t and Represents time information; frame .
8. A cloud-edge-device collaborative service device, characterized in that, include: Create modules; This is used to establish cloud-edge latency models and edge-end latency models; the cloud-edge latency model is used to minimize the transmission latency between the cloud and the edge; the edge-end latency model is used to minimize the transmission latency between the edge and the terminal. The determination module is used to determine cloud-edge routing information and edge-side resource allocation results based on the cloud-edge latency model and the edge latency model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the cloud-edge-device collaborative service method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cloud-edge-device collaborative service method as described in any one of claims 1 to 7.