Scheduling system of content distribution network based on edge computing
Through the scheduling system based on edge computing, combined with multi-dimensional state perception and semantic analysis, intelligent scheduling and resource allocation of edge nodes are realized, solving the problem of insufficient support for computing-intensive requests in existing CDNs and improving the resource efficiency and service quality of CDNs.
Patent Information
- Application Number
- CN202510963550.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-09-12
AI Technical Summary
The existing content delivery network (CDN) scheduling system does not fully utilize edge computing capabilities, especially in providing insufficient support for computationally intensive requests such as dynamic rendering and real-time transcoding.
An edge computing-based scheduling system is adopted, including an edge resource perception module, an NLP semantic parsing module, a reinforcement learning scheduling engine, a federated collaboration module, etc. Through multi-dimensional state perception, semantic parsing and classification, reinforcement learning-driven scheduling and resource linkage, intelligent scheduling and resource allocation of edge nodes are realized.
It improves the resource efficiency and service quality of CDN, meets user needs while fully utilizing the cache and computing power of edge nodes, realizes the technical upgrade from traditional surface feature matching to intent understanding, reduces resource idleness and improves response speed.
Smart Images

Figure CN120639869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer network technology, and in particular to a scheduling system for a content distribution network based on edge computing. Background Art
[0002] CDN (Content Delivery Network), the Chinese name is content distribution network. The main task of CDN is to deliver content from the source station to the user end as quickly as possible. The basic idea of CDN is to avoid bottlenecks and links on the Internet that may affect the speed and stability of data transmission as much as possible, so as to make content transmission faster and better. By placing edge node servers in various parts of the network to form a content distribution network, it can redirect user access requests to the closest and best edge node based on comprehensive information such as network traffic, the load of each edge node, the distance to the user, and the response time in real time.
[0003] In the existing content delivery network (CDN) scheduling system, CDN edge nodes only serve as content caching nodes, which do not fully utilize edge computing capabilities (such as CPU / GPU computing power and local processing functions), and lack support for computationally intensive requests such as dynamic rendering and real-time transcoding. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a scheduling system for a content distribution network based on edge computing, which solves the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a scheduling system for a content distribution network based on edge computing, comprising the following functional modules;
[0006] Edge resource perception module, which is deployed on edge nodes to collect computing, cache, network resource status and task queue information;
[0007] The NLP semantic parsing module is used to segment and semantically encode user requests and output a semantic label vector containing resource requirements;
[0008] A reinforcement learning scheduling engine, which includes:
[0009] State space, which is used to integrate edge node states and semantic labels;
[0010] A decision engine that generates request scheduling and resource allocation policies based on the PPO algorithm. The PPO algorithm generates scheduling policies in the action space based on the state space, where the action space contains request scheduling actions and resource scheduling actions.
[0011] The semantic and resource linkage module is used to perform resource pre-configuration such as computing power allocation and cache pre-fetching based on semantic tags, and to feed back the processing results to the NLP semantic parsing module;
[0012] The federated collaboration module uses the FedAvg algorithm to aggregate the policy gradients of edge nodes and achieves network-wide policy synchronization through the gossip protocol. It is used to achieve policy gradient aggregation and global policy synchronization between edge nodes to ensure the consistency of the distributed system.
[0013] Furthermore, the following process steps are included:
[0014] S1. Multi-dimensional state perception of edge nodes:
[0015] Deploy edge resource perception modules on edge nodes to collect computing resources, cache resources, network resources, and task queues in real time;
[0016] S2. Semantic parsing and classification:
[0017] The user request text is processed through the NLP semantic parsing module, first cleaning the URL parameters and word segmentation annotation, and then using the BERT pre-training model to generate a semantic label vector T containing the user's intention t ;
[0018] S3. Reinforcement Learning-Driven Scheduling:
[0019] Constructing a state space for a reinforcement learning scheduling engine, integrating the multi-dimensional state perception data of edge nodes with semantic label vectors to construct a complete state space that includes the real-time state of the node and the user's request intent. Within this complete state space, request scheduling and resource scheduling allocation strategies are generated. A multi-objective reward function is introduced as the "evaluation criterion" for reinforcement learning, converting the real-time processing performance of edge nodes into a quantifiable reward signal to drive policy iteration.
[0020] S4. Linking semantics and resources:
[0021] The semantic and resource linkage module performs resource pre-configuration based on semantic tags:
[0022] Static file requests: dispatched to nodes with a cache hit rate > 70% and free memory space > 50%, triggering file preloading;
[0023] Dynamic rendering requests: Assigned to nodes with ≥8 CPU cores and load <60%, and deployed as lightweight container instances through Kubernetes;
[0024] Real-time video stream request: dispatched to GPU nodes that support hardware-accelerated decoding, and pre-pushed the first key frame through the edge P2P link;
[0025] Based on the processing results of the edge nodes, the resource demand weights of the semantic tags are automatically adjusted, thus forming a semantic-driven dynamic allocation closed loop of edge resources.
[0026] Furthermore, in step S1, the computing resources include the number of CPU cores and their utilization, GPU memory usage, and remaining memory space;
[0027] Cache resources include cache hit rates and storage space usage for static files, dynamic components, and video slices;
[0028] Network resources include upstream and downstream bandwidth, RTT delay, and packet loss rate;
[0029] The task queue includes the number of pending requests and the distribution of request types, including static files, dynamic rendering, and real-time video streaming.
[0030] Furthermore, in step S3, the state space is constructed as follows:
[0031] S t =[C t , B t , Q t , L t , T t ]
[0032] Among them, C t is the computing resource vector, B t is the network resource vector, Q t Task queue characteristics, L t is the historical load index, T t is the semantic label vector.
[0033] Furthermore, in step S3, the request scheduling actions include local processing, neighboring node collaborative processing, and back-to-source processing;
[0034] Local processing: It assigns the request directly to the current edge node for processing and utilizes local cache or computing power;
[0035] Collaborative processing by adjacent nodes: forwarding requests to adjacent edge nodes, or triggering collaborative processing by multiple edge nodes;
[0036] Back-to-source processing: If the edge node cannot process the request due to reasons such as cache miss or insufficient computing power, the request will be sent back to the central server.
[0037] Furthermore, in step S3, resource scheduling includes dynamically allocating computing resources, pre-fetching cache content, and allocating dedicated channels;
[0038] Dynamically allocate computing resources: temporarily allocate CPU cores, GPU memory, or memory space for requests;
[0039] Prefetch cache content: Load relevant content into the edge node cache in advance based on semantic tags;
[0040] Dedicated channels: Allocate independent bandwidth resources to high-priority requests to avoid network congestion affecting the user experience. High-priority requests include but are not limited to live streaming and dynamic rendering.
[0041] Furthermore, the state space S t =[C t , B t , Q t , L t , T t ] and action space A t =[a1, a2, a3, a4, a5, a6] through the strategy network πθ(A t |S t ) implementation, where a1, a2, a3, a4, a5, and a6 represent local processing, collaborative processing by adjacent nodes, and back-to-source processing in the request scheduling action, as well as dynamic allocation of computing resources, pre-fetching of cache content, and division of dedicated channels in the resource scheduling action.
[0042] Furthermore, in step S3, the multi-objective reward function is as follows:
[0043] R t =αR latency +βR resource +γR qos
[0044] Among them, R t is the total reward value; α, β, γ are weight coefficients, α+β+γ=1; R latency is the response time reward; R resource is the resource utilization reward; R qos Reward for quality of service.
[0045] Furthermore, the PPO algorithm uses the total reward R t The specific process of performing policy gradient updates is as follows:
[0046] Total Reward: R t =αR latency +βR resource +γR qos
[0047] Advantage function: O t =R t +γ·V(S t+1 )-V(S t )
[0048] Among them, V(S t ) is the state value function;
[0049] Clipping probability ratio: clip(r t (θ), 1-∈, 1+∈), where
[0050] Furthermore, when the decision engine generates a scheduling strategy based on the PPO algorithm, the advantage function O t Fusion of multi-objective rewards R t The clipping probability ratio is used to prevent the strategy from updating too much. The specific gradient update formula is:
[0051] Among them, ∈ is the clipping threshold.
[0052] The present invention provides a scheduling system for a content distribution network based on edge computing, which has the following beneficial effects:
[0053] 1. This edge computing-based content distribution network scheduling system performs semantic understanding of user requests to parse the user's deep intent, accurately understand user needs and generate semantic tags. It also integrates the multi-dimensional state perception data of edge nodes with the semantic tag vector, and generates scheduling actions based on node status and user intent. This satisfies user needs while making full use of each edge node, giving edge nodes caching, computing, and scheduling capabilities, thus breaking through the limitations of traditional CDN's "surface feature matching" and achieving a technical upgrade from traditional "file type identification" to "intent understanding", and from "fixed rules" to "dynamic intelligence", ultimately improving CDN's resource efficiency and service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a schematic diagram of the process steps of the scheduling system of the content distribution network based on edge computing of the present invention;
[0055] Figure 2 A schematic diagram of dimension and action association rules of a scheduling system for a content distribution network based on edge computing of the present invention;
[0056] Figure 3 Schematic diagram comparing the effect parameters of embodiments 1 to 3 of the scheduling system of the content distribution network based on edge computing of the present invention and the traditional technical solution. DETAILED DESCRIPTION
[0057] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0058] like Figure 1-Figure 3 As shown, the present invention provides a technical solution: a scheduling system for a content distribution network based on edge computing, including the following functional modules;
[0059] Edge resource perception module, which is deployed on edge nodes to collect computing, cache, network resource status and task queue information;
[0060] The NLP semantic parsing module is used to segment and semantically encode user requests and output a semantic label vector containing resource requirements;
[0061] A reinforcement learning scheduling engine, which includes:
[0062] State space, which is used to integrate edge node states and semantic labels;
[0063] A decision engine that generates request scheduling and resource allocation policies based on the PPO algorithm. The PPO algorithm generates scheduling policies in the action space based on the state space, where the action space contains request scheduling actions and resource scheduling actions.
[0064] The semantic and resource linkage module is used to perform resource pre-configuration such as computing power allocation and cache pre-fetching based on semantic tags, and to feed back the processing results to the NLP semantic parsing module;
[0065] The federated collaboration module uses the FedAvg algorithm to aggregate the policy gradients of edge nodes and achieves network-wide policy synchronization through the gossip protocol. It is used to achieve policy gradient aggregation and global policy synchronization between edge nodes to ensure the consistency of the distributed system.
[0066] The process steps include:
[0067] S1. Multi-dimensional state perception of edge nodes:
[0068] Deploy edge resource perception modules on edge nodes to collect computing resources, cache resources, network resources, and task queues in real time;
[0069] Computing resources include the number of CPU cores and their utilization, GPU memory usage, and remaining memory space;
[0070] Cache resources include cache hit rates and storage space usage for static files, dynamic components, and video slices;
[0071] Network resources include upstream and downstream bandwidth, RTT delay, and packet loss rate;
[0072] The task queue includes the number of pending requests and the distribution of request types, including static files, dynamic rendering, and real-time video streaming;
[0073] S2. Semantic parsing and classification:
[0074] The user request text is processed through the NLP semantic parsing module. First, the URL parameters and word segmentation annotations (such as Jieba word segmentation) are cleaned, and then the BERT pre-trained model is used to generate a semantic label vector T containing the user's intention. t ;
[0075] S3. Reinforcement Learning-Driven Scheduling:
[0076] Constructing a state space for a reinforcement learning scheduling engine, integrating the multi-dimensional state perception data of edge nodes with semantic label vectors to construct a complete state space that includes the real-time state of the node and the user's request intent. Within this complete state space, request scheduling and resource scheduling allocation strategies are generated. A multi-objective reward function is introduced as the "evaluation criterion" for reinforcement learning, converting the real-time processing performance of edge nodes into a quantifiable reward signal to drive policy iteration.
[0077] Request scheduling includes local processing, collaborative processing by adjacent nodes, and back-to-source processing.
[0078] Local processing: This involves assigning requests directly to the current edge node for processing and utilizing local cache or computing power. Collaborative processing with adjacent nodes: This involves forwarding requests to adjacent edge nodes or triggering collaborative processing among multiple edge nodes. Back-to-source processing: If an edge node is unable to process a request due to reasons such as a cache miss or insufficient computing power, the request is forwarded back to the central server.
[0079] Resource scheduling includes dynamically allocating computing resources, pre-fetching cache content, and allocating dedicated channels;
[0080] Dynamically allocate computing resources: temporarily allocate CPU cores, GPU memory, or memory space for requests; pre-fetch cache content: load relevant content into the edge node cache in advance based on semantic tags; divide dedicated channels: allocate independent bandwidth resources for high-priority requests (such as 4K live streaming) to avoid network congestion affecting the experience. High-priority requests include but are not limited to live streaming and dynamic rendering.
[0081] The state space is constructed as follows:
[0082] S t =[C t , B t , Q t , L t , T t ]
[0083] Among them, C t is the computing resource vector, B t is the network resource vector, Q t Task queue characteristics, L t is the historical load index, T t is the semantic label vector;
[0084] State space S t =[C t , B t , Q t , L t , T t ] and action space A t =[a1, a2, a3, a4, a5, a6] through the strategy network πθ(A t |S t ) implementation, where a1, a2, a3, a4, a5, and a6 represent local processing, neighboring node collaborative processing, and back-to-source processing in the request scheduling action, and dynamic allocation of computing resources, pre-fetching cache content, and division of dedicated channels in the resource scheduling action;
[0085] The multi-objective reward function is as follows:
[0086] R t =αR latency +βR resource +γR qos
[0087] Among them, R t is the total reward value; α, β, γ are weight coefficients, α+β+γ=1; R latency is the response time reward; R resource is the resource utilization reward; R qos Rewards for service quality;
[0088] The PPO algorithm uses the total reward R t The specific process of performing policy gradient updates is as follows:
[0089] Total Reward: R t =αR latency +βR resource +γR qos
[0090] Advantage function: O t =R t +γ·V(S t+1 )-V(S t )
[0091] Among them, V(S t ) is the state value function;
[0092] Clipping probability ratio: clip(r t (θ), 1-∈, 1+∈), where
[0093] When the decision engine generates a scheduling strategy based on the PPO algorithm, it uses the advantage function O tFusion of multi-objective rewards R t The clipping probability ratio is used to prevent the strategy from updating too much. The specific gradient update formula is:
[0094] Among them, ∈ is the clipping threshold;
[0095] S4. Linking semantics and resources:
[0096] The semantic and resource linkage module performs resource pre-configuration based on semantic tags:
[0097] For user static file requests: dispatch to a node with a cache hit rate > 70% and free memory space > 50%, triggering file preloading;
[0098] For users' dynamic rendering requests: they are assigned to nodes with ≥8 CPU cores and a load <60%, and lightweight container instances are deployed through Kubernetes.
[0099] For users' real-time video streaming requests: dispatch to GPU nodes that support hardware-accelerated decoding, and pre-push the first key frame through the edge P2P link;
[0100] Based on the processing results of the edge nodes, the resource demand weights of the semantic tags are automatically adjusted, thus forming a semantic-driven dynamic allocation closed loop of edge resources.
[0101] Example 1, static resource acceleration scenario:
[0102] When a user requests "Get homepage product images," the NLP semantic parsing module interprets the request as "Static file - frequently accessed."
[0103] The scheduling engine matches edge node A, which has a cache hit rate of 80% and a remaining memory of 60%, and preloads other popular images in the directory into the cache of node A.
[0104] Response time is shortened by 40% compared to traditional CDN, and the cache hit rate of node A is increased to 85%;
[0105] Thus, semantic-based proactive resource scheduling replaces the traditional passive response mode and reduces resource idleness to fully utilize resources;
[0106] Example 2: Dynamic API request processing:
[0107] When a user initiates a dynamic API request to "query user orders," the NLP semantic parsing module outputs the label "dynamic rendering - medium computing power requirement."
[0108] The scheduling engine detects that the CPU load of edge node B is 50% (16 cores), allocates 4 CPU cores to the request, and creates a separate container instance.
[0109] Processing time was reduced from 150ms in the traditional solution to 90ms, and Node B's CPU utilization increased to 75% without overload;
[0110] By linking semantic tags with resource status, a closed loop is achieved from understanding user intent to pre-configuring resources and then to feedback on results, avoiding resource mismatches caused by superficial feature matching in traditional solutions.
[0111] Example 3: 4K live streaming distribution:
[0112] When a user requests "Watch 4K sports live," the NLP semantic parsing module generates the tag "Live video stream - High bandwidth - Hardware decoding."
[0113] The scheduling engine selects edge node C, which supports GPU hardware acceleration and has an uplink bandwidth of 200 Mbps. It also pre-pushes the first key frame to adjacent nodes D / E via a P2P link.
[0114] The live stream freeze rate has been reduced from 8% in traditional solutions to 2%, and the content transmission delay between edge nodes has been reduced by 50%;
[0115] The above traditional solutions adopt a centralized CDN architecture, a static caching strategy, a DNS round-robin scheduling mechanism, a stateless load balancing mechanism, fixed resource allocation for resource management, and centralized content distribution for content distribution.
[0116] Based on the above description, the present invention performs semantic understanding on user requests to parse the user's deep intentions, accurately understand user needs and generate semantic tags, and integrates the multi-dimensional state perception data of edge nodes with semantic tag vectors, and generates scheduling actions based on node status and user intentions, thereby meeting user needs while making full use of each edge node, so that the edge nodes have caching, computing, and scheduling capabilities. In this way, semantic-based active resource scheduling replaces the traditional passive response mode, reduces the resource idle rate to fully utilize resources, and through the linkage between semantic tags and resource status, realizes a closed loop from user intent understanding to resource pre-configuration to effect feedback, avoiding resource mismatch caused by surface feature matching in traditional solutions, thereby breaking through the limitations of traditional CDN "surface feature matching", and realizing technical upgrades from traditional "file type recognition" to "intention understanding", from "fixed rules" to "dynamic intelligence", and ultimately improving the resource efficiency and service quality of CDN.
[0117] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.
Claims
1. A scheduling system for a content distribution network based on edge computing, characterized by: Includes the following functional modules; Edge resource perception module, which is deployed on edge nodes to collect computing, cache, network resource status and task queue information; The NLP semantic parsing module is used to segment and semantically encode user requests and output a semantic label vector containing resource requirements; A reinforcement learning scheduling engine, which includes: State space, which is used to integrate edge node states and semantic labels; A decision engine that generates request scheduling and resource allocation policies based on the PPO algorithm. The PPO algorithm generates scheduling policies in the action space based on the state space, where the action space contains request scheduling actions and resource scheduling actions. The semantic and resource linkage module is used to perform resource pre-configuration such as computing power allocation and cache pre-fetching based on semantic tags, and to feed back the processing results to the NLP semantic parsing module; The federated collaboration module uses the FedAvg algorithm to aggregate the policy gradients of edge nodes and achieves network-wide policy synchronization through the gossip protocol. It is used to achieve policy gradient aggregation and global policy synchronization between edge nodes to ensure the consistency of the distributed system.
2. The scheduling system for a content distribution network based on edge computing according to claim 1, characterized in that: The process steps include: S1. Multi-dimensional state perception of edge nodes: Deploy edge resource perception modules on edge nodes to collect computing resources, cache resources, network resources, and task queues in real time; S2. Semantic parsing and classification: The user request text is processed through the NLP semantic parsing module, first cleaning the URL parameters and word segmentation annotation, and then using the BERT pre-training model to generate a semantic label vector T containing the user's intention t ; S3. Reinforcement Learning-Driven Scheduling: Constructing a state space for a reinforcement learning scheduling engine, integrating multi-dimensional state perception data from edge nodes with semantic label vectors to create a complete state space that includes the node's real-time state and user request intent. Within this complete state space, request scheduling and resource scheduling allocation strategies are generated. A multi-objective reward function is introduced as an evaluation criterion for reinforcement learning, converting the real-time processing performance of edge nodes into quantifiable reward signals to drive policy iteration. S4. Linking semantics and resources: The semantic and resource linkage module performs resource pre-configuration based on semantic tags: Static file requests: dispatched to nodes with a cache hit rate > 70% and free memory space > 50%, triggering file preloading; Dynamic rendering requests: Assigned to nodes with ≥8 CPU cores and load <60%, and deployed as lightweight container instances through Kubernetes; Real-time video stream request: dispatched to GPU nodes that support hardware-accelerated decoding, and pre-pushed the first key frame through the edge P2P link; Based on the processing results of the edge nodes, the resource demand weights of the semantic tags are automatically adjusted, thus forming a semantic-driven dynamic allocation closed loop of edge resources.
3. The scheduling system for a content distribution network based on edge computing according to claim 2, characterized in that: In step S1, the computing resources include the number of CPU cores and their utilization, GPU memory usage, and remaining memory space; Cache resources include cache hit rates and storage space usage for static files, dynamic components, and video slices; Network resources include upstream and downstream bandwidth, RTT delay, and packet loss rate; The task queue includes the number of pending requests and the distribution of request types, including static files, dynamic rendering, and real-time video streaming.
4. The scheduling system for a content distribution network based on edge computing according to claim 2, characterized in that: In step S3, the state space is constructed as follows: S t =[C t ,B t ,Q t ,L t ,T t ] Among them, C t is the computing resource vector, B t is the network resource vector, Q t Task queue characteristics, L t is the historical load index, T t is the semantic label vector.
5. The scheduling system for a content distribution network based on edge computing according to claim 4, characterized in that: In step S3, the request scheduling actions include local processing, neighboring node collaborative processing, and back-to-source processing; Local processing: It assigns the request directly to the current edge node for processing and utilizes local cache or computing power; Collaborative processing by adjacent nodes: forwarding requests to adjacent edge nodes, or triggering collaborative processing by multiple edge nodes; Back-to-source processing: If the edge node cannot process the request due to reasons such as cache miss or insufficient computing power, the request will be sent back to the central server.
6. The scheduling system for a content distribution network based on edge computing according to claim 5, characterized in that: In step S3, resource scheduling includes dynamically allocating computing resources, pre-fetching cache content, and allocating dedicated channels; Dynamically allocate computing resources: temporarily allocate CPU cores, GPU memory, or memory space for requests; Prefetch cache content: Load relevant content into the edge node cache in advance based on semantic tags; Dedicated channels: Allocate independent bandwidth resources to high-priority requests to avoid network congestion affecting the user experience. High-priority requests include but are not limited to live streaming and dynamic rendering.
7. The scheduling system for a content distribution network based on edge computing according to claim 6, characterized in that: State space S t =[C t , B t , Q t , L t , T t ] and action space A t =[a1, a2, a3, a4, a5, a6] through the strategy network πθ(A t |S t ) implementation, where a1, a2, a3, a4, a5, and a6 represent local processing, collaborative processing by adjacent nodes, and back-to-source processing in the request scheduling action, as well as dynamic allocation of computing resources, pre-fetching of cache content, and division of dedicated channels in the resource scheduling action.
8. The scheduling system for a content distribution network based on edge computing according to claim 2, characterized in that: In step S3, the multi-objective reward function is as follows: R t =αR latency +βR resource +γR qos Among them, R t is the total reward value; α, β, γ are weight coefficients, α+β+γ=1; R latency is the response time reward; R resource is the resource utilization reward; R qos Reward for quality of service.
9. The scheduling system for a content distribution network based on edge computing according to claim 8, characterized in that: The PPO algorithm uses the total reward R t The specific process of performing policy gradient updates is as follows: Total Reward: R t =αR latency +βR resource +γR qos Advantage function: O t =R t +γ·V(S t+1 )-V(S t ) Among them, V(S t ) is the state value function; Clipping probability ratio: clip(r t (θ), 1-∈, 1+∈), where 10. The scheduling system for a content distribution network based on edge computing according to claim 9, characterized in that: When the decision engine generates a scheduling strategy based on the PPO algorithm, it uses the advantage function O t Fusion of multi-objective rewards R t The clipping probability ratio is used to prevent the strategy from updating too much. The specific gradient update formula is: Among them, ∈ is the clipping threshold.
Citation Information
Patent Citations
Heterogeneous network multi-dimensional resource collaborative optimization method based on deep reinforcement learning
CN116132304A
Edge calculation method based on AI
CN119046010A
Policy information intelligent pushing system and method based on big data analysis
CN119493902A
Information consultation service system and method based on cloud computing platform
CN119739910A
Generative AI endogenous communication network architecture and hierarchical collaborative scheduling method
CN119815557A
Cited By
Sea area sensing task-oriented task scheduling method and device
CN121092294A
Container mirror image pre-distribution method and system based on semantic layering and collaborative scheduling
CN121309674A
GPU (Graphics Processing Unit) resource scheduling method based on data awareness and computer equipment
CN121433921A
A data-aware-based GPU resource scheduling method and computer device
CN121433921B
CDN edge node collaborative optimization method and system based on lightweight AI model
CN121585732A