Multi-node cloud edge collaborative large model distributed dynamic scheduling method based on global historical perception
Patent Information
- Application Number
- CN202610827351.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]为了克服现有云边协同大语言模型推理框架在处理多并发、多轮对话任务时面临的高维状态感知困难、历史上下文空间分布不均导致的语义断层、以及传统强化学习算法感知维度单一且训练易震荡等技术不足,本发明提供了一种基于全局历史感知的云边协同大模型分布式动态调度方法
本发明通过构建全域历史感知异构图,打破了传统分布式环境下的语义数据孤岛,实现了物理资源态势与长程语义关联的结构化融合感知。引入的门控势能增强机制能够自适应地权衡以通信代价换取精度增益的调度逻辑,有效抑制了动态网络波动带来的语义噪声。相比于传统启发式算法或扁平化强化学习模型,本发明在保障多轮对话语义连贯性的前提下,显著降低了跨节点协同的无效开销,并大幅提升了多目标约束下的任务完成率。此外,采用指数型时延惩罚与近端策略裁剪机制,有效解决了大模型推理随机性导致的训练不收敛难题,具备更强的环境适应性与鲁棒性。
Smart Images

Figure CN122846181A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of artificial intelligence and edge computing, specifically to a distributed dynamic scheduling method for a large cloud-edge collaborative model based on a global historical perception heterogeneous graph and a near-end policy optimization algorithm. It is applicable to complex network interaction scenarios with high requirements for contextual coherence and real-time performance, such as multi-turn dialogues. Background Technology
[0002] In edge computing environments, cloud-edge collaborative inference provides an effective paradigm for deploying large-scale language models. By rationally distributing computational tasks between the cloud and the edge, bandwidth bottlenecks and transmission latency faced by pure cloud deployments can be alleviated. However, existing distributed cloud-edge collaborative frameworks still face significant challenges when dealing with multi-turn dialogue tasks with high concurrency.
[0003] Multi-turn dialogue tasks are highly dependent on historical context. However, due to the random offloading and dynamic evolution of computational tasks, historical context is scattered across multiple inference nodes in different physical locations, forming data silos. Existing scheduling algorithms mostly focus on real-time computational load balancing and lack cross-node semantic collaboration mechanisms. This makes it difficult to accurately extract high-value historical fragments scattered across the entire network, easily leading to semantic gaps and a decrease in inference accuracy.
[0004] Furthermore, the system state encompasses both underlying physical resource metrics (such as computing power, bandwidth, and queue length) and upper-level high-dimensional semantic features. Traditional heuristic scheduling strategies and conventional deep reinforcement learning algorithms often employ one-dimensional, flattened state representations, lacking the ability to structurally represent complex heterogeneous network topologies and failing to effectively capture the coupling relationship between physical states and semantic features. This makes it difficult for the system to establish an effective mapping between communication costs and semantic benefits, and the policy network is prone to oscillations in dynamic, non-stationary environments, making it difficult to meet strict end-to-end latency constraints.
[0005] Existing cross-node historical information synchronization mechanisms are mostly mechanical, failing to fully balance the trade-off between the additional transmission latency caused by synchronizing long-range contexts and the improvement in inference accuracy. This scheduling logic, lacking coordination between physical and semantic considerations, restricts the overall performance of the system. Therefore, there is an urgent need for a distributed dynamic scheduling mechanism that can integrate graph topology awareness and reinforcement learning, accurately match fragmented historical information, and adapt to dynamic changes in physical resources. Summary of the Invention
[0006] To overcome the shortcomings of existing cloud-edge collaborative large language model inference frameworks in handling multi-concurrency, multi-turn dialogue tasks, such as difficulties in high-dimensional state perception, semantic discontinuities caused by uneven distribution of historical context space, and the single perception dimension and easy oscillation during training of traditional reinforcement learning algorithms, this invention provides a distributed dynamic scheduling method for cloud-edge collaborative large models based on global history perception.
[0007] This invention aims to jointly optimize average end-to-end latency and multi-round interactive reasoning accuracy by constructing a global history-aware heterogeneous graph and a graph-guided near-end policy optimization mechanism.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: a distributed dynamic scheduling method for a cloud-edge collaborative large-scale model based on global historical awareness, characterized in that the method includes the following steps: S1. Obtain multi-turn dialogue task requests from users in the cloud-edge collaborative reasoning system, as well as the system status of each heterogeneous reasoning model node; S2. Each heterogeneous inference model node maintains two queues locally and independently: a task buffer queue for caching tasks to be processed to represent the real-time load, and a historical information queue for storing the complete historical dialogue sequence after inference is completed. S3. Based on the task request and the system status, a joint scheduling decision is made using a near-end strategy optimization method guided by a pre-constructed global history-aware heterogeneous graph to obtain the target unloading node and cross-node historical information synchronization instructions. S4. The task request and the required historical data to be synchronized are routed to the target unloading node for inference execution, and the task parsing results are fed back to the user.
[0009] Furthermore, step S3, which describes joint scheduling decision processing using a pre-constructed global history-aware heterogeneous graph-guided near-end policy optimization method, includes: S31. Map the semantic features and performance constraints of the task request, the real-time statistical information of the heterogeneous inference model nodes, and the interaction records in the historical information queue to construct a global historical perception heterogeneous graph. S32. Extract and output the graph state representation vector of the current system through the feature aggregation mechanism in the global historical perception heterogeneous graph; S33. Optimize the agent's processing of the graph state representation vector through a proximal strategy, and output a global gating signal and node selection probability distribution; S34. Based on the global gating signal and node selection probability distribution, determine the target unloading node and whether to trigger cross-node historical information synchronization.
[0010] Furthermore, step S31, which maps and constructs a global history-aware heterogeneous graph by mapping the semantic features and performance constraints of the task request, the real-time statistical information of the heterogeneous inference model nodes, and the interaction records in the historical information queue, includes: S311. Map the multi-turn dialogue task request to a task node, map the interaction record in the historical information queue to a historical node, and map the heterogeneous inference model node to an inference node; S312. Establish a storage edge between the historical node and the physically associated inference node, and establish a candidate edge between the task node and all available inference nodes; S313. Introduce a virtual synchronization edge between the historical node and the non-local inference node to characterize the feasibility of cross-node semantic collaboration and data migration.
[0011] Furthermore, after introducing a virtual synchronization edge between the historical node and the non-local inference node as described in step S313, the method further includes: S3131. Obtain the actual transmission delay between the source node and the target node corresponding to the virtual synchronization edge; S3132. Calculate the dynamic exponential decay weight based on the ratio of the actual transmission delay to the maximum end-to-end delay constraint of the current task request; S3133. The dynamic exponential decay weight is assigned to the virtual synchronization edge so that when network congestion leads to excessive transmission costs, attention to remote historical nodes is automatically suppressed in the topology.
[0012] Furthermore, step S32, which involves extracting and outputting the graph state representation vector of the current system through the feature aggregation mechanism in the global historical perception heterogeneous graph, includes: S321. Generate a three-dimensional state-aware preference vector using an agent based on deep reinforcement learning, and use it as a global gating signal; S322. Utilize the global gating signal to drive the gating potential enhancement mechanism to dynamically adjust the feature fusion weights of the inference node's inherent physical state characteristics, the local cache history characteristics based on the storage edge connection, and the remote history characteristics based on the virtual synchronization edge connection, respectively. S323. The feature embedding representation of the inference node is updated by adaptive weighted fusion calculation, and pooled and merged into a graph state representation vector representing the current system.
[0013] Furthermore, in step S322, when performing feature weighted fusion using the global gating signal-driven gating potential enhancement mechanism, the fusion term of the remote historical features based on the virtual synchronization edge connection is subject to dual constraints: S3221. The first constraint is the semantic attention coefficient calculated based on the semantic association strength between historical nodes and the current task node; S3222. The second constraint is the dynamic exponential decay weight corresponding to the virtual synchronization edge, in order to balance the physical communication cost of cross-node state synchronization with the semantic reasoning gain.
[0014] Furthermore, step S33, which involves optimizing the agent's processing of the graph state representation vector through a proximal policy to output a global gating signal and a node selection probability distribution, includes: S331. Construct a hierarchical and decoupled action space, and use a policy network to receive the graph state representation vector; S332. In the high-level action space, the global gating signal that satisfies the simplex constraint is generated through the Softmax classification reparameterization sampling mechanism to alleviate the policy training oscillation in the high-dimensional continuous network state; S333. In the low-level action space, based on the enhanced graph features modulated by the global gating signal, the node selection probability distribution is calculated by a multilayer perceptron, and classification sampling is performed based on this distribution during the exploration phase to determine the target unloading node.
[0015] Furthermore, step S34, which involves determining the target unloading node and whether to trigger cross-node historical information synchronization based on the global gating signal and node selection probability distribution, includes: S341. Select the target unloading node based on the node selection probability distribution; S342. Based on the cooperative transmission preference weights in the global gating signal, evaluate and trigger cross-node historical information synchronization; S343. The proximal policy optimization agent used to output the gating signal and probability distribution is trained based on an immediate composite utility function, which is composed of an accuracy gain term and a delay violation penalty term. When the target unloading node meets the minimum inference accuracy requirement, the accuracy gain term provides a positive utility evaluation. S344. The delay violation penalty item adopts an exponential penalty mechanism, and its penalty coefficient is determined by the ratio of the actual end-to-end total delay to the maximum delay constraint of the task; wherein, the actual end-to-end total delay includes task upload delay, cross-node historical information synchronization transmission delay, queue waiting delay, and inference delay after strengthening historical context.
[0016] The second aspect of this invention relates to a distributed dynamic scheduling device for a cloud-edge collaborative large-scale model based on global historical awareness, comprising:
[0017] The task request and status acquisition module is used to acquire multi-turn dialogue task requests from users in the cloud-edge collaborative inference system, as well as the system status of each heterogeneous inference model node. Each heterogeneous inference model node independently maintains a task buffer queue and a historical information queue locally. The decision instruction determination module is used to perform joint scheduling decision processing based on the task request and the system state, through a pre-built global history-aware heterogeneous graph-guided near-end policy optimization algorithm, to obtain the target unloading node and cross-node historical information synchronization instructions. The task execution and result feedback module is used to route the task request and the required historical data to the target unloading node for inference execution, and to provide feedback on the task parsing results to the user.
[0018] The third aspect of this application relates to an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a distributed dynamic scheduling method for a cloud-edge collaborative large model based on global history awareness, as proposed in this invention.
[0019] The beneficial effects of this invention are: This invention breaks down semantic data silos in traditional distributed environments by constructing a heterogeneous graph of global historical perception, achieving structured fusion perception of physical resource status and long-range semantic associations. The introduced gating potential enhancement mechanism adaptively balances the scheduling logic of trading communication costs for accuracy gains, effectively suppressing semantic noise caused by dynamic network fluctuations. Compared to traditional heuristic algorithms or flattened reinforcement learning models, this invention significantly reduces the ineffective overhead of cross-node collaboration while ensuring the semantic coherence of multi-turn dialogues, and greatly improves the task completion rate under multi-objective constraints. Furthermore, the use of exponential latency penalties and near-end policy pruning mechanisms effectively solves the training non-convergence problem caused by the randomness of large model inference, exhibiting stronger environmental adaptability and robustness. Attached Figure Description
[0020] Figure 1 This is a system framework diagram of the present invention.
[0021] Figure 2 This is a HAN topology view of the present invention.
[0022] Figure 3 This is the HAN calculation view of the present invention.
[0023] Figure 4 This is a view of the GH-HG-PPO algorithm framework of the present invention. Detailed Implementation
[0024] The following detailed description of the distributed dynamic scheduling method for a multi-node cloud-edge collaborative large model based on full-domain historical perception, proposed in this invention, is provided in conjunction with the accompanying drawings and specific embodiments.
[0025] Example 1
[0026] like Figure 1As shown, this embodiment relates to a distributed dynamic scheduling method for a multi-node cloud-edge collaborative large model based on full-domain historical perception, including: Step 1: Construct a distributed cloud-edge collaborative inference system environment. The system consists of edge access points (APs), a cloud-based large language model server, and... It consists of several heterogeneous inference nodes (SLMs). The edge nodes deploy the Doubao-1.5 series models, while the cloud deploys the DeepSeek-V3 model to simulate asymmetric computing power distribution. Each inference node $m$ maintains a task buffer queue locally. and historical information queue The history queue stores the complete historical dialogue sequence for which reasoning has been completed. This is used to characterize the prior semantic accumulation of a node in a specific domain.
[0027] Step 2: Feature definition of multi-turn dialogue task requests. (In discrete time slots) Inside, users Generate request The input text is mapped to semantic feature vectors using a RoBERTa-base pre-trained model. The task request is processed via a quintuple. Full description, in which The number of text tokens to input. As a minimum accuracy constraint, This defines the maximum end-to-end latency constraint. This definition provides the data foundation for subsequent physical load calculations and semantic matching metrics.
[0028] Step 3: Establish an end-to-end total latency model. The total system latency consists of five parts: wireless transmission, forwarding, synchronization, queuing, and inference. The latency for a task to be transmitted from the wireless channel to the access point is... The newly arrived task is at the node. Queuing delay Based on the cumulative workload calculation, the formula is as follows:
[0029] in Characterizes the real-time hardware performance of a node in processing a single token. Actual inference latency. The distinction is made based on whether history enhancement is triggered. When historical context is used, the inference time is calculated based on the semantic correlation strength between the current task and the historical records.
[0030] Step 4: Propose the core optimization problem to be solved by this invention. Addressing the conflict between heterogeneous physical resources and semantic fragmentation in distributed scenarios, this invention formalizes scheduling decisions into a constrained Markov decision process (CMDP), establishing the following mathematical optimization model. :
[0031] In the formula, the decision variable is the unloading target node. And whether to trigger cross-node historical synchronization decision The challenge of this problem lies in satisfying hard delay constraints. Under the premise of intelligent path selection and synchronization strategy, the overall system utility is maximized. .
[0032] Step 5: The problem is transformed into a solvable reinforcement learning proposition P2. Since the hard constraints in P1 are difficult to solve directly using traditional convex optimization methods in dynamic environments, this invention utilizes an exponential barrier function to transform the constraints... and This is transformed into a penalty term in the reward function, thus converting the original problem into an unconstrained Markov decision process optimization proposition P2:
[0033] in, This is a composite reward function that integrates a latency penalty term and a precision gain term. This transformation ensures the smoothness of the policy gradient during training, enabling end-to-end solution of scheduling problems under complex constraints through deep reinforcement learning.
[0034] Step 6: Construct a global history-aware heterogeneous graph (GH-HG). To address the challenge of high-dimensional state awareness in P1, the system maps the global state to a graph structure. Node set Includes task, history, and reasoning nodes. Edge set. Besides including physical storage edges and task candidate edges, the key feature is the introduction of virtual synchronization edges. These edges establish a logical connection between remote historical nodes and the current inference node, with their weights... The consumption of inter-node communication bandwidth on latency budget is directly perceived, and communication cost considerations are embedded in the graph construction stage.
[0035] Step 7: Projection and Alignment of Heterogeneous Features in the Latent Space. Since the semantic features and physical load features have different dimensions, this invention utilizes a feature projection matrix. Projecting heterogeneous node features onto a latent space of uniform dimension. The projection process is formally represented as follows: This eliminates the dimensional differences, laying an algebraic foundation for subsequent semantic aggregation based on meta-paths. Then, the first stage of hierarchical shared two-stage graph aggregation is executed. The system performs a gateless basic graph convolutional aggregation based on the current graph topology, generating an initial set of node embeddings and constructing the basic observation state. The basic observation space encompasses environmental information such as mission requirements, node load, and network status.
[0036] Step 8: Generate gating preference signals. The reinforcement learning agent (Actor network) generates gating preference signals based on the observed states. The Gaussian distribution in the output continuous action space is sampled and projected to generate a gated preference vector that satisfies the simplex constraint. .
[0037] Step 9: Three-way Gated Potential Feature Aggregation. Using this preference vector as the global control signal, the Gated Potential Feature Aggregation (GPA) mechanism is driven to perform the second-stage feature fusion. The update formula is as follows: , in, represents the semantic attention coefficients between nodes. This mechanism dynamically adjusts the weights of the physical state preservation term, the local storage enhancement term, and the virtual collaborative injection term.
[0038] Step 10: Determine the optimal cooperative scheduling decision. Calculate the matching score between the task embedding and the enhanced embedding of each inference node using a multilayer perceptron (MLP), mapping it to the discrete selection probability distribution of the target node: The system uses this information to determine the optimal unloading inference node. And whether to trigger cross-node historical synchronization decisions.
[0039] Step 11: Scheduling and Status Feedback. The access point routes tasks based on the decision. During the inference execution phase, the total end-to-end latency cost is considered. This includes transmission latency, queuing latency, and inference latency, which includes historical context fusion overhead. Queuing latency is based on the node's real-time computing power. Calculation of cumulative workload.
[0040] Step 12: Construct a composite utility reward function. To improve overall system performance while satisfying time delay constraints, a composite utility reward function is constructed. Represented as: ,in, For inference accuracy gain, This study employs an exponential barrier penalty mechanism to penalize delay violations, thereby characterizing the system losses caused by breaching delay constraints.
[0041] Step 13: Policy Optimization and Gradient Backpropagation. The policy network parameters are updated using the Proximal Policy Optimization (PPO) algorithm. With value network parameters The policy update step size is limited by minimizing the pruning objective function: Simultaneously, the advantage function is calculated using generalized advantage estimation (GAE). To balance the bias and variance.
[0042] Step 14: End-to-end closed-loop optimization. The action gradient is directly transmitted back to the underlying global history-aware heterogeneous graph network using the gradient backpropagation mechanism, realizing the synchronous and coordinated evolution of feature dimensionality reduction, semantic aggregation, and resource scheduling strategies.
[0043] Step 15: Model Deployment and Inference Path Selection. In actual deployment, the system, based on the trained parameters, perceives the heterogeneous graph state in the current environment, selects the optimal unloading path and synchronization strategy according to the highest probability, and updates the historical records of the nodes after inference is completed.
[0044] Step 16: Maximize long-term cumulative returns. By iterating through the above steps, guide the agent to optimize the cumulative discounted composite utility of the system over an infinite time domain.
[0045] This embodiment addresses the challenges of low-sample historical dependencies and resource heterogeneity in complex distributed cloud-edge environments with multiple concurrent tasks. For example... Figure 1 and Figure 4 As shown, the system achieves a multi-objective balance between task routing, historical data reuse, and collaborative reasoning through a joint decision-making architecture. In this scheme, the system is configured with M=4 heterogeneous nodes, including the Doubao-1.5 series models located at the edge and the DeepSeek-V3 model located in the cloud. By introducing a global history-aware heterogeneous graph, the system breaks down data silos in traditional distributed environments, enabling node features to perceive the semantic correlation strength across nodes.
[0046] For example Figure 2 , Figure 3 and Figure 4 The graph aggregation and PPO optimization logic shown in the figure involves setting an initial learning rate during training. A discount factor of 0.99 is used to estimate the generalized advantage, and the precision gain weight is incorporated into the reward function. Set to 10, latency penalty weight Set to 1.0. This configuration effectively guides the agent to maintain low end-to-end latency while ensuring inference accuracy. Experimental data shows that, compared with baseline algorithms such as Vanilla DRL, the GH-HG-PPO algorithm proposed in this embodiment achieves optimal matching between computational resources and dynamic semantic requirements while ensuring high inference accuracy and reducing average end-to-end latency.
[0047] Example 2
[0048] This embodiment relates to a distributed dynamic scheduling device for a cloud-edge collaborative large-scale model based on global history awareness, used to implement the distributed dynamic scheduling method for a cloud-edge collaborative large-scale model based on global history awareness described in Embodiment 1, including:
[0049] The task request and status acquisition module is used to acquire multi-turn dialogue task requests from users in the cloud-edge collaborative inference system, as well as the system status of each heterogeneous inference model node. Each heterogeneous inference model node independently maintains a task buffer queue and a historical information queue locally. The decision instruction determination module is used to perform joint scheduling decision processing based on the task request and the system state, through a pre-built global history-aware heterogeneous graph-guided near-end policy optimization algorithm, to obtain the target unloading node and cross-node historical information synchronization instructions. The task execution and result feedback module is used to route the task request and the required historical data to the target unloading node for inference execution, and to provide feedback on the task parsing results to the user.
[0050] Example 3
[0051] This embodiment relates to an electronic device for implementing the distributed dynamic scheduling method of cloud-edge collaborative large model based on global history awareness as described in Embodiment 1. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the distributed dynamic scheduling method of cloud-edge collaborative large model based on global history awareness of the present invention.
[0052] The embodiments described in this specification are merely illustrative examples of implementations of the inventive concept. The scope of protection of this invention should not be limited to the specific forms presented in these embodiments, but also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A distributed dynamic scheduling method for a cloud-edge collaborative large-scale model based on global historical awareness, characterized in that... ,include: S1 acquires multi-turn dialogue task requests from users in the cloud-edge collaborative reasoning system, as well as the system status of each heterogeneous reasoning model node; S2 In this context, each heterogeneous inference model node independently maintains two queues locally: a task buffer queue for caching tasks to be processed to represent the real-time load, and a historical information queue for storing the complete historical dialogue sequence after inference is completed. Based on the task request and the system status, S3 performs joint scheduling decision processing through a near-end strategy optimization method guided by a pre-constructed global history-aware heterogeneous graph, and obtains the target unloading node and cross-node historical information synchronization instructions. S4 routes the task request and the required historical data to the target unloading node for inference execution, and provides feedback on the task parsing results to the user.
2. The method according to claim 1, characterized in that... Step S3, which describes joint scheduling decision processing using a pre-constructed global history-aware heterogeneous graph-guided near-end policy optimization method, includes: S31 maps and constructs a global historical perception heterogeneous graph by mapping the semantic features and performance constraints of the task request, the real-time statistical information of the heterogeneous inference model nodes, and the interaction records in the historical information queue. S32 extracts and outputs the graph state representation vector of the current system through the feature aggregation mechanism in the global historical perception heterogeneous graph; S33 optimizes the agent's processing of the graph state representation vector through a near-end strategy, and outputs a global gating signal and a node selection probability distribution. Based on the global gating signal and the node selection probability distribution, S34 determines the target unloading node and whether to trigger cross-node historical information synchronization.
3. The method according to claim 2, characterized in that... Step S31, which involves mapping and constructing a global historical awareness heterogeneous graph by the semantic features and performance constraints of the task request, the real-time statistical information of the heterogeneous inference model nodes, and the interaction records in the historical information queue, includes: S311 maps the multi-turn dialogue task request to a task node, the interaction record in the historical information queue to a historical node, and the heterogeneous reasoning model node to a reasoning node. S312 establishes a storage edge between the historical node and the physically associated inference node, and establishes a candidate edge between the task node and all available inference nodes; S313 introduces a virtual synchronization edge between the historical node and the non-local inference node to characterize the feasibility of cross-node semantic collaboration and data migration.
4. The method according to claim 3, characterized in that... After introducing a virtual synchronization edge between the historical node and the non-local inference node as described in step S313, the method further includes: S3131 Obtains the actual transmission delay between the source node and the target node corresponding to the virtual synchronization edge; S3132 calculates the dynamic exponential decay weight based on the ratio of the actual transmission delay to the maximum end-to-end delay constraint of the current task request. S3133 assigns the dynamic exponential decay weight to the virtual synchronization edge so that when network congestion leads to excessive transmission costs, attention to remote historical nodes is automatically suppressed in the topology.
5. The method according to claim 2, characterized in that... Step S32, which involves extracting and outputting the graph state representation vector of the current system through the feature aggregation mechanism in the global history-aware heterogeneous graph, includes: S321 uses a deep reinforcement learning-based agent to generate a three-dimensional state-aware preference vector, which is then used as a global gating signal. S322 utilizes the global gating signal to drive the gating potential enhancement mechanism, dynamically adjusting the feature fusion weights of the inference node's inherent physical state characteristics, the local cache history characteristics based on the storage edge connection, and the remote history characteristics based on the virtual synchronization edge connection. S323 updates the feature embedding representation of the inference node through adaptive weighted fusion calculation and pools and merges it into a graph state representation vector representing the current system.
6. The method according to claim 5, characterized in that... In step S322, when performing feature weighted fusion using the global gating signal-driven gating potential enhancement mechanism, the fusion term of the remote historical features based on the virtual synchronization edge connection is subject to dual constraints: The first constraint in S3221 is the semantic attention coefficient calculated based on the semantic association strength between historical nodes and current task nodes; The second constraint S3222 is the dynamic exponential decay weight corresponding to the virtual synchronization edge, in order to balance the physical communication cost of cross-node state synchronization with the semantic reasoning gain.
7. The method according to claim 2, characterized in that... Step S33, which involves optimizing the agent's processing of the graph state representation vector using a near-end policy to output a global gating signal and a node selection probability distribution, includes: S331 constructs a hierarchical and decoupled action space, and uses a policy network to receive the graph state representation vector; In the high-level action space, S332 generates the global gating signal that satisfies the simplex constraint through the Softmax classification reparameterization sampling mechanism to alleviate policy training oscillations in high-dimensional continuous network states. In the low-level action space, S333 calculates the node selection probability distribution based on the enhanced graph features modulated by the global gating signal using a multilayer perceptron, and performs classification sampling based on this distribution during the exploration phase to determine the target unloading node.
8. The method according to claim 2, characterized in that... Step S34, which involves determining the target unloading node and whether to trigger cross-node historical information synchronization based on the global gating signal and node selection probability distribution, includes: S341 Selects the target unloading node based on the node selection probability distribution; S342 evaluates and triggers cross-node historical information synchronization based on the cooperative transmission preference weights in the global gating signal; The proximal policy optimization agent used by S343 to output the gating signal and probability distribution is trained based on an immediate composite utility function, which is composed of an accuracy gain term and a delay violation penalty term. When the target unloading node meets the minimum inference accuracy requirement, the accuracy gain term provides a positive utility evaluation. The delay violation penalty item described in S344 adopts an exponential penalty mechanism, and its penalty coefficient is determined by the ratio of the actual end-to-end total delay to the maximum delay constraint of the task; wherein, the actual end-to-end total delay includes task upload delay, cross-node historical information synchronization transmission delay, queue waiting delay, and inference delay after strengthening historical context.
9. A distributed dynamic scheduling device for a cloud-edge collaborative large-scale model based on global historical perception, characterized in that... ,include: The task request and status acquisition module is used to acquire multi-turn dialogue task requests from users in the cloud-edge collaborative inference system, as well as the system status of each heterogeneous inference model node. Each heterogeneous inference model node independently maintains a task buffer queue and a historical information queue locally. The decision instruction determination module is used to perform joint scheduling decision processing based on the task request and the system state, through a pre-built global history-aware heterogeneous graph-guided near-end policy optimization algorithm, to obtain the target unloading node and cross-node historical information synchronization instructions. The task execution and result feedback module is used to route the task request and the required historical data to the target unloading node for inference execution, and to provide feedback on the task parsing results to the user.
10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that... When the processor executes the computer program, it implements a distributed dynamic scheduling method for a cloud-edge collaborative large model based on global history awareness, as described in any one of claims 1-8.