Distribution automation terminal edge cluster adaptive access method and device based on cloud edge fusion, terminal equipment and storage medium

By constructing an adaptive access method for edge clusters of power distribution automation terminals based on cloud-edge fusion, Markov decision-making and quadruple experience playback pools optimize queuing delay and load balancing, the problem of poor delay and balance in the existing technology is solved, and the system's response performance and efficiency are improved.

CN120343022APending Publication Date: 2025-07-18ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394886.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the power distribution automation business, the prior art cannot learn and optimize targetedly when the delay and equalization are poor, making it difficult for the system to maintain stable millisecond-level response performance.

Method used

A power distribution automation terminal edge cluster adaptive access method is constructed based on cloud-edge fusion. By obtaining the data queue and average data arrival rate of the power distribution equipment, the state space, action space and reward function of Markov's decision are determined, and the four-fold experience playback pool is used to store data classification and explore weight adjustments, combining local and global network parameters updates to optimize queuing delay and load balancing.

Benefits of technology

It achieves targeted learning and optimization when facing different performance challenges, and improves the response speed and system efficiency of power distribution automation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343022A_ABST
    Figure CN120343022A_ABST
Patent Text Reader

Abstract

The invention discloses a distribution automation terminal edge cluster adaptive access method and device based on cloud edge fusion, terminal equipment and a storage medium, and relates to the field of electric power automation, and the method comprises the steps: obtaining related data to construct a Markov decision; obtaining a joint decision of the action space from the data of the state space through a local network, updating the data of the state space after executing the decision, and storing the data in a quadruple experience playback pool in a classified manner; the exploration weight of the quadruple experience playback pool is determined, sampling is carried out based on the exploration weight, a predicted action space is obtained through a local target network, target values are calculated, and all local network parameters are updated; and the cloud server updates the global network parameters and issues the parameters to each edge server, and each edge server obtains and executes a joint decision based on the global network parameters. By implementing the method and the system, the problem that the distribution automation service cannot be learned and optimized in a targeted manner when the time delay and the balance degree are relatively poor can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power automation, and in particular to a method, device, terminal device and storage medium for adaptive access of an edge cluster of a distribution automation terminal based on cloud-edge integration. Background Art

[0002] With the rapid development of the power Internet of Things, the large-scale deployment of distribution automation terminals and their high-frequency data collection requirements not only drive the geometric growth of power service data, but also emphasize the importance of interconnection between terminals. Limited by the bandwidth of the network transmission channel, the traditional cloud computing architecture faces problems such as high transmission delay and insufficient data processing timeliness in the distribution automation scenario, directly affecting the real-time response ability of the system. To address this challenge, an edge cluster constructed by multiple edge servers, relying on its spatial deployment advantage close to power terminal devices, can sink data computing tasks to network edge nodes to achieve efficient local processing of data, and further promote the interconnection within the edge cluster and with the cloud. The cloud-edge collaborative architecture realizes the elastic expansion and optimal scheduling of computing resources by integrating cloud computing resources and real-time processing capabilities on the edge side. However, in the existing scenario of dynamic access of distribution terminals, due to the lack of a dynamic coordination mechanism for computing resource allocation between the cloud and the edge side, server queue congestion occurs within the edge cluster, and the data packet loss rate of distribution terminals increases, making it difficult for the system to maintain stable millisecond-level response performance.

[0003] Therefore, there is an urgent need for a method for adaptive access of an edge cluster of a distribution automation terminal based on cloud-edge integration, which is of great significance for optimizing system performance, improving operation efficiency, and ensuring the real-time performance of distribution automation. Deep reinforcement learning has powerful environmental perception and decision-making capabilities and has been widely used for optimizing the adaptive access of an edge cluster of a distribution automation terminal. In the face of a dynamic environment, deep reinforcement learning can significantly improve the optimization effect and achieve the efficient operation of the distribution network and the optimal allocation of resources.

[0004] The prior art one proposes a method for optimizing the access of a distribution terminal based on a federated deep Q network. Through federated deep Q network reinforcement learning, the environment is perceived and decisions are made. By considering queue backlog, the load balance degree between the distribution terminal and the edge cluster is quantified, and the efficient scheduling of edge Internet of Things resources and the optimization of terminal access are realized. The prior art two proposes a method for optimizing the access of a distribution terminal based on DAC. Through the cooperation of the actor network and the critic network, the access strategy of the distribution terminal is intelligently adjusted to achieve the efficient operation of the power system and the optimal allocation of resources. None of the above prior arts constructs different experience replay pools for delay and balance degree performance, and it is difficult to solve the problem that distribution automation services cannot learn and optimize targeted when facing poor delay and balance degree. Summary of the Invention

[0005] The embodiments of the present invention provide a method, device, terminal device and storage medium for adaptive access of an edge cluster of a distribution automation terminal based on cloud-edge integration. The present invention can solve the problem that distribution automation services cannot learn and optimize targeted when facing poor delay and balance degree.

[0006] An embodiment of the present invention provides a method for adaptive access of an edge cluster of a distribution automation terminal based on cloud-edge integration, including:

[0007] Obtain the data queues and data average arrival rates of each distribution device;

[0008] Based on the data queues and data average arrival rates of each distribution device, determine the state space, action space and reward function of Markov decision-making;

[0009] Input the data in the state space into the initial local actor network of each edge server to obtain a joint decision in the action space. After each edge server executes the joint decision, calculate the value of the reward function and update the data in the state space, and classify and store the updated data in the state space into a quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, a balance degree replay pool, a delay improvement replay pool and a balance degree improvement replay pool;

[0010] According to the data classification in the quadruple experience replay pool, determine the exploration weights of the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor network to obtain a predicted action space, and based on the sampled state space data and the predicted action space, and update all local network parameters of each edge server according to the value of the reward function;

[0011] Upload all the updated local network parameters of each edge server to the cloud server, the cloud server performs global network parameter update, and send the updated global network parameters to each edge server, and each edge server obtains and executes the joint decision in the action space based on the global network parameters.

[0012] Further, the determining the state space, action space and reward function of Markov decision-making based on the data queues and data average arrival rates of each distribution device includes:

[0013] Based on the data queues and data average arrival rates of each distribution device, with the goal of minimizing the weighted sum of queuing delay and load balance degree, construct a delay balance degree optimization model, and impose comprehensive condition constraints on the delay balance degree optimization model; wherein, the comprehensive condition constraints include: distribution terminal access constraints, data allocation ratio constraints, edge server computing resource allocation constraints and cloud server computing resource allocation constraints;

[0014] Model the delay equalization degree optimization model as a Markov decision, and determine the state space, action space, and reward function of the Markov decision.

[0015] Furthermore, the delay equalization degree optimization model includes:

[0016]

[0017] In the formula, P is the minimum value of the weighted sum of the queuing delay and the load balancing degree, and x m,n (t) is the access state of the distribution terminal, is the proportion of the data of distribution terminal m processed by edge server n at the t-th time slot, is the proportion of the data of distribution terminal m transmitted from edge server n to the cloud server for processing at the t-th time slot, is the CPU cycle frequency of edge server n for processing the data of distribution terminal m, is the CPU cycle frequency of the cloud server for processing the data of distribution terminal m, k is the weight of the original optimization target, is the weighted sum of the queuing delay and the load balancing degree at the t-th time slot, M is M distribution terminals, is the data queue maintained by distribution terminal m at the t-th time slot, is the amount of acquisition data uploaded from the acquisition layer to distribution terminal m at the t-th time slot, is the amount of data transmitted from distribution terminal m to the edge cluster layer at the t-th time slot, N is N edge servers, is the data queue maintained by edge server n from distribution terminal m at the t-th time slot, is the amount of data of distribution terminal m received by edge server n at the t-th time slot, is the amount of data of distribution terminal m processed by edge server n at the t-th time slot, is the data queue of distribution terminal m maintained by the cloud server at the t-th time slot, is the amount of data of distribution terminal m received by the cloud server after being forwarded by the edge server at the t-th time slot, is the amount of data of distribution terminal m processed by the cloud server at the t-th time slot.

[0018] Furthermore, the state space of the Markov decision includes: the data queue and the average data arrival rate of each power distribution device; among them, the power distribution device includes: distribution terminal, edge server, and cloud server;

[0019] The action space includes: the joint decision of the access state of the distribution terminal, the data processing allocation ratio, and the computing resource allocation;

[0020] The reward function is the negative value of the minimum value of the weighted sum of the queuing delay and the load balancing degree.

[0021] Further, the delay replay pool includes: a first delay replay pool, a second delay replay pool, a third delay replay pool, and a fourth delay replay pool; the balance degree replay pool includes: a first balance degree replay pool, a second balance degree replay pool, a third balance degree replay pool, and a fourth balance degree replay pool; the delay improvement replay pool includes: a first delay improvement replay pool, a second delay improvement replay pool, and a third delay improvement replay pool; the balance degree improvement replay pool includes: a first balance degree improvement replay pool, a second balance degree improvement replay pool, and a third balance degree improvement replay pool;

[0022] The storing the data of the updated state space into the quadruple experience replay pool class by class includes:

[0023] Based on the data of the updated state space, calculate the corresponding queuing delay and load balance degree; according to the queuing delay and the load balance degree, calculate the delay improvement state and the balance degree improvement state;

[0024] Compare the queuing delay with each preset classification threshold of the delay replay pool. When the queuing delay is less than the first delay threshold, store the data of the corresponding updated state space into the first delay replay pool; when the queuing delay is between the first delay threshold and the second delay threshold, store the data of the corresponding updated state space into the second delay replay pool; when the queuing delay is between the second delay threshold and the third delay threshold, store the data of the corresponding updated state space into the third delay replay pool; when the queuing delay is greater than the third delay threshold, store the data of the corresponding updated state space into the fourth delay replay pool;

[0025] Compare the load balance degree with each preset classification threshold of the balance degree replay pool. When the load balance degree is less than the first balance degree threshold, store the data of the corresponding updated state space into the first balance degree replay pool; when the load balance degree is between the first balance degree threshold and the second balance degree threshold, store the data of the corresponding updated state space into the second balance degree replay pool; when the load balance degree is between the second balance degree threshold and the third balance degree threshold, store the data of the corresponding updated state space into the third balance degree replay pool; when the load balance degree is greater than the third balance degree threshold, store the corresponding data into the fourth balance degree replay pool;

[0026] Compare the delay improvement state with each preset classification threshold of the delay improvement replay pool. When the delay improvement state is less than the first delay improvement threshold, store the data of the corresponding updated state space into the first delay improvement replay pool; when the delay improvement state is between the first delay improvement threshold and the second delay improvement threshold, store the data of the corresponding updated state space into the second delay improvement replay pool; when the delay improvement state is greater than the second delay improvement threshold, store the data of the corresponding updated state space into the third delay improvement replay pool;

[0027] Compare the state of balance improvement with each preset classification threshold in the balance improvement playback pool. When the state of balance improvement is less than the first balance improvement threshold, store the data in the corresponding updated state space into the first balance improvement playback pool; when the state of balance improvement is between the first balance improvement threshold and the second balance improvement threshold, store the data in the corresponding updated state space into the second balance improvement playback pool; when the state of balance improvement is greater than the second balance improvement threshold, store the data in the corresponding updated state space into the third balance improvement playback pool.

[0028] Further, calculate the queuing delay through the following formula:

[0029]

[0030] where

[0031]

[0032] In the formula, is the queuing delay of the data of distribution terminal m at the t-th time slot, is the queuing delay of the data of distribution terminal m at the t-th time slot in distribution terminal m, is the queuing delay of the data of distribution terminal m at the t-th time slot in the first edge server, is the queuing delay of the data of distribution terminal m at the t-th time slot in the n-th edge server, is the queuing delay of the data of distribution terminal m at the t-th time slot in the N-th edge server, is the queuing delay of the data of distribution terminal m at the t-th time slot in the cloud server, is the average data arrival rate of distribution terminal m, is the average data arrival rate of distribution terminal m in edge server n, is the average data arrival rate of distribution terminal m in the cloud server, is the forwarding delay for the data of distribution terminal m to be transmitted from edge server n to the cloud server, R n (t) is the data transmission rate for edge server n to transmit data to the cloud server.

[0033] Further, calculate the load balancing degree through the following formula:

[0034] η(t) = η dec (t) + ω edg η edg (t);

[0035] where

[0036]

[0037] In the formula, η(t) is the load balancing degree of the t-th time slot, and η dec (t) is the load balancing degree of the distribution terminal at the t-th time slot, and ω edg is the average load balancing degree weight of the edge server, and η edg (t) is the load balancing degree of the edge server at the t-th time slot.

[0038] Another embodiment of the present invention provides an edge cluster adaptive access device for distribution automation terminals based on cloud-edge integration, including:

[0039] A data acquisition module, a Markov decision module, a data classification and storage module, a parameter update module, and a joint decision determination module:

[0040] The data acquisition module is used to acquire the data queue and the average data arrival rate of each distribution device;

[0041] The Markov decision module is used to determine the state space, action space, and reward function of the Markov decision based on the data queue and the average data arrival rate of each distribution device;

[0042] The data classification and storage module is used to input the data in the state space into the initial local actor network of each edge server to obtain a joint decision in the action space. After each edge server executes the joint decision, calculate the value of the reward function and update the data in the state space, and classify and store the updated data in the state space into a quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, an equilibrium degree replay pool, a delay improvement replay pool, and an equilibrium degree improvement replay pool;

[0043] The parameter update module is used to determine the exploration weights of the quadruple experience replay pool according to the data classification in the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor network to obtain a predicted action space, and update all local network parameters of each edge server based on the sampled state space data and the predicted action space, and according to the value of the reward function;

[0044] The joint decision determination module is used to upload all the updated local network parameters of each edge server to the cloud server, the cloud server performs global network parameter update, and issues the updated global network parameters to each edge server, and each edge server obtains and executes the joint decision in the action space based on the global network parameters.

[0045] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for adaptive access of the edge cluster of the distribution automation terminal based on cloud-edge integration as described in any one of the embodiments.

[0046] Another embodiment of the present invention provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the method for adaptive access of the edge cluster of the distribution automation terminal based on cloud-edge integration as described in any one of the above embodiments.

[0047] By implementing the present invention, the following beneficial effects are achieved:

[0048] Obtain the data queues and data average arrival rates of the distribution terminals, edge servers, and cloud servers; construct a delay balance degree optimization model, model the delay balance degree optimization model as a Markov decision, and determine the state space, action space, and reward function of the Markov decision; input the data in the state space into the initial local actor networks of each edge server to obtain a joint decision in the action space. After each edge server executes the joint decision, update the data in the state space and classify and store it in a quadruple experience replay pool; determine the exploration weights of the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor networks to obtain the predicted action space, calculate each target Q value, and update all local network parameters of each edge server based on each target Q value; upload all the updated local network parameters of each edge server to the cloud server, the cloud server performs global network parameter updates, and sends the updated global network parameters to each edge server. Each edge server obtains and executes the joint decision in the action space based on the global network parameters. By classifying and storing the data in the Markov decision state space in a quadruple experience replay pool, and determining the exploration weights of the quadruple experience replay pool according to the reward function, queuing delay, and load balance degree, and sampling based on the exploration weights, the adjustment of the experience replay pool sampling can be realized, and the experience replay pool corresponding to the weak points of the queuing delay and load balance degree can be dynamically enhanced in a targeted manner to ensure that the distribution automation service can learn and optimize in a targeted manner when facing different performance challenges. Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of the method for adaptive access of the edge cluster of the distribution automation terminal based on cloud-edge integration provided by an embodiment of the present invention.

[0050] Figure 2It is a schematic structural diagram of an edge cluster adaptive access device for a distribution automation terminal based on cloud-edge integration provided by an embodiment of the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0053] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, "a plurality of" means two or more unless otherwise specifically defined.

[0054] As Figure 1 shown, to solve the problem that distribution automation services cannot learn and optimize targeted when facing poor latency and balance, an embodiment of the present invention provides a cloud-edge integration-based distribution automation terminal edge cluster adaptive access method, including the following steps:

[0055] Step S1, obtain the data queues and data average arrival rates of each distribution device; wherein, the distribution devices include: distribution terminals, edge servers, and cloud servers;

[0056] In the present invention, obtain the data queues and data average arrival rates of distribution terminals, edge servers, and cloud servers;

[0057] Specifically, for the data queue and data average arrival rate of the distribution terminal, define as the data queue maintained by distribution terminal m at the t-th time slot, then the data queue maintained by distribution terminal m at the (t + 1)-th time slot can be calculated by the following formula:

[0058]

[0059] Among them,

[0060]

[0061] In the formula, is the data queue maintained by the distribution terminal m in the (t + 1)-th time slot, is the amount of collected data uploaded by the sensor to the distribution terminal m in the t-th time slot, is the amount of data transmitted by the distribution terminal m to the edge cluster layer in the t-th time slot, x m,n (t) is the access status of the distribution terminal, R m,n (t) is the data transmission rate of the distribution terminal m transmitted to the edge server n, B m,n (t) is the channel bandwidth for the distribution terminal m to transmit data to the edge server n in the t-th time slot, g m,n (t) is the channel gain in the t-th time slot, V m is the transmission power of the distribution terminal m, σ0 is the noise power spectral density, σ m,n (t) is the electromagnetic interference power of the channel in the t-th time slot;

[0062] The average data arrival rate of the distribution terminal is calculated through the following formula:

[0063]

[0064] In the formula, is the average data arrival rate of the distribution terminal m, t - 1 is the (t - 1)-th time slot, is the amount of collected data uploaded by the sensor to the distribution terminal m in the i-th time slot;

[0065] For the data queue and average data arrival rate of the edge server, it is defined that is the data queue maintained by the edge server n from the distribution terminal m in the t-th time slot. Then, the data queue maintained by the edge server n from the distribution terminal m in the (t + 1)-th time slot can be calculated through the following formula:

[0066]

[0067] Among them:

[0068]

[0069] In the formula, is the data queue maintained by the edge server n from the distribution terminal m in the (t + 1)-th time slot, is the amount of data of the distribution terminal m received by the edge server n in the t-th time slot, The data volume of distribution terminal m processed by edge server n in the t-th time slot is the proportion of the data of distribution terminal m in the t-th time slot processed by edge server n, and τ is the duration of one time slot is the CPU cycle frequency of edge server n for processing the data of distribution terminal m, λ m is the computational complexity of the data of distribution terminal m

[0070] The average data arrival rate of the edge server is calculated by the following formula

[0071]

[0072] where is the average data arrival rate of distribution terminal m at edge server n is the data volume of distribution terminal m received by edge server n in the i-th time slot

[0073] For the data queue and average data arrival rate of the cloud server, define as the data queue of distribution terminal m maintained by the cloud server in the t-th time slot. Then, the data queue of distribution terminal m maintained by the cloud server in the (t + 1)-th time slot can be calculated by the following formula

[0074]

[0075] where

[0076]

[0077] where is the data queue of distribution terminal m maintained by the cloud server in the (t + 1)-th time slot is the data volume of distribution terminal m received by the cloud server from the edge server in the t-th time slot is the data volume of distribution terminal m processed by the cloud server in the t-th time slot, and N is N edge servers is the proportion of the data of distribution terminal m transmitted from edge server n to the cloud server for processing in the t-th time slot is the CPU cycle frequency of the cloud server for processing the data of distribution terminal m

[0078] The average data arrival rate of the cloud server is calculated by the following formula

[0079]

[0080] where is the average data arrival rate of distribution terminal m at the cloud server is the data volume of distribution terminal m received by the cloud server from the edge server in the i-th time slot

[0081] Step S2: Determine the state space, action space, and reward function of the Markov decision based on the data queues and data average arrival rates of each power distribution device;

[0082] In a preferred embodiment, the determining the state space, action space, and reward function of the Markov decision based on the data queues and data average arrival rates of each power distribution device includes:

[0083] Based on the data queues and data average arrival rates of each power distribution device, with the goal of minimizing the weighted sum of the queuing delay and the load balancing degree, construct a delay balance degree optimization model, and impose comprehensive condition constraints on the delay balance degree optimization model; among them, the comprehensive condition constraints include: power distribution terminal access constraints, data allocation ratio constraints, edge server computing resource allocation constraints, and cloud server computing resource allocation constraints;

[0084] Model the delay balance degree optimization model as a Markov decision, and determine the state space, action space, and reward function of the Markov decision.

[0085] Specifically, the power distribution terminal access constraint is:

[0086] where M is M power distribution terminals and T is T equal-length time slots;

[0087] The data allocation ratio constraint is:

[0088] The edge server computing resource allocation constraint is:

[0089] where is the maximum available CPU cycle frequency of edge server n at the t-th time slot;

[0090] The cloud server computing resource allocation constraint is:

[0091] where is the maximum available CPU cycle frequency of the cloud server at the t-th time slot.

[0092] In a preferred embodiment, the delay balance degree optimization model includes:

[0093]

[0094] where P is the minimum value of the weighted sum of the queuing delay and the load balancing degree, k is the original optimization target weight, is the weighted sum of the queuing delay and the load balancing degree at the t-th time slot.

[0095] In a preferred embodiment, the state space of the Markov decision includes: the data queues and data average arrival rates of each power distribution device; wherein, the power distribution devices include: power distribution terminals, edge servers, and cloud servers;

[0096] The action space includes: the access status of the power distribution terminal, the joint decision of the data processing allocation ratio and the computing resource allocation;

[0097] The reward function is the negative value of the minimum of the weighted sum of the queuing delay and the load balancing degree.

[0098] Specifically, the state space is expressed as:

[0099]

[0100] In the formula, x(t) represents the state space;

[0101] The action space is expressed as:

[0102]

[0103] In the formula, V(t) represents the action space;

[0104] The reward function is -P.

[0105] Step S3: Input the data of the state space into the initial local actor network of each edge server to obtain the joint decision of the action space. After each edge server executes the joint decision, calculate the value of the reward function and update the data of the state space, and classify and store the updated data of the state space into the quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, a balance degree replay pool, a delay improvement replay pool, and a balance degree improvement replay pool;

[0106] In a preferred embodiment, the delay replay pool includes: a first delay replay pool, a second delay replay pool, a third delay replay pool, and a fourth delay replay pool; the balance degree replay pool includes: a first balance degree replay pool, a second balance degree replay pool, a third balance degree replay pool, and a fourth balance degree replay pool; the delay improvement replay pool includes: a first delay improvement replay pool, a second delay improvement replay pool, and a third delay improvement replay pool; the balance degree improvement replay pool includes: a first balance degree improvement replay pool, a second balance degree improvement replay pool, and a third balance degree improvement replay pool;

[0107] The classifying and storing the updated data of the state space into the quadruple experience replay pool includes:

[0108] Based on the updated data of the state space, calculate the corresponding queuing delay and load balancing degree; according to the queuing delay and the load balancing degree, calculate the delay improvement state and the balance degree improvement state;

[0109] Compare the queuing delay with each preset classification threshold of the delay replay pool. When the queuing delay is less than the first delay threshold, store the data of the corresponding updated state space in the first delay replay pool; when the queuing delay is between the first delay threshold and the second delay threshold, store the data of the corresponding updated state space in the second delay replay pool; when the queuing delay is between the second delay threshold and the third delay threshold, store the data of the corresponding updated state space in the third delay replay pool; when the queuing delay is greater than the third delay threshold, store the data of the corresponding updated state space in the fourth delay replay pool;

[0110] Compare the load balancing degree with each preset classification threshold of the balancing degree replay pool. When the load balancing degree is less than the first balancing degree threshold, store the data of the corresponding updated state space in the first balancing degree replay pool; when the load balancing degree is between the first balancing degree threshold and the second balancing degree threshold, store the data of the corresponding updated state space in the second balancing degree replay pool; when the load balancing degree is between the second balancing degree threshold and the third balancing degree threshold, store the data of the corresponding updated state space in the third balancing degree replay pool; when the load balancing degree is greater than the third balancing degree threshold, store the corresponding data in the fourth balancing degree replay pool;

[0111] Compare the delay improvement state with each preset classification threshold of the delay improvement replay pool. When the delay improvement state is less than the first delay improvement threshold, store the data of the corresponding updated state space in the first delay improvement replay pool; when the delay improvement state is between the first delay improvement threshold and the second delay improvement threshold, store the data of the corresponding updated state space in the second delay improvement replay pool; when the delay improvement state is greater than the second delay improvement threshold, store the data of the corresponding updated state space in the third delay improvement replay pool;

[0112] Compare the balancing degree improvement state with each preset classification threshold of the balancing degree improvement replay pool. When the balancing degree improvement state is less than the first balancing degree improvement threshold, store the data of the corresponding updated state space in the first balancing degree improvement replay pool; when the balancing degree improvement state is between the first balancing degree improvement threshold and the second balancing degree improvement threshold, store the data of the corresponding updated state space in the second balancing degree improvement replay pool; when the balancing degree improvement state is greater than the second balancing degree improvement threshold, store the data of the corresponding updated state space in the third balancing degree improvement replay pool.

[0113] Specifically, the time delays in the first time-delay playback pool, the second time-delay playback pool, the third time-delay playback pool, and the fourth time-delay playback pool increase in sequence; the balance degrees in the first balance-degree playback pool, the second balance-degree playback pool, the third balance-degree playback pool, and the fourth balance-degree playback pool increase in sequence; the time-delay performances reflected by the time-delay improvement states in the first time-delay improvement playback pool, the second time-delay improvement playback pool, and the third time-delay improvement playback pool decrease in sequence, that is, the first time-delay improvement playback pool has the best time-delay performance; the balance performances reflected by the balance-degree improvement states in the first balance-degree improvement playback pool, the second balance-degree improvement playback pool, and the third balance-degree improvement playback pool decrease in sequence, that is, the first balance-degree improvement playback pool has the best balance performance.

[0114] In the present invention, according to Little's law, the queuing delay is defined as the ratio of the queue backlog to the average data arrival rate. The data of the distribution terminal m is processed in the edge server and the cloud server respectively at a certain distribution ratio simultaneously. Therefore, the total queuing delay of the data of the distribution terminal m depends on the maximum value of the queuing delays of the edge layer and the cloud layer. In a preferred embodiment, the following formula is used to calculate the queuing delay:

[0115]

[0116] Wherein,

[0117]

[0118] In the formula, is the queuing delay of the data of the distribution terminal m at the t-th time slot, is the queuing delay of the data of the distribution terminal m at the distribution terminal m at the t-th time slot, is the queuing delay of the data of the distribution terminal m at the first edge server at the t-th time slot, is the queuing delay of the data of the distribution terminal m at the n-th edge server at the t-th time slot, is the queuing delay of the data of the distribution terminal m at the N-th edge server at the t-th time slot, is the queuing delay of the data of the distribution terminal m in the cloud server at the t-th time slot, is the forwarding delay for the data of the distribution terminal m to be transmitted from the edge server n to the cloud server, R n (t) is the data transmission rate for the edge server n to transmit data to the cloud server.

[0119] In the present invention, the standard deviation of the data queue is used to quantify the load balance degree between the distribution terminal and the edge cluster. In a preferred embodiment, the following formula is used to calculate the load balance degree:

[0120] η(t) = η dec (t) + ω edg ηedg (t);

[0121] Among them,

[0122]

[0123] In the formula, η(t) is the load balancing degree of the t-th time slot, η dec (t) is the load balancing degree of the distribution terminal at the t-th time slot, ω edg is the average load balancing degree weight of the edge server, η edg (t) is the load balancing degree of the edge server at the t-th time slot.

[0124] By calculating the delay improvement state, and calculating the balance degree improvement state by η(t) - η(t - 1); in the formula, is the queuing delay of the data of the distribution terminal m at the (t - 1)-th time slot, and η(t - 1) is the load balancing degree at the (t - 1)-th time slot;

[0125] Schematically, the edge server n initializes the local actor network ψ n , the local target actor network two local critic networks θ n,1 and θ n,2 as well as two local target critic networks and Input the current system state X(t) of the state space into the initial local actor network ψ n of the edge server n, and output all possible continuous action probability distributions ψ n (χ(t)) that satisfy the constraints. After considering the exploration noise ε, an action is obtained, denoted as V n (t) ~ ψ n (X(t)) + ε; obtain the joint decision of the action space. After the edge server n executes the joint decision, calculate the value of the reward function and update the data of the state space to χ(t + 1), and classify and store the updated data of the state space into the quadruple experience replay pool.

[0126] Step S4, according to the data classification situation in the quadruple experience replay pool, determine the exploration weight of the quadruple experience replay pool, sample based on the exploration weight, input the sampled state space data into the corresponding local target actor network, obtain the predicted action space, and update all local network parameters of each edge server based on the sampled state space data and the predicted action space and according to the value of the reward function;

[0127] In the present invention, according to the data classification in the quadruple experience replay pool, it is judged whether the queuing delay and load balancing degree of the power distribution system meet the preset requirements. When the classification meets the requirements, it is considered that the system performance is good, and it is determined that the quadruple experience replay pools all have the same exploration weight, and equiprobable sampling is performed on different situations of the quadruple experience replay pool to ensure the global convergence of the algorithm under the quadruple experience replay pool; when the classification does not meet the requirements, the performance weak points of the queuing delay and load balancing degree are determined, the exploration weight is adjusted, and enhanced processing of learning and exploration is carried out for different states in the quadruple experience replay pool in a targeted manner.

[0128] Specifically, when the system balance performance is poor, by increasing the exploration weight of the corresponding experience replay pool, the learning samples that can promote better balance performance are emphasized, and the experience of the balance improvement replay pool is used to quickly converge, so that the algorithm can more efficiently improve the balance of the overall performance. When the system delay performance is poor, by increasing the exploration weight of the corresponding experience replay pool, the learning of the learning samples that can promote better delay performance is enhanced, and the experience of the delay improvement replay pool is used to quickly converge, so that the algorithm can more effectively reduce the processing delay, optimize the communication efficiency, and improve the response speed. The corresponding exploration weight is adjusted through the following formula:

[0129]

[0130] In the formula, ρ Δη+ (t + 1) is the exploration weight of the first balance improvement replay pool at the (t + 1)-th time slot, α η is the balance exploration weight adjustment parameter, ρ Δη+ (t) is the exploration weight of the first balance improvement replay pool at the t-th time slot, η max is the maximum threshold of the load balancing degree, β η is the historical load balancing degree weight, η(i) is the load balancing degree at the i-th time slot, ρ Δτ- (t + 1) is the exploration weight of the first delay improvement replay pool at the (t + 1)-th time slot, α τ is the queuing delay exploration weight adjustment parameter, ρ Δτ- (t) is the exploration weight of the first delay improvement replay pool at the t-th time slot, τ max is the maximum threshold of the queuing delay, β τ is the historical delay weight, τ que (i) is the queuing delay of the data of the power distribution terminal m at the i-th time slot.

[0131] The sampled state space data is input into the corresponding local target actor network to obtain the predicted action space. Based on the sampled state space data and the predicted action space, and according to the value of the reward function, each target Q value is obtained. Based on each target Q value, all local network parameters of each edge server are updated;

[0132] Specifically, taking the sampling data of the first balance improvement replay pool as an example, the sampling state data χ(t + 1) is input into the local target actor network and noise is added to obtain the predicted action, expressed as:

[0133] where represents the predicted action;

[0134] χ(t + 1) and are input into the local target critic network and to obtain the Q-values of and and According to the Q-values of and the target Q-value is obtained based on the value of the reward function and is expressed by the following formula:

[0135]

[0136] In the formula, G n is the target Q-value, reward n is the value of the reward function of edge server n, and λ is the discount factor.

[0137] Edge server n updates the local network parameters and updates the parameters β n,1 and β n,2 of the two local critic networks θ n,1 (t) and β n,2 (t) using the mini-batch gradient descent method; it is expressed as:

[0138]

[0139] In the formula, β n,i (t + 1) is the parameter of the i-th local critic network at the (t + 1)-th time slot, Z is the number of empirical samples, is the predicted action at the t-th time slot, and X(t) is the state data at the t-th time slot;

[0140] The parameters α n of the local actor network ψ n are updated using the deterministic policy gradient method, expressed as:

[0141]

[0142] In the formula, α n (t + 1) is the parameter of the local actor network ψ n at the (t + 1)-th time slot, V​n (t) is the state value function;

[0143] For the local target critic network and and the local target actor network Update the parameters, which is expressed as:

[0144]

[0145] In the formula, is the parameter of the local target critic network at the (t + 1)-th time slot , μ is the soft update coefficient, β n,1 (t) is the parameter of the local critic network θ at the t-th time slot n,1 , is the parameter of the local target critic network at the t-th time slot , is the parameter of the local target critic network at the (t + 1)-th time slot , β n,2 (t) is the parameter of the local critic network θ at the t-th time slot n,2 , is the parameter of the local target critic network at the t-th time slot , is the parameter of the local target actor network at the (t + 1)-th time slot , α n (t) is the parameter of the local actor network ψ at the t-th time slot n , is the parameter of the local target actor network at the t-th time slot .

[0146] Step S5: Upload all the updated local network parameters of each edge server to the cloud server. The cloud server updates the global network parameters and distributes the updated global network parameters to each edge server. Each edge server obtains and executes the joint decision of the action space based on the global network parameters.

[0147] Specifically, each edge server uploads the updated local actor network ψ n , the local target actor network the two local critic networks θ n,1 and θ n,2 , the two local target critic networks and Upload it to the cloud server. The cloud server performs federated aggregation to update the global network parameters, and distributes the updated global network parameters to each edge server. Each edge server obtains the global network parameters as the new local network parameters and performs a joint decision in the action space. The joint decision made by the edge server based on the global information can improve the operating efficiency of the power distribution system.

[0148] As Figure 2 shown, it is a schematic structural diagram of an edge cluster adaptive access device for a power distribution automation terminal based on cloud-edge fusion provided by an embodiment of the present invention, including:

[0149] A data acquisition module, a Markov decision module, a data classification and storage module, a parameter update module, and a joint decision determination module:

[0150] The data acquisition module is used to acquire the data queues and data average arrival rates of each power distribution device;

[0151] The Markov decision module is used to determine the state space, action space, and reward function of the Markov decision based on the data queues and data average arrival rates of each power distribution device;

[0152] The data classification and storage module is used to input the data in the state space into the initial local actor network of each edge server to obtain a joint decision in the action space. After each edge server executes the joint decision, it calculates the value of the reward function and updates the data in the state space, and classifies and stores the updated data in the state space into a quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, an equilibrium degree replay pool, a delay improvement replay pool, and an equilibrium degree improvement replay pool;

[0153] The parameter update module is used to determine the exploration weights of the quadruple experience replay pool according to the data classification situation in the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor network to obtain the predicted action space, and update all local network parameters of each edge server based on the sampled state space data and the predicted action space and according to the value of the reward function;

[0154] The joint decision determination module is used to upload all the updated local network parameters of each edge server to the cloud server. The cloud server updates the global network parameters and distributes the updated global network parameters to each edge server. Each edge server obtains and executes the joint decision in the action space based on the global network parameters.

[0155] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention, and can implement the cloud-edge fusion-based adaptive access method for the edge cluster of the distribution automation terminal provided by any one of the above method item embodiments of the present invention.

[0156] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.

[0157] Those skilled in the art can clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0158] Another preferred embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the cloud-edge fusion-based adaptive access method for the edge cluster of the distribution automation terminal as described in any one of the above embodiments.

[0159] It should be noted that the terminal device mentioned here can be computing devices such as desktop computers, notebooks, palm computers, and cloud servers. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that, for example, it may also include input / output devices, network access devices, buses, etc.

[0160] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device through various interfaces and lines.

[0161] The memory can be used to store the computer program. The processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0162] Another preferred embodiment of the present invention provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the edge cluster adaptive access method for a distribution automation terminal based on cloud-edge integration according to any one of the present invention.

[0163] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0164] The foregoing is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications are also regarded as the protection scope of the present invention.

Claims

1. An edge cluster adaptive access method for distribution automation terminals based on cloud-edge integration, characterized in that Including: Obtain the data queues and data average arrival rates of each power distribution device; Based on the data queues and data average arrival rates of each power distribution device, determine the state space, action space, and reward function of the Markov decision-making; Input the data in the state space into the initial local actor network of each edge server to obtain the joint decision in the action space. After each edge server executes the joint decision, calculate the value of the reward function and update the data in the state space, and classify and store the updated data in the state space into the quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, a balance degree replay pool, a delay improvement replay pool, and a balance degree improvement replay pool; According to the data classification in the quadruple experience replay pool, determine the exploration weights of the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor network to obtain the predicted action space, and update all local network parameters of each edge server based on the sampled state space data and the predicted action space and according to the value of the reward function; Upload all the updated local network parameters of each edge server to the cloud server. The cloud server performs global network parameter updates, and distributes the updated global network parameters to each edge server. Each edge server obtains and executes the joint decision in the action space based on the global network parameters.

2. The edge cluster adaptive access method for the distribution automation terminal based on cloud-edge integration according to claim 1, wherein The determining the state space, action space, and reward function of the Markov decision-making based on the data queues and data average arrival rates of each power distribution device includes: Based on the data queues and data average arrival rates of each power distribution device, with the goal of minimizing the weighted sum of the queuing delay and the load balance degree, construct a delay balance degree optimization model and impose comprehensive condition constraints on the delay balance degree optimization model; wherein, the comprehensive condition constraints include: power distribution terminal access constraints, data distribution ratio constraints, edge server computing resource allocation constraints, and cloud server computing resource allocation constraints; Model the delay balance degree optimization model as a Markov decision-making to determine the state space, action space, and reward function of the Markov decision-making.

3. The edge cluster adaptive access method for the distribution automation terminal based on cloud-edge integration according to claim 2, wherein, The delay balance degree optimization model includes: where P is the minimum value of the weighted sum of the queuing delay and the load balancing degree, and x m,n (t) is the access status of the distribution terminal, is the proportion of the data of distribution terminal m in the t-th time slot processed by edge server n, is the proportion of the data of distribution terminal m in the t-th time slot transmitted from edge server n to the cloud server for processing, is the CPU cycle frequency of edge server n for processing the data of distribution terminal m, is the CPU cycle frequency of the cloud server for processing the data of distribution terminal m, and κ is the weight of the original optimization objective, is the weighted sum of the queuing delay and the load balancing degree in the t-th time slot, M is M distribution terminals, is the data queue maintained by distribution terminal m in the t-th time slot, is the amount of acquisition data uploaded from the acquisition layer to distribution terminal m in the t-th time slot, is the amount of data transmitted by distribution terminal m to the edge cluster layer in the t-th time slot, N is N edge servers, is the data queue maintained by edge server n from distribution terminal m in the t-th time slot, is the amount of data of distribution terminal m received by edge server n in the t-th time slot, is the amount of data of distribution terminal m processed by edge server n in the t-th time slot, is the data queue of distribution terminal m maintained by the cloud server in the t-th time slot, is the amount of data of distribution terminal m received by the cloud server after being forwarded by the edge server in the t-th time slot, is the amount of data of distribution terminal m processed by the cloud server in the t-th time slot.

4. The edge cluster adaptive access method for the distribution automation terminal based on cloud-edge integration according to claim 3, wherein, The state space of the Markov decision-making includes: the data queues and data average arrival rates of each power distribution device; wherein, the power distribution devices include: power distribution terminals, edge servers, and cloud servers; The action space includes: the joint decision of the access state of the power distribution terminal, the data processing distribution ratio, and the computing resource allocation; The reward function is the negative value of the minimum of the weighted sum of the queuing delay and the load balance degree.

5. The edge cluster adaptive access method for the distribution automation terminal based on cloud-edge integration according to claim 4, characterized in that, The delay replay pool includes: a first delay replay pool, a second delay replay pool, a third delay replay pool, and a fourth delay replay pool; the balance degree replay pool includes: a first balance degree replay pool, a second balance degree replay pool, a third balance degree replay pool, and a fourth balance degree replay pool; the delay improvement replay pool includes: a first delay improvement replay pool, a second delay improvement replay pool, and a third delay improvement replay pool; the balance degree improvement replay pool includes: a first balance degree improvement replay pool, a second balance degree improvement replay pool, and a third balance degree improvement replay pool; Classifying and storing the data in the updated state space into a quadruple experience replay pool includes: Based on the data in the updated state space, calculating the corresponding queuing delay and load balancing degree; calculating the delay improvement state and the balancing degree improvement state according to the queuing delay and the load balancing degree; Comparing the queuing delay with each preset classification threshold of the delay replay pool. When the queuing delay is less than the first delay threshold, storing the data in the corresponding updated state space into the first delay replay pool; when the queuing delay is between the first delay threshold and the second delay threshold, storing the data in the corresponding updated state space into the second delay replay pool; when the queuing delay is between the second delay threshold and the third delay threshold, storing the data in the corresponding updated state space into the third delay replay pool; when the queuing delay is greater than the third delay threshold, storing the data in the corresponding updated state space into the fourth delay replay pool; Comparing the load balancing degree with each preset classification threshold of the balancing degree replay pool. When the load balancing degree is less than the first balancing degree threshold, storing the data in the corresponding updated state space into the first balancing degree replay pool; when the load balancing degree is between the first balancing degree threshold and the second balancing degree threshold, storing the data in the corresponding updated state space into the second balancing degree replay pool; when the load balancing degree is between the second balancing degree threshold and the third balancing degree threshold, storing the data in the corresponding updated state space into the third balancing degree replay pool; when the load balancing degree is greater than the third balancing degree threshold, storing the corresponding data into the fourth balancing degree replay pool; Comparing the delay improvement state with each preset classification threshold of the delay improvement replay pool. When the delay improvement state is less than the first delay improvement threshold, storing the data in the corresponding updated state space into the first delay improvement replay pool; when the delay improvement state is between the first delay improvement threshold and the second delay improvement threshold, storing the data in the corresponding updated state space into the second delay improvement replay pool; when the delay improvement state is greater than the second delay improvement threshold, storing the data in the corresponding updated state space into the third delay improvement replay pool; Comparing the balancing degree improvement state with each preset classification threshold of the balancing degree improvement replay pool. When the balancing degree improvement state is less than the first balancing degree improvement threshold, storing the data in the corresponding updated state space into the first balancing degree improvement replay pool; when the balancing degree improvement state is between the first balancing degree improvement threshold and the second balancing degree improvement threshold, storing the data in the corresponding updated state space into the second balancing degree improvement replay pool; when the balancing degree improvement state is greater than the second balancing degree improvement threshold, storing the data in the corresponding updated state space into the third balancing degree improvement replay pool.

6. The edge cluster adaptive access method for the distribution automation terminal based on cloud-edge integration according to claim 5, wherein Calculating the queuing delay through the following formula: Among them, Wherein, is the queuing delay of the data of the distribution terminal m in the t-th time slot, is the queuing delay of the data of the distribution terminal m in the distribution terminal m in the t-th time slot, is the queuing delay of the data of the distribution terminal m in the first edge server in the t-th time slot, is the queuing delay of the data of the distribution terminal m in the n-th edge server in the t-th time slot, is the queuing delay of the data of the distribution terminal m in the N-th edge server in the t-th time slot, is the queuing delay of the data of the distribution terminal m in the cloud server in the t-th time slot, is the average data arrival rate of the distribution terminal m, is the average data arrival rate of the distribution terminal m in the edge server n, is the average data arrival rate of the distribution terminal m in the cloud server, is the forwarding delay of the data of the distribution terminal m transmitted from the edge server n to the cloud server, R n (t) is the data transmission rate of the edge server n transmitted to the cloud server.

7. The edge cluster adaptive access method for a distribution automation terminal based on cloud-edge integration according to claim 6, wherein Calculating the load balancing degree through the following formula: η(t) = η dec (t) + ω edg η edg (t); Among them, where η(t) is the load balancing degree at the t-th time slot, and η dec (t) is the load balancing degree of the distribution terminal at the t-th time slot, ω edg is the average load balancing degree weight of the edge server, and η edg (t) is the load balancing degree of the edge server at the t-th time slot.

8. The edge cluster adaptive access device for distribution automation terminals based on cloud-edge integration is characterized in that, Including: A data acquisition module, a Markov decision-making module, a data classification and storage module, a parameter update module, and a joint decision-making determination module: The data acquisition module is used to acquire the data queue and the data average arrival rate of each power distribution device; The Markov decision-making module is used to determine the state space, action space, and reward function of the Markov decision based on the data queues and data average arrival rates of each power distribution device; The data classification and storage module is used to input the data in the state space into the initial local actor networks of each edge server to obtain a joint decision in the action space. After each edge server executes the joint decision, it calculates the value of the reward function and updates the data in the state space, and classifies and stores the updated data in the state space into a quadruple experience replay pool; wherein, the quadruple experience replay pool includes: a delay replay pool, a balance degree replay pool, a delay improvement replay pool, and a balance degree improvement replay pool; The parameter update module is used to determine the exploration weights of the quadruple experience replay pool according to the data classification in the quadruple experience replay pool, sample based on the exploration weights, input the sampled state space data into the corresponding local target actor networks to obtain the predicted action space, and update all local network parameters of each edge server based on the sampled state space data and the predicted action space, and according to the value of the reward function; The joint decision determination module is used to upload all the updated local network parameters of each edge server to the cloud server. The cloud server performs global network parameter updates and distributes the updated global network parameters to each edge server. Each edge server obtains and executes the joint decision in the action space based on the global network parameters.

9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the cloud-edge fusion-based power distribution automation terminal edge cluster adaptive access method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute the cloud-edge fusion-based power distribution automation terminal edge cluster adaptive access method according to any one of claims 1 to 7.