Privacy-Enhanced Cloud-Edge Collaborative Computing Method and System Based on Speculative Pricing Mechanism
The bidirectional reinforcement learning method using predicted pricing optimizes resource allocation in cloud-edge collaborative computing, addressing privacy and complexity issues by simulating resource scheduling and reducing computational complexity.
Patent Information
- Application Number
- CN202411789953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing cloud-edge collaborative computing methods have problems such as data re-identification risks, complex computing processes and difficult configuration in terms of privacy protection, making it difficult to take into account privacy and efficient scheduling in a dynamic environment.
A two-end reinforcement learning method based on the speculative price mechanism is adopted to collect and simulate cloud edge collaboratively calculate network parameters, generate speculative prices, and make resource scheduling decisions to achieve privacy enhancement.
It improves the accuracy and efficiency of privacy protection in cloud-edge collaborative computing network, generates a resource purchase strategy with the best service quality, and adapts to dynamic environmental changes.
Smart Images

Figure CN119668874B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of resource scheduling in collaborative computing networks, and particularly to a cloud-edge collaborative computing privacy enhancement method and system based on a speculative price mechanism. Background Art
[0002] As an emerging computing mode, cloud-edge collaborative computing realizes the efficient utilization of resources and the improvement of service quality by integrating the high performance of cloud computing and the low latency characteristics of edge computing. In the cloud-edge collaborative computing environment, resource scheduling has become one of the research hotspots. Existing research mainly focuses on three aspects: computation offloading, resource allocation, and resource provisioning. Computation offloading reduces the device burden and improves efficiency by offloading terminal tasks to the edge or the cloud; resource allocation focuses on the efficient allocation of resources such as computing, storage, and communication; resource provisioning responds to real-time workloads by dynamically adjusting the resource location. However, existing methods mostly emphasize system performance and often neglect privacy protection. In addition, the problem of high computational complexity also makes it difficult for these methods to balance privacy and efficient scheduling in a dynamic cloud-edge environment, which to a certain extent limits the practicality and popularization.
[0003] Privacy enhancement technology has become one of the key means to ensure the security of cloud-edge collaborative computing by preventing the leakage of sensitive information during data transmission and calculation. Among them, common technical methods include data encryption, differential privacy, and split learning. However, in actual collaborative networks, there are still problems such as the risk of data re-identification, complex calculation processes, and difficult configuration. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the present invention provides a cloud-edge collaborative computing privacy enhancement method and system based on a speculative price mechanism. By synthesizing multi-dimensional network sensitive information into a speculative price through double-ended reinforcement learning, and then obtaining the optimal resource purchase decision for service time and cost, it solves the problems of data re-identification risk, complex calculation process, and difficult configuration that still exist in actual collaborative networks.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows: In the first aspect, a cloud-edge collaborative computing privacy enhancement method based on a speculative price mechanism is provided, including:
[0006] S1. Collect the cloud server data and edge server data of the cloud-edge collaborative computing network whose network privacy needs to be enhanced as the cloud-edge collaborative computing network parameter set;
[0007] S2. Perform action state modeling based on the cloud-edge collaborative computing network parameter set, and generate a cloud server state set and an edge server state set respectively;
[0008] S3. Based on the Q-learning method, simulate the resource scheduling process of cloud-edge collaborative computing, and use the cloud server state set and the edge server state set for double-ended reinforcement learning training to obtain a first type of Q-table and a second type of Q-table;
[0009] S4. Complete privacy-enhanced cloud-edge collaborative computing according to the first type of Q-table and the second type of Q-table.
[0010] Furthermore: S1 includes:
[0011] S11. Collect the unit time usage price of all types of computing resources in the network , frequency f and supply distribution ;
[0012] S12. Collect the total task computing demand distribution received by each edge server per unit time D , the purchase range and over-purchase range of all types of computing resources;
[0013] S13. Determine the distribution range of the speculated price according to the unit time usage price ; ;
[0014] S14. Use the unit time usage price , frequency f , supply distribution , demand distribution D , the purchase range of all types of computing resources , over-purchase range and the distribution range of the price together as the cloud-edge collaborative computing network parameter set.
[0015] Furthermore: In S1, the unit time usage price of all types of computing resources includes the unit time usage price of the general computing resources of the cloud server , the unit time usage price of the dynamic computing resources of the cloud server and the unit time usage price of the local computing resources of each edge server ; ; where max{.} represents the operation of taking the maximum value;
[0016] The frequency of all types of computing resources in the network f includes the computing frequency of the general resources of the cloud server , the computing frequency of the dynamic resources of the cloud server , the computing frequency of the local resources of each edge server ;
[0017] The supply distribution of all types of computing resources in the network , where represents the possible maximum supply of the dynamic resources of the cloud server per unit time.
[0018] Furthermore: In S12, the distribution of the total task computing demand received by each edge server per unit time , where represents the total task computing demand that a certain edge server may receive per unit time;
[0019] The purchase range of all types of computing resources includes the purchasable range of the general computing resources of the cloud server by the edge server and the purchasable range of the dynamic computing resources of the cloud server ;
[0020] The over-purchase range of all types of computing resources is the over-purchase range of the general computing resources of the cloud server by the edge server .
[0021] Furthermore, in S2, there is a corresponding action set for the cloud servers in the cloud server state set ; where represents the space of the speculated value of the resource purchase price of the cloud server for a certain edge server; when the cloud server is in the state , the action taken represents the speculated value of the resource purchase price of the cloud server for a certain edge server in the current round;
[0022] There is a corresponding action set for the edge servers in the edge server state set , where represents the space of the general resource quantity and dynamic resource quantity that a certain edge server can purchase; when the edge server is in the state , the action taken represents the general resource quantity and dynamic resource quantity that the edge server attempts to purchase.
[0023] Furthermore: The cloud server state includes the real-time maximum supply of the dynamic computing resources and the transaction information with a certain edge server in the previous round;
[0024] The edge server state is a triple, including the speculated price informed by the cloud server, the total computing demand for tasks received at the current moment, and the general resource quantity that is currently over-purchased and not yet used.
[0025] Further: The resource scheduling process of cloud-edge collaborative computing in S3 includes:
[0026] S301. Determine the resource purchase volume of the cloud server based on the initial task status of the edge server;
[0027] S302. Use the cloud server to allocate resources according to the resource purchase volume of each edge server;
[0028] S303. Based on the resource volume allocated by the cloud server, the edge server distributes tasks to the general resources of the cloud server, the dynamic resources of the cloud server, and the local resources of the edge server for calculation.
[0029] Further: The action process of the Q-learning method in S3 includes:
[0030] S311. Use a Q-table to calculate and store the Q-values of each state-action pair;
[0031] S312. Gradually converge the policy by iteratively updating the Q-values;
[0032] S313. Search the iterated Q-table and select the action corresponding to the maximum Q-value according to the current state.
[0033] Further: In S3, the steps of performing double-ended reinforcement learning training include:
[0034] S321. Initialize the experience pool of the cloud server , the experience pool of the edge server , the Q-table of the cloud server and the Q-table of the edge server;
[0035] S322. Simulate the resource scheduling process of cloud-edge collaborative computing, store the cloud server samples generated in each round into the cloud server experience pool , and store the edge server samples generated in each round into the edge server experience pool ;
[0036] S323. Randomly select from the cloud server experience pool , the edge server experience pool minibatch a cloud server sample and an edge server sample;
[0037] Among them, the cloud server sample and the edge server sample come from the same simulation round;
[0038] S324. Calculate the target value of each state in the randomly selected cloud server sample and edge server sample respectively;
[0039] S325. Update the Q values in the Q tables of the cloud server and the edge server according to the target value of each state, and use them as the first type of Q table and the second type of Q table.
[0040] In a second aspect, the present application provides a cloud-edge collaborative computing privacy enhancement system based on a speculative price mechanism, which performs privacy enhancement according to the cloud-edge collaborative computing privacy enhancement method based on the speculative price mechanism described in the first aspect.
[0041] The beneficial effects of the present invention are as follows:
[0042] 1. The present invention conducts research in a cloud-edge collaborative computing network, and verifies the feasibility and applicability of the solution in the scenario by simulating the resource scheduling process of cloud-edge collaborative computing.
[0043] 2. Use the double-ended reinforcement learning method for training to generate speculative prices more accurately and quickly, and obtain the optimal resource purchase strategy for service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. is a flowchart of an implementation of a cloud-edge collaborative computing privacy enhancement method based on a speculative price mechanism.
[0045] Figure 2 FIG. is a schematic diagram of a privacy-enhanced cloud-edge collaborative computing simulation.
[0046] Figure 3 FIG. is a schematic diagram of the Q learning and its action process.
[0047] Figure 4 FIG. is a flowchart of double-ended reinforcement learning training.
[0048] Figure 5 FIG. is another flowchart of an implementation of a cloud-edge collaborative computing privacy enhancement method based on a speculative price mechanism.
[0049] Figure 6 FIG. is an experimental parameter diagram of an embodiment.
[0050] Figure 7 FIG. is a graph showing the relationship between the number of training cycles and the normalized service time and service cost in an embodiment.
[0051] Figure 8 FIG. is a comparison graph of the speculative price and the actual price in an embodiment.
[0052] Figure 9 FIG. is a comparison graph of an embodiment with a random policy and a general resource single-purchase policy.
[0053] Figure 10 FIG. is a diagram showing the action of the method for detecting distribution drift based on hypothesis testing in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0055] Embodiment 1:
[0056] As Figure 1 shown, in an embodiment of the present invention, a cloud-edge collaborative computing privacy enhancement method based on a speculative price mechanism is provided, including:
[0057] S1. Collect the cloud server data and edge server data of the cloud-edge collaborative computing network whose network privacy needs to be enhanced as the cloud-edge collaborative computing network parameter set;
[0058] S2. Perform action state modeling based on the cloud-edge collaborative computing network parameter set, and respectively generate the cloud server state set and the edge server state set ;
[0059] S3. Based on the Q-learning method, simulate the resource scheduling process of cloud-edge collaborative computing, and use the cloud server state set and the edge server state set to perform double-ended reinforcement learning training to obtain the first type of Q-table and the second type of Q-table;
[0060] S4. Complete the privacy-enhanced cloud-edge collaborative computing according to the first type of Q-table and the second type of Q-table. The specific method can be summarized as: input the initial states into the first type of Q-table respectively, select the cloud server action (i.e., the speculative price of resources) corresponding to the maximum Q value, so as to construct the initial task states of the edge servers; input the initial states of the edge servers into the second type of Q-table to obtain the actions of the edge servers (i.e., the resource purchase quantities), and the edge servers complete resource transactions and calculations according to their respective purchase quantities.
[0061] Specifically, S1 includes:
[0062] S11. Collect the unit time usage price , frequency f and supply distribution of all types of computing resources in the network;
[0063] S12. Collect the distribution of the total task computing demand received by each edge server per unit time D , the purchase range and over-purchase range of all types of computing resources;
[0064] S13. Determine the distribution range of the speculated price according to the per-unit-time usage price ; ;
[0065] S14. Take the per-unit-time usage price , frequency f , supply distribution , demand distribution D , the purchase range of all types of computing resources , over-purchase range and the distribution range of the price together as the cloud-edge collaborative computing network parameter set.
[0066] Specifically, in S1, the per-unit-time usage price of all types of computing resources includes the per-unit-time usage price of the general computing resources of the cloud server , the per-unit-time usage price of the dynamic computing resources of the cloud server and the per-unit-time usage price of the local computing resources of each edge server ;
[0067] ; where max{.} represents the operation of taking the maximum value;
[0068] The frequency of all types of computing resources in the network f includes the computing frequency of the general resources of the cloud server , the computing frequency of the dynamic resources of the cloud server , the computing frequency of the local resources of each edge server ;
[0069] The supply distribution of all types of computing resources in the network , where represents the possible maximum supply of the dynamic resources of the cloud server per unit time, and there is no upper limit on the maximum supply of the general resources of the cloud server and the local resources of the edge server
[0070] Specifically, in S12, the distribution of the total task computing demand received by each edge server per unit time , where represents the total task computing demand that a certain edge server may receive per unit time;
[0071] The purchase range of all types of computing resources includes the purchasable range of the general computing resources of the cloud server by the edge server and the purchasable range of the dynamic computing resources of the cloud server ; there is no upper limit on the purchase of local resources;
[0072] The over-purchase range of all types of computing resources is the over-purchase range of the general computing resources of the edge server for the cloud server , (that is, the amount of resources that exceeds the actual usage in the current unit time and allows reservation), and over-purchase of the dynamic resources of the cloud server and the local resources of the edge server is not allowed.
[0073] Specifically, in S2, there is a corresponding action set for the cloud servers in the cloud server status set ; among them, represents the estimated value space of the purchase price of the resources of a certain edge server by the cloud server; when the cloud server is in the state , the action taken represents the estimated value of the purchase price of the resources of a certain edge server by the cloud server in the current round, that is , the closer the estimated price is to the actual price after the edge server makes a purchase decision, the more rewards will be obtained; otherwise, penalties will be obtained;
[0074] There is a corresponding action set for the edge servers in the edge server status set , among them, represents the space of the general resource quantity and the dynamic resource quantity that a certain edge server can purchase; when the edge server is in the state , the action taken represents the general resource quantity and the dynamic resource quantity that the edge server attempts to purchase, that is , when the two types of purchase quantities are determined, the lower the final transaction cost and the shorter the unit task time, the more rewards will be obtained; otherwise, penalties will be obtained; in the formula, , .
[0075] The cloud server status includes the real-time maximum supply of dynamic computing resources and the transaction information of the previous round with a certain edge server, that is:
[0076] ;
[0077] The edge server status is a triple, including the estimated price informed by the cloud server, the total computing demand of the received tasks at the current moment, and the general resources that are currently over-purchased and not yet used, that is:
[0078] .
[0079] Specifically, referring to Figure 2 , the resource scheduling process of cloud-edge collaborative computing in S3 includes:
[0080] S301. Determine the cloud server resource purchase quantity based on the initial task status of the edge server;
[0081] The initial task status of the edge server Through the initial task status of the cloud server , inform the initial resource speculation price of each edge server Determine;
[0082] The cloud server resource purchase quantity includes the general resource purchase quantity of the cloud server and the dynamic resource purchase quantity of the cloud server ;
[0083] S302. Use the cloud server to allocate resources according to the resource purchase quantity of each edge server;
[0084] The general resource quantity allocated to each edge server is , and the dynamic resource quantity allocated to each edge server ; where represents the upper limit of the supply of dynamic resources at the current moment
[0085] S303. Based on the resource quantity allocated by the cloud server, the edge server distributes tasks to the general resources of the cloud server, the dynamic resources of the cloud server, and the local resources of the edge server for calculation. The specific calculation formula is:
[0086]
[0087] Among them, respectively represent the task calculation quantities of the edge server allocated to the general resources of the cloud server, the dynamic resources of the cloud server, and the local resources of the edge server, represents the total task calculation demand received by the edge server per unit time at the current moment, represents the general resource quantity reserved by the edge server per unit time at the current moment and not yet used; the actual resource purchase cost of each edge server in this round is , and the average execution time of this round of tasks is , is the average transmission delay of the task, is the average computing resource quantity required for the task, is the actual computing frequency of the dynamic resources allocated to each edge server;
[0088] Optionally, to pursue the cloud-edge collaborative computing accuracy, the calculation round E can be set, and these S301 - S303 are repeatedly executed for E rounds.
[0089] Specifically, referring to Figure 3 , the action process of the Q-learning method in S3 includes:
[0090] S311. Use a Q-table to calculate and store the Q-values for each state-action pair;
[0091] S312. Gradually converge the policy by iteratively updating the Q-values. The expression is:
[0092]
[0093] where, s represents the state, a represents the state s corresponding action, represents the state s next state, a represents the state corresponding action, r represents the reward, ε is the learning rate, γ is the decay coefficient;
[0094] S313. Search the iterated Q-table and select the action corresponding to the maximum Q-value according to the current state. The expression is:
[0095] .
[0096] Specifically, referring to Figure 4 , in S3, the steps for double-ended reinforcement learning training include:
[0097] S321. Initialize the cloud server experience pool , the edge server experience pool , the cloud server Q-table, and the edge server Q-table;
[0098] Among them, the cloud server Q-table is set sheets, and the dimension is set to ; the initialized edge server Q-table is set sheets, and the dimension is set to ; represents the size of the set space;
[0099] S322. Simulate the resource scheduling process of cloud-edge collaborative computing, store the cloud server samples generated in each round to the cloud server experience pool , store the storage edge server samples generated in each round to the edge server experience pool , t is an integer greater than or equal to 1, representing the number of simulation rounds;
[0100] S323. From the cloud server experience pool , the edge server experience pool Randomly select minibatch cloud server samples and edge server samples , j denote the randomly selected samples;
[0101] Among them, the cloud server samples and the edge server samples come from the same simulation round;
[0102] S324. Calculate the objective value of each state in the randomly selected cloud server samples and edge server samples respectively;
[0103] Specifically, update the Q value as the objective value through the reward after executing : If the next state is an absorbing state, then , otherwise: ;
[0104] Among them, y ( j ) is the objective value of the state corresponding to the selected j th sample, and this objective value is related to whether the next state is an absorbing state; r ( j ) is the reward value in the selected j th sample;
[0105] S325. Update the Q values in the cloud server Q table and the edge server Q table according to the objective value of each state, as the first type of Q table and the second type of Q table.
[0106] This embodiment also provides a cloud-edge collaborative computing privacy enhancement system for privacy enhancement according to the cloud-edge collaborative computing privacy enhancement method based on the speculative price mechanism.
[0107] Embodiment 2:
[0108] As Figure 2 shown, this embodiment provides another implementation method of the cloud-edge collaborative computing privacy enhancement method based on the speculative price mechanism, including:
[0109] S1. Collect the cloud server data and edge server data of the cloud-edge collaborative computing network whose network privacy needs to be enhanced, as the cloud-edge collaborative computing network parameter set;
[0110] S2. Perform action state modeling according to the cloud-edge collaborative computing network parameter set, and generate a cloud server state set and an edge server state set respectively;
[0111] S3. Based on the Q-learning method, simulate the resource scheduling process of cloud-edge collaborative computing, and use the cloud server state set and the edge server state set for double-sided reinforcement learning training to obtain the first type of Q-table and the second type of Q-table;
[0112] S4. Complete privacy-enhanced cloud-edge collaborative computing according to the first type of Q-table and the second type of Q-table. The specific method can be summarized as follows: Input initial states into the first type of Q-table respectively, and select the cloud server action (i.e., the speculative price of resources) corresponding to the largest Q-value, so as to construct initial task states of edge servers ; Input initial states of edge servers into the second type of Q-table to obtain actions of edge servers (i.e., the resource purchase volume). The edge servers complete resource transactions and calculations according to their respective purchase volumes;
[0113] S5. Each edge server records and statistics the time and cost of each service in step S4. If there are outliers, report anomalies; if there are no anomalies, report normal; Statistically analyze the reported information of edge servers within the time range and conduct a hypothesis test. If the hypothesis test of resource supply distribution drift is successful, immediately return to S1 for retraining; otherwise, return to S4 for the next round of collaborative computing; is an integer greater than 1.
[0114] Specifically, S1 includes:
[0115] S11. Collect the unit-time usage price , frequency f and supply volume distribution of all types of computing resources in the network;
[0116] S12. Collect the total task calculation demand distribution D received by each edge server per unit time, the purchase range and over-purchase range of all types of computing resources;
[0117] S13. Draw up the distribution range of speculative prices according to the unit-time usage price ;
[0118] S14. The unit-time usage price , frequency f , supply volume distribution , demand distribution D , the purchase range of all types of computing resources, and the over-purchase range and the distribution range of prices Together they serve as the cloud-edge collaborative computing network parameter set.
[0119] Specifically, in S1, the unit-time usage price of all types of computing resources includes the unit-time usage price of the general computing resources of the cloud server , the unit-time usage price of the dynamic computing resources of the cloud server and the unit-time usage price of the local computing resources of each edge server ;
[0120] ; where max{.} represents the operation of taking the maximum value;
[0121] The frequencies of all types of computing resources in the network f include the computing frequency of the general resources of the cloud server , the computing frequency of the dynamic resources of the cloud server , and the computing frequency of the local resources of each edge server ;
[0122] The supply distribution of all types of computing resources in the network , where represents the possible maximum supply of the dynamic resources of the cloud server within a unit time, and there is no upper limit on the maximum supply of the general resources of the cloud server and the local resources of the edge server.
[0123] Specifically, in S12, the distribution of the total task computing demand received by each edge server per unit time , where represents the total task computing demand that a certain edge server may receive per unit time;
[0124] The purchase range of all types of computing resources includes the purchasable range of the general computing resources of the cloud server by the edge server and the purchasable range of the dynamic computing resources of the cloud server , and there is no upper limit on the purchase of local resources;
[0125] The over-purchase range of all types of computing resources is the over-purchase range of the general computing resources of the cloud server by the edge server , (that is, the amount of resources that exceeds the actual usage within the current unit time and is allowed to be reserved), and over-purchase of dynamic resources of the cloud server and local resources of the edge server is not allowed.
[0126] Specifically, in S2, there is a corresponding action set for the cloud servers in the cloud server status set ; where Represents the speculation value space of the cloud server for the resource purchase price of a certain edge server; when the cloud server is in the state the action taken Represents the speculation value of the cloud server for the current round of resource purchase price of a certain edge server, that is , the closer the speculation price is to the actual price after the edge server makes a purchase decision, the more rewards will be obtained; otherwise, penalties will be obtained;
[0127] There is a corresponding action set for the edge server in the edge server state set where Represents the general resource quantity and dynamic resource quantity space that a certain edge server can purchase; when the edge server is in the state the action taken Represents the general resource quantity and dynamic resource quantity that the edge server attempts to purchase, that is , when the two types of purchase quantities are determined, the lower the final transaction cost and the shorter the unit task time, the more rewards will be obtained; otherwise, penalties will be obtained; in the formula, , .
[0128] The cloud server state includes the real-time maximum supply of dynamic computing resources and the transaction information with a certain edge server in the previous round, that is:
[0129] ;
[0130] The edge server state is a triple, including the speculation price informed by the cloud server, the total computing demand for tasks received at the current moment, and the general resource quantity that is currently over-purchased and not yet used, that is:
[0131] .
[0132] Specifically, the resource scheduling process of cloud-edge collaborative computing in S3 includes:
[0133] S301. Determine the cloud server resource purchase quantity based on the initial task state of the edge server;
[0134] The initial task state of the edge server Through the initial task state of the cloud server inform the initial resource speculation price of each edge server determined;
[0135] The cloud server resource purchase quantity includes the cloud server general resource purchase quantity and the cloud server dynamic resource purchase quantity ;
[0136] S302. Use a cloud server to allocate resources according to the resource purchase volume of each edge server;
[0137] The amount of general resources allocated to each edge server is , and the amount of dynamic resources allocated to each edge server ; In the formula, represents the upper limit of the supply of dynamic resources at the current moment
[0138] S303. Based on the amount of resources allocated by the cloud server, the edge server distributes tasks to the general resources of the cloud server, the dynamic resources of the cloud server, and the local resources of the edge server for calculation. The specific calculation formula is:
[0139]
[0140] Among them, respectively represent the task calculation amounts of the general resources of the cloud server, the dynamic resources of the cloud server, and the local resources of the edge server allocated by the edge server, represents the total calculation demand of tasks received by the edge server per unit time at the current moment, represents the amount of general resources reserved and not yet used by the edge server per unit time at the current moment; The actual resource purchase cost of each edge server in this round is , and the average execution time of this round of tasks is , is the average transmission delay of the task, is the average amount of computing resources required for the task, is the actual computing frequency of the dynamic resources allocated to each edge server;
[0141] Optionally, to pursue the accuracy of cloud-edge collaborative computing, the number of calculation rounds E can be set, and these S301 - S303 are repeatedly executed for E rounds.
[0142] Specifically, the action process of the Q-learning method in S3 includes:
[0143] S311. Use a Q-table to calculate and store the Q values of each state-action pair;
[0144] S312. Gradually converge the policy by iteratively updating the Q value. The expression is:
[0145]
[0146] Among them, s represents the state, a represents the state s corresponding action, represents the state s of the next state, a represents the state The corresponding action r Indicates a reward ε Is the learning rate γ Is the decay coefficient;
[0147] S313. Search for the Q-table after iteration, and select the action corresponding to the largest Q-value according to the current state. Its expression is:
[0148] .
[0149] Specifically, in S3, the steps for double-ended reinforcement learning training include:
[0150] S321. Initialize the cloud server experience pool , the edge server experience pool , the cloud server Q-table and the edge server Q-table;
[0151] Among them, the cloud server Q-table is set sheets, and the dimension is set to ; Initialize the edge server Q-table and set sheets, and the dimension is set to ; Represents the size of the set space;
[0152] S322. Simulate the resource scheduling process of cloud-edge collaborative computing, and store the cloud server samples generated in each round into the cloud server experience pool , and store the storage edge server samples generated in each round into the edge server experience pool , t is an integer greater than or equal to 1, representing the number of simulation rounds;
[0153] S323. Randomly select from the cloud server experience pool , the edge server experience pool minibatch cloud server samples and the edge server samples , j represents the randomly selected samples;
[0154] Among them, the cloud server samples and the edge server samples come from the same simulation round;
[0155] S324. Calculate the target value of each state in the randomly selected cloud server samples and edge server samples respectively;
[0156] Specifically, the Q-value is updated as the target value by the reward after executing : If the next state is an absorbing state, then , otherwise: ;
[0157] Wherein, y ( j ) is the target value corresponding to the state of the selected j th sample, and this target value is related to whether the next state is an absorbing state; r ( j ) is the reward value in the selected j th sample;
[0158] S325. Update the Q values in the Q tables of the cloud server and the edge server according to the target value of each state, and use them as the first type of Q table and the second type of Q table.
[0159] In particular, the main steps of the distribution drift detection method based on hypothesis testing in S5 are as follows:
[0160] E1. Through edge servers, respectively count the service time and cost in each state within a time range;
[0161] E2. If the service time and cost in a certain state exceed the maximum or minimum range within the previous statistical range, report an anomaly; otherwise, report normal;
[0162] E3. Count the reporting results of edge servers within a time range, set the null hypothesis that the distribution Q of the maximum supply of dynamic resources has not changed, that Q has changed, and then perform the following test:
[0163]
[0164] Wherein, α is the probability that the edge server correctly issues an anomaly alarm when Q changes, β is the probability that the edge server wrongly issues an anomaly alarm when Q has not changed, is the total amount of anomaly reports of edge servers counted within the time range.
[0165] Embodiment 3:
[0166] This embodiment is simulated based on Embodiment 1 or Embodiment 2 to implement a cloud-edge collaborative computing network including four edge servers and a central cloud, and the specific parameters are as Figure 6 shown.
[0167] The maximum dynamic resource supply of the cloud server per hour ranges from 0 to 1.6 MG CPU rounds, and the estimated price ranges from 1.5 ¥ to 3.0 ¥ per hour. The total computing resource required for the edge server to receive tasks per time unit is random and follows a distribution, and the experiment is repeated more than 10 times. Using the privacy-enhanced method proposed in this application, the relationship diagram of the training cycle number, normalized service cost and time of this solution is as shown in Figure 7 Figure [1]. In the figure, the horizontal axis represents the training cycle number, and the vertical axis represents the relationship between the normalized service cost and time. It can be seen that as the training cycle number increases, both the service cost and time decrease significantly and finally converge. Figure 8 Figure [2] is a comparison chart of the estimated price and the real resource price. The circumferential layers of the radar chart represent the normalized real resource price. It can be seen that there is an obvious positive correlation between the estimated price and the real price under different resource demand levels of the edge server and the over-purchase volume of general resources . This shows that the estimated price can be sufficiently realistic on the basis of enhancing privacy.
[0168] The method of this application is compared with the random strategy and the general resource single-purchase strategy respectively, and the comparison chart of the normalized service cost and time is as shown in Figure 9 Figure [3]. The abscissa in the figure represents the resource scheduling simulation of cloud-edge collaboration using the random strategy, the general resource single-purchase strategy and the method of this application respectively, and the left and right vertical axes represent the normalized service cost and time respectively. It can be seen that the method of this application has obvious advantages in service quality. The single purchase strategy does not consider the diversity of network resources. Although only purchasing general resources makes the service time relatively shorter, it is at the cost of a huge service cost.
[0169] Figure 10 Figure [4] shows the effect diagram of the hypothesis test distribution drift detection method proposed in this application in the actual deployment environment. The abscissa in the figure represents the system operation rounds, and the vertical axes represent the normalized service cost and time respectively. It can be seen that when the system runs to about 220 rounds, the distribution of the maximum supply of dynamic resources of the cloud server drifts. The method of this application detects this change in a small number of rounds in time and restarts the policy learning. After re-training, the new policy is more adaptable to the actual network environment after the distribution drift than the old policy, and both the service time and cost have decreased, successfully improving the service quality.
[0170] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A privacy-enhanced method for cloud-edge collaborative computing based on a speculative pricing mechanism, characterized in that Including: S1. Collect the cloud server data and edge server data of the cloud-edge collaborative computing network with privacy to be enhanced as the cloud-edge collaborative computing network parameter set, specifically: S11. Collect the unit - time usage prices, frequencies, f and supply - quantity distributions of all types of computing resources in the network; S12. Collect the total task calculation demand distribution received by each edge server per unit time D The purchase range and over-purchase range of all types of computing resources; S13. According to the usage price per unit time Formulate the distribution range of the estimated price ; S14. Use the price per unit time , frequency f , supply distribution , demand distribution D , the purchase range of all types of computing resources , the over-purchase range and the distribution range of prices together as the cloud-edge collaborative computing network parameter set; S2. Perform action state modeling based on the cloud-edge collaborative computing network parameter set, and generate a cloud server state set and an edge server state set respectively; S3. Based on the Q-learning method, simulate the resource scheduling process of cloud-edge collaborative computing, and use the cloud server state set and the edge server state set for double-ended reinforcement learning training to obtain a first type of Q-table and a second type of Q-table; In S3, the steps of performing double-ended reinforcement learning training include: S321. Initialize the experience pool of the cloud server . Initialize the experience pool of the edge server . Initialize the Q-table of the cloud server and the Q-table of the edge server; S322. Simulate the resource scheduling process of cloud-edge collaborative computing, and store the cloud server samples generated in each round into the cloud server experience pool , and store the storage edge server samples generated in each round into the edge server experience pool ; S323. Randomly select from the experience pool of the cloud server and the experience pool of the edge server respectively minibatch a cloud server sample and an edge server sample; wherein, the cloud server sample and the edge server sample come from the same simulation round; S324. Calculate the target value of each state in the randomly selected cloud server sample and edge server sample respectively; S325. Update the Q values in the cloud server Q-table and the edge server Q-table according to the target value of each state as the first type of Q-table and the second type of Q-table; S4. Complete the privacy-enhanced cloud-edge collaborative computing according to the first type of Q-table and the second type of Q-table.
2. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 1, wherein In S1, the unit time usage price of all types of computing resources including the unit time usage price of general computing resources of cloud servers , the unit time usage price of dynamic computing resources of cloud servers and the unit time usage price of local computing resources of each edge server , where max{.} represents the maximum value operation; Frequencies of all types of computing resources in the network f Including the computing frequencies of general resources of cloud servers , the computing frequencies of dynamic resources of cloud servers , the computing frequencies of local resources of each edge server ; Supply distribution of all types of computing resources in the network , where represents the possible maximum supply of cloud server dynamic resources per unit time.
3. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 1, wherein In S12, the distribution of the total task calculation requirements received by each edge server per unit time , where represents the total task calculation requirements that a certain edge server may receive per unit time; The purchase scope of all types of computing resources includes the purchasable scope of general computing resources of edge servers for cloud servers and the purchasable scope of dynamic computing resources of cloud servers ; The over-purchase range of all types of computing resources is the over-purchase range of general computing resources of edge servers for cloud servers .
4. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 1, wherein In S2, there is a corresponding action set for the cloud server status set ; where represents the speculation value space of the cloud server for the resource purchase price of a certain edge server; when the cloud server is in the state , the action taken represents the speculation value of the cloud server for the current round of resource purchase price of a certain edge server; There is a corresponding action set for the edge server status set , where represents the general resource quantity and dynamic resource quantity space that a certain edge server can purchase; when the edge server is in the state , the action taken represents the general resource quantity and dynamic resource quantity that the edge server attempts to purchase.
5. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 4, wherein The cloud server state includes the real-time maximum supply of dynamic computing resources and the transaction information of the previous round with a certain edge server; The edge server state is a triple, including the speculated price informed by the cloud server, the total computing demand for receiving tasks at the current moment, and the general resource amount that is currently over-purchased and not yet used.
6. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 1, characterized in that The resource scheduling process of cloud-edge collaborative computing in S3 includes: S301. Determine the cloud server resource purchase amount based on the initial task state of the edge server; S302. Use the cloud server to perform resource allocation according to the resource purchase amounts of each edge server; S303. Based on the resource amount allocated by the cloud server, the edge server distributes tasks to the general resources, dynamic resources of the cloud server, and local resources of the edge server for calculation.
7. The privacy-enhanced cloud-edge collaborative computing method based on the speculative price mechanism according to claim 1, characterized in that The action process of the Q-learning method in S3 includes: S311. Use a Q-table to calculate and store the Q values of each state-action pair; S312. Gradually converge the policy by iteratively updating the Q value; S313. Search the iterated Q-table and select the action corresponding to the maximum Q value according to the current state.
8. A cloud-edge collaborative computing privacy enhancement system based on a speculative price mechanism, characterized in that, Perform privacy enhancement according to the cloud-edge collaborative computing privacy enhancement method based on the speculated price mechanism according to any one of claims 1-7.
Citation Information
Patent Citations
Industrial Internet of Things cloud edge collaborative unloading and resource allocation method based on DDPG-D3QN
CN116390125A
Intelligent production line management and control cloud edge-end system resource collaboration method
CN116708577A