Cloud security migration method based on federated architecture fusion diffusion model
By integrating the diffusion model and transformer generation model in the federated architecture, the cloud security migration method is solved, and the problems of high cost and poor stability in deep reinforcement learning cloud migration are achieved, efficient and secure data migration strategy optimization is achieved, reducing the risk of redundant data transmission and interruption.
Patent Information
- Application Number
- CN202510471170.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
The existing cloud migration method based on deep reinforcement learning has the problem of excessive migration investment cost and unstable migration, especially in data transmission and processing, and it is difficult to ensure the security and stability of data.
The cloud security migration method based on the federated architecture is adopted. By performing data processing and training locally on the client, the diffusion model is used to generate high-quality compact data representations, and combined with transformer generation models, reducing the amount of data, and at the same time, the distributed training method of federated learning is used to avoid centralized storage and transmission, and enhancing the stability of migration.
It reduces the response cost of cloud migration, improves the stability of the migration process, ensures the security and integrity of data transmission, reduces redundant data transmission, and improves the optimization efficiency of migration strategies.
Smart Images

Figure CN120342875A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud migration, and in particular relates to a cloud security migration method based on a federated architecture fusion diffusion model. Background Art
[0002] With the continuous maturity and popularization of cloud computing technology, cloud migration has become an important strategic choice for many enterprises to optimize resource allocation, improve business flexibility and reduce costs. Under this general trend, cloud migration of social network data is also particularly necessary and urgent. Therefore, how to efficiently and securely complete the cloud migration of social network data has become a key issue that needs to be solved. Especially in terms of migration cost and stability, Dawa makes the challenges faced by cloud migration more complex. Therefore, designing an efficient, secure and stable cloud migration strategy that adapts to different task requirements not only helps to reduce migration costs and ensure the stability of the migration process, but is also the key to achieving successful cloud migration, optimizing resource allocation and coping with complex business environments.
[0003] Existing cloud migration methods are mainly divided into traditional migration methods and machine learning-based migration methods. Traditional migration methods usually rely on a lot of manual operations, and due to the lack of automated tools, there are many repetitive tasks, which increases costs. For example, the migration method based on virtual machine images requires a lot of storage space to generate and transmit huge virtual machine image files, especially when containing a large amount of data. This will directly increase the storage and transmission costs; the migration optimization method based on ant colony algorithm uses parallel optimization to improve the migration speed, but it requires a lot of iterations and calculations to find a satisfactory solution, which means that a lot of computing resources need to be invested, including server running time and computing power, which directly increases the cost of migration. The migration method based on machine learning accurately captures the dependencies between data through intelligent algorithms, solves the data transmission problem in the traditional migration process, and provides better resource utilization and cost-effectiveness. For example, the software migration optimization strategy based on transfer learning selects the optimal data transmission path and resource allocation scheme by reusing the parameters and features of the source model, but because its optimization process involves the movement and redistribution of a large amount of data, the data cannot be transmitted stably.
[0004] Deep Reinforcement Learning (DRL), as a machine learning paradigm that combines deep learning and reinforcement learning, has shown great potential in the field of cloud migration. DRL enables the agent to learn from experience by continuously learning, trial and error, and optimizing its own behavior in the interaction with the environment. It is particularly suitable for optimizing long-term decision-making strategies for cloud migration. In the field of optimization strategies, DRL shows strong adaptability and generalization capabilities, bringing new solutions to efficient migration.
[0005] However, DRL still faces challenges in terms of cost and stability. DRL is usually considered based on a specific cost model, while the cloud migration cost structure is complex and varies depending on the cloud service provider, resulting in an increase in migration costs. Similarly, during the data migration process, problems such as data loss, data corruption, or data inconsistency may occur, and it is difficult for DRL to handle these problems, leading to unstable migration. Therefore, reducing the migration cost of DRL in cloud migration and enhancing the stability of DRL in dealing with different problems have become the focus of current research.
[0006] Diffusion Models (DM), as a generative model, its core idea is to simulate the physical diffusion process, gradually transform data into noise and then learn the reverse process, and gradually recover the original data from the noise to achieve high-quality generation effects. This high-quality generation can be used to generate more compact and efficient data representations to remove redundant data and reduce the amount of data that needs to be transmitted, thereby reducing the migration cost.
[0007] Federated Learning, as an emerging machine learning method, allows multiple devices or computing nodes to perform model training without sharing the original data. Federated Learning allows data to be retained locally without centralized storage or transmission to the cloud, thus avoiding data loss during the migration process and enhancing stability. Moreover, Federated Learning adopts a distributed training method, and this distributed training method further ensures the stability of migration. Summary of the Invention
[0008] Aiming at the problems of excessive migration input cost and unstable migration existing in the existing cloud migration method based on deep reinforcement learning, the present invention provides a cloud security migration method based on a federated architecture integrating a diffusion model.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A cloud security migration method based on a federated architecture integrating a diffusion model, the method comprising the following steps:
[0011] Step 1: Design a federated architecture including multiple clients and a central server, randomly divide the migration data into several parts and transmit them into the corresponding clients;
[0012] The federated architecture in Step 1 is composed of a central server (Server) and multiple clients (Client 1, Client 2,...), and inside the client, diffusion model-based deep reinforcement learning (DMDRL) is used;
[0013] The diffusion model-based reinforcement learning (DMDRL) of each client is divided into two parts: a local actor and a global critic;
[0014] The local actor is composed of the interaction between the diffusion model and the environment. It remains local throughout the entire training process and generates an action a through interaction with the local environment i ;
[0015] The global critic is composed of the interaction between the state space and the diffusion model.
[0016] Step 2: Use the deep reinforcement learning method to update the diffusion model inside the client, interact with the environment, and then aggregate the results of all clients to the central server;
[0017] The diffusion model in Step 2 includes forward diffusion and reverse denoising, specifically:
[0018] Step 2.1: The generation of cloud migration workloads is used as the continuous state space S = {S1, S2, ……, S n} of the diffusion model, where S i represents the i-th state variable; at each time step i, the diffusion model uses the input state S i (the state is defined by the API information provided by the migration software, including the expenditure cost and response time of the physical node, etc.) to predict the action a i , and then executes the action a i , interacts with the environment within the time from 1 to T i , and generates an accumulated reward r and the next state S i+1 ;
[0019] Step 2.2: Forward diffusion:
[0020] First, the original migration data input passes through a 4×4 and a 2×2 convolutional kernel to extract the input state S i , and the ReLU activation function is used to accelerate the training speed. Then, through the flattening layer (Flatten), while keeping the original order of the input data unchanged, all elements of the input data are traversed, and all elements are arranged in a one-dimensional array in the original order;
[0021] Then the data is passed into a Gated Recurrent Unit (GRU) layer. In the GRU layer, the Update Gate and the Reset Gate capture long-term dependencies in sequential data by controlling the flow of information. Here, the data is processed using GRU to eliminate corresponding incomplete redundant data while preserving the state-action reward correspondence, which can reduce the traffic of migrated data when accessing the API later. The data processed by the GRU layer is embedded into an embedding sequence through a separate Multilayer Perceptron (MLP), and then passed through a Dropout layer, which randomly shuts down a part of the neurons, that is, the outputs of these nodes are set to 0, to prevent the neural network from overfitting to the training data. Then, the diffusion time step h Ti is added to the return time step h R to strengthen the condition of the stacked input and formulate the input of GPT2. The formula is as follows:
[0022] h tokens = LN(h Ti + h R + E pos )
[0023] where E pos represents the positional embedding, LN represents layer normalization for stable training, and h tokens represents the input of GPT2;
[0024] The GPT2 is a decoder-only transformer that combines self-attention mechanisms to further generate more compact data to reduce the data volume. The GPT2 architecture is used as a trainable backbone in DMDRL to process sequential inputs, and GPT2 outputs an updated representation h out . The formula is:
[0025] h out = transformer(h tokens )
[0026] Finally, a prediction head composed of fully connected layers is used to randomly predict the corresponding noise at the diffusion time step k. The prediction head is a two-layer MLP with Mish activation. The Mish activation function has a faster convergence speed than the traditional ReLU activation function, and the predicted noise shares the same dimensional space as the original input. This noise is used for reverse denoising during the inference process.
[0027] During the forward diffusion process of the diffusion model, the model randomly samples the diffusion time step k ∼ U(1, K) in each iteration to obtain a noise sequence x k ;
[0028] Step 2.3: Reverse Denoising:
[0029] In the reverse denoising stage, the goal of the diffusion model is to recover the original data from the data with added noise. In the reverse denoising stage of DMDRL, it is necessary to recover the action sequence of this step length from the noise sequence;
[0030] First, gradually add noise to the data in K steps from the forward diffusion chain to obtain the noise sequence x k , and the formula is:
[0031]
[0032] where τ is the sampling trajectory from k ∼ U(1, K), N represents the normal distribution, and I represents the identity matrix. β max = 10, β min = 0.1. The noise sequence includes the original data, the state-action and reward of each round of training in addition to the noise;
[0033] Then, the diffusion model recovers the data by iteratively removing the noise. The reverse diffusion chain for removing the noise is constructed as:
[0034] p θ ((x k (τ)|x k-1 (τ)), x0): = N(x k-1 (τ)|μ θ (x k (τ), x0, k))
[0035] where μ θ (x k (τ), x0, k) represents the mean learned by the model. At each step, the model tries to estimate the denoised data and gradually reduces the influence of the noise.
[0036] The goal of the diffusion model is to remove the noise and separate the action sequence A{a1, a2, ……, a i} and the state-action-reward sequence T{S1, a1, r1, ……, S i , a i , r i} containing the transitions during the training process. Finally, calculate the loss L dnoise and update the model to enter the next round of iteration. The formula for the loss L dnoise is:
[0037]
[0038] where ε is the noise added during the forward diffusion process, ε kis the noise predicted by the prediction head at the k-th diffusion time step.
[0039] The reverse denoising process of the diffusion model is an iterative process that uses the noise distribution and conditional information learned by the model to recover high-quality data from the noise. This process has been successfully applied to cloud migration tasks to reduce communication costs.
[0040] Step 3: After aggregation by the central server, the new diffusion model is distributed to each client, and the optimal model is iteratively updated to find the optimal cloud migration strategy.
[0041] In each round of training, the global critic Critic is updated based on the state-action pairs {S i , a i} generated by the interaction between the state space and the diffusion model, so as to obtain the model parameter h out and is sent to the central server for aggregation.
[0042] The client is updated in each round of training. First, the central server samples the data of client C i and participates in the training. Client C i obtains the latest model parameter h out . Subsequently, the local actor Actor is initialized. The local actor Actor is a policy network responsible for mapping the environmental state into the action space. Client C i uses reinforcement learning based on the diffusion model to update the probability distribution of the local actor Actor to generate actions based on the model parameter h out The formula is:
[0043]
[0044] where J(h out ) is the expectation of the cumulative reward, a t and s t are the action and state at time step t respectively, and A t is the advantage function, which represents the degree of superiority or inferiority of taking action a t relative to the average performance in state s t ;
[0045] After updating the local actor Actor, client C i interacts with the local environment through the local actor Actor updated in t phases and collects the new action a t ;
[0046] Finally, client C i starts the optimization of the global critic Critic and calculates the update of the loss parameter of the global critic Critic. The formula is:
[0047]
[0048] Among them, c represents the total number of clients, N i refers to the total number of samples of client C i and is the total number of samples of all clients. The central server performs random subsampling on a group of clients C calculates and updates the model parameters generated thereby. The formula is: i
[0049]
[0050] Among them, α is the learning rate, is the local gradient of client C i at the nth iteration h out ;
[0051] Finally, the central server implements local training and transmits the latest model parameters to all clients in c. After t iterations, a globally applicable global model is obtained in a globally averaged aggregation manner. The formula is:
[0052]
[0053] Among them, π is the action distribution when a given state is given, Q refers to the state-action value of the state-action pair S-a, c represents the total number of clients, N i refers to the total number of samples of client C i ;
[0054] Compared with the prior art, the present invention has the following advantages:
[0055] (1) Reducing cloud migration response costs: The method generates high-quality data by introducing noise in the diffusion model and uses a transformer in the diffusion model to generate more compact and efficient data, reducing additional costs;
[0056] (2) More stable cloud migration: Introducing a federated architecture into DMDRL, the distributed operation mode of the federated architecture and its characteristic of not directly accessing the original data are used to jointly achieve stable data transmission with the diffusion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a framework diagram of cloud migration based on deep reinforcement learning;
[0058] Figure 2 is a diagram of the DMDRL training process;
[0059] Figure 3 is a forward diffusion architecture diagram of the diffusion model;
[0060] Figure 4 Reverse denoising graph for the diffusion model;
[0061] Figure 5 Is the federated learning process. Detailed implementation method
[0062] To understand the present invention in depth, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples is intended to deepen the comprehensive understanding of the disclosed content of the present invention.
[0063] A cloud security migration method based on a federated architecture integrating a diffusion model, as Figure 1 shown, the method includes the following steps:
[0064] Step 1: Design a federated architecture including multiple clients and a central server, randomly divide the migration data into several parts and transmit them into the corresponding clients;
[0065] The federated architecture in Step 1 consists of a central server (Server) and multiple clients (Client 1, Client 2,...), and within the client, deep reinforcement learning (DMDRL) based on the diffusion model is used;
[0066] The deep reinforcement learning (DMDRL) based on the diffusion model for each client is divided into two parts: a local actor Actor and a global critic Critic;
[0067] The local actor Actor is composed of the interaction between the diffusion model and the environment, and remains local throughout the entire training process, generating an action a through interaction with the local environment i ;
[0068] The global critic Critic is composed of the interaction between the state space and the diffusion model.
[0069] Step 2: Use the deep reinforcement learning method to update the diffusion model inside the client, interact with the environment, and then aggregate the results of all clients to the central server; DMDRL guides the generation of cloud migration policies through training the diffusion model, and its training process is as Figure 2 shown.
[0070] The diffusion model in Step 2 includes forward diffusion and reverse denoising, as Figure 3 shown, specifically:
[0071] Step 2.1: The generation of cloud migration workloads is used as the continuous state space S = {S1, S2,..., S n} of the diffusion model, where S irepresents the i-th state variable; at each time step i, the diffusion model takes the input state S i (The state is defined by the API information provided by the migration software, including the expenditure cost, response time, etc. of the physical node) to predict the action a i , and then executes the action a i , and interacts with the environment within the time from 1 to T i , generates the cumulative reward r and the next state S after the action is executed i+1 ;
[0072] Step 2.2: Forward diffusion:
[0073] First, the original data input passes through a 4×4 and a 2×2 convolutional kernel to extract the input state S i , and the ReLU activation function is used to accelerate the training speed. Then, through the Flatten layer, while keeping the original order unchanged, all elements of the input data are traversed, and all elements are arranged in a one-dimensional array in the original order;
[0074] Then the data is passed into a Gated Recurrent Unit (GRU) layer. The Update Gate and Reset Gate in the GRU layer capture the long-term dependencies in the sequential data by controlling the flow of information. Here, the data is processed by GRU to eliminate the corresponding incomplete redundant data while retaining the correspondence between the state-action rewards, which can reduce the traffic of migrated data when accessing the API later. The data processed by the GRU layer is embedded into the embedding sequence through a separate Multilayer Perceptron (MLP), and then through the Dropout layer, a part of the neurons are randomly turned off, that is, the outputs of these nodes are set to 0, to prevent the neural network from overfitting to the training data. Then, the diffusion time step h Ti is added to the return time step h R to strengthen the conditions of the stacked input and formulate the input of GPT2. The formula is as follows:
[0075] h tokens = LN(h Ti + h R + E pos )
[0076] where E pos represents the positional embedding, LN represents the layer normalization for stable training, and h tokens represents the input of GPT2;
[0077] The GPT2 is a decoder-only Transformer that incorporates self-attention mechanism to further generate more compact data to reduce the data volume. The GPT2 architecture is adopted as the trainable backbone in DMDRL to process sequential inputs, and the GPT2 outputs the updated representation h out , and the formula is:
[0078] h out = transformer(h tokens )
[0079] Finally, a prediction head consisting of fully connected layers is used to randomly predict the corresponding noise at the diffusion time step k. The prediction head is a two-layer MLP with Mish activation. The Mish activation function has a faster convergence rate than the traditional ReLU activation function, and the predicted noise shares the same dimensional space with the original input. This noise is used for reverse denoising during the inference process.
[0080] During the forward diffusion process of the diffusion model, the model randomly samples the diffusion time step k~U(1, K) in each iteration to obtain the noise sequence x k ;
[0081] Step 2.3: Reverse denoising:
[0082] In the reverse denoising stage, the goal of the diffusion model is to recover the original data from the data with added noise. In the reverse denoising stage of DMDRL, it is necessary to recover the action sequence of this step from the noise sequence. The reverse denoising process is as Figure 4 shown.
[0083] First, noise is gradually added to the data in K steps from the forward diffusion chain to obtain the noise sequence x k , and the formula is:
[0084]
[0085] where τ is the sampling trajectory from k~U(1, K), N represents the normal distribution, I represents the identity matrix, β max = 10, β min = 0.1, and the noise sequence includes the original data, the state-action and reward of each round of training in addition to the noise;
[0086] Then, the diffusion model recovers the data by iteratively removing the noise. The reverse diffusion chain for removing the noise is constructed as:
[0087] p θ ((x k (τ)|x k-1 (τ)), x0): = N(x k-1 (τ)|μθ (x k (τ), x0, k))
[0088] Among them, μ θ (x k (τ), x0, k) represents the mean value learned by the model. In each step, the model tries to estimate the denoised data and gradually reduces the influence of noise.
[0089] The goal of the diffusion model is to remove noise and separate the action sequence A{a1, a2, ……, a i}} and the state-action return sequence T{S1, a1, r1, ……, S i , a i , r i}} that contains the transitions during the training process, and finally calculates the loss L dnoise and updates the model to enter the next iteration. The formula for the loss L dnoise is:
[0090]
[0091] Among them, ε is the noise added during the forward diffusion process, and ε k is the noise predicted by the prediction head at the k-th diffusion time step.
[0092] The reverse denoising process of the diffusion model is an iterative process that uses the learned noise distribution and conditional information of the model to recover high-quality data from the noise. This process has been successfully applied to cloud migration tasks to reduce communication costs.
[0093] Step 3: After the central server aggregates, it distributes the new diffusion model to each client, iteratively updates to find the optimal model, and uses this to find the optimal cloud migration strategy.
[0094] In each round of training, the global critic Critic is updated according to the state-action pairs {S i , a i} generated by the interaction between the state space and the diffusion model, so as to obtain the model parameter h out and is sent to the central server for aggregation.
[0095] Figure 5 In, the client is updated in each round of training. First, the central server samples the data of client C i and participates in the training. Client C i obtains the latest model parameter h out , and then initializes the local actor Actor. The local actor Actor is a policy network responsible for mapping the environmental state to the action space. Client C iUtilize reinforcement learning based on the diffusion model, based on the model parameter h out Update the probability distribution of the local participant Actor to generate actions. The formula is as follows:
[0096]
[0097] where J(h out ) is the expectation of the cumulative reward, a t and s t are the action and state at time step t respectively, and A t is the advantage function, indicating the degree of superiority or inferiority of taking action a t under state s t compared to the average performance;
[0098] After updating the local participant Actor, the client C i interacts with the local environment through the local participant Actor updated in t stages and collects the new action a t ;
[0099] Finally, the client C i initiates the optimization of the global critic Critic and calculates the update of the loss parameter of the global critic Critic. The formula is as follows:
[0100]
[0101] where c represents the total number of clients, N i refers to the total number of samples of client C i , is the total number of samples of all clients. The central server performs random subsampling on a group of clients C i and calculates and updates the generated model parameters. The formula is as follows:
[0102]
[0103] where α is the learning rate, is the local gradient of client C i at the nth iteration h out ;
[0104] The central server implements local training and transmits the latest model parameters to all clients in c. After iterating t times, a global model with universality is obtained in a globally averaged aggregation manner. The formula is as follows:
[0105]
[0106] Among them, π is the action distribution when a given state, Q refers to the state-action value of the state-action pair S-a, c represents the total number of clients, and N i refers to the total number of samples of client C i .
[0107] To verify the effectiveness of the deep reinforcement learning method (Fed-DMDRL) based on the federated architecture fusion diffusion model proposed in the present invention, experiments evaluate the cost and stability for two different APIs of the social network.
[0108] Table 1 Comparison between Fed-DMDRL and single-plan methods
[0109]
[0110] The two methods of REMaP and IntMA focus on the resource consumption of a single component. The migration results are shown in Table 1. When Compose is the key API (i.e., giving priority to responding to Compose), Fed-DMDRL is the cheapest in terms of cost ($9.6 per minute), while REMaP and IntMA reach $39.4 and $39.6 respectively, which are more than 4 times that of Fed-DMDRL; when User-timeline is the key API (i.e., giving priority to responding to User-timeline), Fed-DMDRL is also the cheapest in terms of cost ($10.4 per minute), while REMaP and IntMA reach $41.5 and $42.5 respectively, which are also more than 4 times that of Fed-DMDRL. This is because there is at least one traffic service in a non-backend data center for one request after the migration of REMaP and IntMA, resulting in the deterioration of API latency. Moreover, although their goal is to reduce the traffic between data centers (a part of the cost), the plan to find the minimized target is based on a simple heuristic, which leads to more resource waste in terms of cost. In addition, it can be observed that the cost is slightly greater when giving priority to responding to User-timeline than when giving priority to responding to Compose. This is because the amount of application or service data involved in the migration of Compose is relatively small, and the data structure is simple, so the migration cost is relatively low. The migration of User-timeline involves a large amount of user data, and the data relationship is complex (such as including user interactions, timestamps, multimedia content, etc.), so the migration cost is higher.
[0111] The number of interruptions generated when giving priority to responding to Compose during the migration is the same as that when giving priority to responding to User-timeline. There are no interruptions generated during the migration of Fed-DMDRL, and the other two methods both generate 6 interruptions. Fed-DMDRL has an obvious advantage and the migration process is the most stable.
[0112] Table 2 Comparison between Fed-DMDRL and multi-plan methods (cost optimization)
[0113]
[0114]
[0115] Table 3 Comparison between Fed-DMDRL and multi-plan methods (safety optimization)
[0116]
[0117] Table 2 and Table 3 are the experimental results under cost optimization and safety optimization respectively (note: the bold fonts are the best results in each column).
[0118] When the main goal of the multi-plan method is cost and Compose is the key API, Fed-DMDRL is the cheapest ($9.6 per minute), while the costs of Atlas and NSGA-Ⅲ are $9.8 per minute and $10.2 per minute respectively; when User-timeline is the key API, Fed-DMDRL is also the cheapest ($10.4 per minute), while the costs of Atlas and NSGA-Ⅲ are $10.9 per minute and $11.2 per minute respectively. This is mainly because the DRL model in Fed-DMDRL has learned to screen available plans and will not waste time on infeasible plans. The costs of random search are $10.5 per minute and $21.9 per minute respectively, and its results are random. In addition, Compose migrates multiple containers to work together, and the dependency relationship between containers is clear, so the migration process is relatively simple; User-timeline migration needs to handle complex data dependency relationships, data consistency, and cross-system data synchronization and other issues, which increases the complexity of migration; this results in the migration cost of User-timeline being slightly higher than that of Compose.
[0119] In both cases, the number of interruptions of Fed-DMDRL is also the least, indicating that Fed-DMDRL has the best stability. This is attributed to the strong security and privacy capabilities of the federated architecture, which allows each node or component to collaborate and communicate while maintaining relative independence. This independence helps to reduce dependencies and complexities during the migration process, thereby reducing the risk of interruption. Each node or component can migrate and test independently to ensure the smooth progress of the migration process.
[0120] When compared with the secure-optimal migration plan, Fed-DMDRL is one of the solutions with the fewest interruptions when Compose is the key API. The number of interruptions for Atlas is also 0. For NSGA-Ⅲ and random search, the number of interruptions reaches 3 and 5 respectively. When User-timeline is the key API, Fed-DMDRL is also the solution with the fewest interruptions. For Atlas, NSGA-Ⅲ and random search, the number of interruptions reaches 1, 3 and 5 respectively. Fed-DMDRL does not need to be centralized to a central server for processing when dealing with data. The data of each software can be processed and calculated locally. Therefore, the most secure solution can be selected. It can be found that although cost is not the main goal of this plan, Fed-DMDRL is also the cheapest in both cases. The strategy of Atlas will slightly increase the cost caused by traffic (13.3 US dollars per minute), and the same is true for NSGA-Ⅲ (11.1 US dollars per minute). For random search, its costs reach 72.7 US dollars and 38.7 US dollars per minute respectively.
[0121] The content not detailedly described in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention are described above to facilitate the understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A cloud security migration method based on a federated architecture integrating diffusion models, characterized in that, The method includes the following steps: Step 1: Design a federated architecture that includes multiple clients and a central server, randomly divide the migration data into several parts, and transfer them into the corresponding clients; Step 2: Use the deep reinforcement learning method to update the diffusion model inside the client, interact with the environment, and then aggregate the results of all clients to the central server; Step 3: After aggregation by the central server, distribute the new diffusion model to each client, iteratively update the optimal model, and use this to find the optimal cloud migration strategy.
2. The cloud security migration method based on the federated architecture integrating the diffusion model according to claim 1, characterized in that The federated architecture in Step 1 consists of a central server and multiple clients, and reinforcement learning based on the diffusion model is used inside the clients; The reinforcement learning based on the diffusion model for each client is divided into two parts: a local actor Actor and a global critic Critic; The local actor Actor is composed of the interaction between the diffusion model and the environment; The global critic Critic is composed of the interaction between the state space and the diffusion model.
3. The cloud security migration method based on the federated architecture integrated diffusion model according to claim 2, wherein, The diffusion model in Step 2 includes forward diffusion and reverse denoising, specifically: Step 2.1: The generation of cloud migration workloads is used as the continuous state space S = {S1, S2, ……, S n}, where S i represents the i-th state variable; at each time step i, the diffusion model takes the input state S i to predict the action a i , and then executes the action a i , interacts with the environment within the time range from 1 to T i , generates the cumulative reward r and the next state S i+1 after the action is executed; Step 2.2: Forward diffusion: First, the original data input passes through a 4×4 and a 2×2 convolutional kernel to extract the input state S i , and the ReLU activation function is used to accelerate the training speed. Then, through the flattening layer, all elements of the input data are traversed while maintaining the original order, and all elements are arranged into a one-dimensional array in the original order; Then the data is passed into a gated recurrent unit layer. In the gated recurrent unit layer, the update gate and the reset gate capture long-term dependencies in sequential data by controlling the flow of information. The data processed by the gated recurrent unit layer is embedded into an embedding sequence through a separate multi-layer perceptron, and then passed through a Dropout layer to randomly deactivate a part of the neurons. Then, the diffusion time step h Ti is added to the return time step h R to strengthen the conditions of the stacked inputs and formulate the input of GPT2. The formula is as follows: h tokens = LN(h Ti + h R + E pos ) Among them, E pos represents positional embedding, LN represents layer normalization for stable training, and h tokens represents the input of GPT2; The GPT2 is a decoder-only transformer that uses the GPT2 architecture as a trainable backbone in reinforcement learning based on diffusion models to process sequential inputs, and the GPT2 outputs an updated representation h out , and the formula is: h out = transformer(h tokens ) Finally, use a prediction head composed of fully connected layers to randomly predict the corresponding noise at the diffusion time step k, and the prediction head is a two-layer MLP with Mish activation; During the forward diffusion process of the diffusion model, the model randomly samples a diffusion time step k ~ U(1, K) at each iteration to obtain a noise sequence x k ; Step 2.3: Reverse denoising: First, add noise to the data gradually in K steps from the forward diffusion chain to obtain the noise sequence x k , and the formula is: Among them, τ is the sampling trajectory from k ∼ U(1, K), N represents the normal distribution, and I represents the identity matrix. β max = 10, β min = 0.
1. The noise sequence includes the original data, the state-action and reward of each round of training in addition to the noise. Then, the diffusion model restores the data by iteratively removing noise, and the reverse diffusion chain for removing noise is constructed as: p θ ((x k (τ)|x k-1 (τ)),x0):=N(x k-1 (τ)|μ θ (x k (τ),x0,k)) Among them, μ θ (x k (τ), x0, k) represents the mean value learned by the model; The goal of the diffusion model is to remove noise and separate the action sequence A{a1, a2, ……, a i} and the state-action return sequence T{S1, a1, r1, ……, S i , a i , r i} that contains the intermediate state-action rewards during training, and finally calculate the loss L dnoise and update the model to enter the next iteration. The formula for the loss L dnoise is as follows: where ε is the noise added in the forward diffusion process, and ε k is the noise predicted by the prediction head at the k-th diffusion time step.
4. A cloud security migration method based on a federated architecture integrating a diffusion model according to claim 3, characterized in that The specific operation of Step 3 is: First, the central server samples client C i The client C i Get the latest model parameters h out , then initialize the local participant Actor, client C i Using reinforcement learning based on diffusion model, based on model parameter h out Update the probability distribution of the local actor's generated action, the formula is: Among them, J(h out ) is the expectation of the cumulative reward, a t and s t are the action and state at time step t respectively, A t is the advantage function, indicating the degree of superiority or inferiority of taking action a t under state s t relative to the average performance; After updating the local participant Actor, client C i The local participant Actor updated through t stages interacts with the local environment and collects the new action a t Collect; Finally, client C i Starts the optimization of the global critic and calculates the loss parameter update of the global critic, with the formula: Among them, c represents the total number of clients, N i refers to the total number of samples of client C i , and is the total number of samples of all clients. The central server performs random subsampling on a group of clients C i and calculates and updates the model parameters generated by it. The formula is as follows: where α is the learning rate, is the client C i at the nth iteration h out of the local gradient; The central server implements local training and transmits the latest model parameters to all clients in c. After t iterations, a globally applicable global model is obtained in a globally averaged aggregation manner. The formula is as follows: Among them, π is the action distribution when a given state, Q refers to the state-action value of the state-action pair S-a, c represents the total number of clients, and N i refers to the client C i is the total number of samples.