Data flow scheduling optimization method and system based on reinforcement learning

By building a federal reinforcement learning environment and dynamic alliance formation system, the privacy protection and collaboration stability problems in data flow scheduling optimization are solved, flexible collaboration and fair incentives are achieved, and resource utilization and task execution efficiency are improved.

CN120342890APending Publication Date: 2025-07-18SHENZHEN ZHUOWEIYA TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510632617.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing data flow scheduling optimization methods have problems such as insufficient privacy protection, rigid collaboration relationships, unfair incentive mechanisms, and inability to adapt to the changing task environment, resulting in the risk of privacy leakage, ineffective collaboration and instability of collaborative relationships.

Method used

Using a method based on reinforcement learning, a federated reinforcement learning environment is built, sensitive data is processed through differential privacy algorithms, a dynamic alliance formation system is established, a Shapley value is used to evaluate the value of participants, a multi-level incentive protocol and reputation evaluation system is introduced, decentralized strategy aggregation is performed, and global data flow scheduling is optimized.

Benefits of technology

It realizes that while protecting the privacy of participants, it improves resource utilization and task execution efficiency, promotes flexible collaboration and long-term stable collaboration relationships, reduces invalid collaboration overhead, and improves the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342890A_ABST
    Figure CN120342890A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data flow scheduling optimization, and discloses a data flow scheduling optimization method and system based on reinforcement learning. The method comprises the steps that a federal reinforcement learning environment is constructed, and each participant maintains a local data flow state representation and scheduling strategy model; a dynamic alliance forming system is realized, and a temporary alliance is formed by evaluating the value of a participant through a Shapley value; establishing a multi-level incentive protocol, and dynamically distributing earnings; a reputation evaluation system is introduced, historical cooperation behaviors are recorded, and reputation scores are calculated; and executing decentration strategy aggregation, and optimizing a global data flow scheduling strategy. The resource utilization rate and the task execution efficiency are improved while privacy of participants is protected, and multi-party collaborative optimization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data stream processing, and more specifically, to a method and system for optimizing data stream scheduling based on reinforcement learning. Background Art

[0002] With the development of distributed computing and cloud computing technologies, data stream processing systems across organizational boundaries have become increasingly popular. In distributed environments such as data centers, cloud computing, and edge computing, multiple independent organizations (such as different enterprises, institutions, or departments) often have their own computing resources and data stream processing tasks, and hope to optimize resource utilization and task execution efficiency through cooperation.

[0003] However, the existing data stream scheduling optimization methods mainly have the following problems: First, traditional centralized scheduling methods need to collect sensitive system status and load data of each participant, resulting in a serious risk of privacy leakage; Second, existing cooperation systems usually adopt a fixed participation structure and lack the ability to dynamically adjust cooperation relationships according to task characteristics and resource status; Third, there is a lack of a fair and effective incentive mechanism, making it difficult to balance the contributions and benefits of each participant, resulting in a widespread "free-rider" problem; Fourth, existing methods lack a long-term reputation evaluation mechanism and it is difficult to establish a lasting and stable cooperation relationship; Finally, centralized global optimization often ignores local characteristics and cannot adapt to changing task environments.

[0004] Therefore, there is an urgent need for a data stream scheduling optimization method that can achieve flexible cooperation, fair incentives, and long-term stability while protecting the privacy of participants. Summary of the Invention

[0005] The present invention provides a method and system for optimizing data stream scheduling based on reinforcement learning, which solves the technical problems such as insufficient privacy protection, rigid cooperation relationships, and unfair incentive mechanisms in related technologies.

[0006] The first aspect of the present invention discloses a method for optimizing data stream scheduling based on reinforcement learning, including the following steps: Construct a federated reinforcement learning environment, where each participant maintains a data stream state representation and a scheduling policy model locally, processes sensitive data using differential privacy algorithms, and uploads encrypted model parameters; Implement a dynamic coalition formation system, and evaluate the value of each participant according to task characteristics and resource status through a contribution metric based on the Shapley value, prompting the participants to autonomously form a temporary coalition; Establish a multi-level incentive protocol, and dynamically allocate benefits based on the data quality, computing contributions, and model improvement degree of the participants; Introduce a reputation evaluation system, collect and record the historical cooperation behaviors of each participant, calculate the reputation score, and use the reputation information as the state input of the reinforcement learning environment; Execute decentralized policy aggregation. The participating parties train the scheduling model based on local data, and the participating parties with high credibility in the alliance execute model aggregation to optimize the global data flow scheduling policy.

[0007] Furthermore, the construction of the federated reinforcement learning environment includes: Use a multi-layer perceptron network to construct a local data flow state representation model, and map input data such as data flow task characteristics, resource utilization, and queue status into a low-dimensional state representation vector; Apply the Laplace noise mechanism to the local state representation vector to achieve differential privacy protection; Construct a local scheduling policy model based on the double deep Q-network algorithm, with the state representation vector as the input and the data flow scheduling decision as the output; Encrypt the local model parameters using a homomorphic encryption scheme and then transmit them.

[0008] Furthermore, the implementation of the dynamic coalition formation system includes: Construct a coalition value evaluation model to quantify the value that can be generated by coalitions formed by different combinations of participating parties; Based on the coalition value evaluation model, apply the Shapley value algorithm to calculate the marginal contribution of each participating party; Use the Monte Carlo sampling method to approximately calculate the Shapley value; Based on the Shapley value of the participating parties and the current task characteristics, construct a coalition formation decision model to determine the optimal coalition combination.

[0009] Furthermore, the establishment of the multi-level incentive protocol includes: Construct a multi-dimensional contribution evaluation model to quantify the contributions of participating parties from three dimensions: data quality, computing contribution, and model improvement; Based on the comprehensive contribution degree of the participating parties, construct a revenue distribution algorithm; Set different incentive levels according to the historical contributions and current values of the participating parties; Formulate smart contract constraint rules based on blockchain to clarify the rights and responsibilities of each participating party.

[0010] Furthermore, the introduction of the reputation evaluation system includes: Construct a multi-dimensional behavior monitoring model to collect and record the historical cooperation behaviors of participating parties; Construct a reputation calculation algorithm based on time decay, making the impact of recent behaviors on reputation greater; Divide the participating parties into different levels according to their reputation scores; Use the reputation information as the state input of the reinforcement learning environment and add a reputation-related reward item to the reward function.

[0011] Furthermore, the execution of decentralized policy aggregation includes: Based on local data, construct and train a scheduling policy model to optimize the local data stream processing performance; Construct a model parameter aggregation algorithm based on reputation weights, and let the parties with high reputation in the alliance execute the aggregation; Adjust the parameter update amplitude according to the characteristics of the parties and task requirements; Dynamically adjust the learning strategy according to the task type and resource status.

[0012] Furthermore, the multi-layer perceptron network in the federated reinforcement learning environment includes: An input layer that receives the original data feature vector, with the dimension being the same as the number of original features; Three hidden layers, which respectively contain 128, 64, and 32 neurons, and each layer uses the ReLU activation function; An output layer that outputs a low-dimensional state representation vector with a dimension of 16.

[0013] Furthermore, the double deep Q network in the federated reinforcement learning environment includes: An evaluation network, which consists of an input layer, three hidden layers, and an output layer. The three hidden layers respectively contain 128, 96, and 64 neurons, and use the LeakyReLU activation function; A target network, with the same structure as the evaluation network; The parameters of the target network are obtained from the evaluation network through soft update, and the update formula is: ; where, is the soft update coefficient, are the parameters of the evaluation network, are the parameters of the target network.

[0014] Furthermore, the reputation calculation algorithm based on time decay in the reputation evaluation system adopts the following formula: ; where, represents the reputation score of party at time , represents the weight of the behavior metric , is the time decay factor, is the time point when the historical behavior occurred, represents the score of the behavior metric of party at time , represents the time interval from the occurrence of the historical behavior to the current time, A weight coefficient that decreases as the time interval increases, ensuring that the most recent behavior has a greater impact on the reputation.

[0015] The second aspect of the present invention discloses a system for optimizing data stream scheduling based on reinforcement learning, which is used to execute the above-mentioned method for optimizing data stream scheduling based on reinforcement learning, including: A federated reinforcement learning environment module, which is used for each participating party to locally maintain the data stream state representation and scheduling policy model, process sensitive data and upload encrypted model parameters; A dynamic coalition formation module, which is used to evaluate the value of each participating party according to the task characteristics and resource status and prompt them to autonomously form a temporary coalition; A multi-level incentive protocol module, which is used to dynamically allocate benefits based on the multi-dimensional contributions of the participating parties; A reputation evaluation module, which is used to collect and record historical cooperation behaviors and calculate reputation scores; A decentralized policy aggregation module, which is used to train a scheduling model based on local data and execute model aggregation to optimize the global data stream scheduling policy.

[0016] The beneficial effects of the present invention are as follows: Through the dynamic coalition formation system, the participating parties can autonomously form a temporary coalition according to the task characteristics and expected benefits, realizing on-demand cooperation. Compared with the federated learning mode with a fixed participating party structure, the cooperation efficiency is improved, and the ineffective cooperation overhead is reduced; The multi-party collaborative scheduling strategy improves the overall performance of the system, improves resource utilization, and reduces the task completion time; The reputation evaluation system and the multi-level incentive protocol ensure that the participating parties obtain benefits matching their contributions, effectively solving the "free-rider" problem; The task-adaptive policy optimization model can flexibly respond to changes in different data stream characteristics and the needs of participating parties, improve the scheduling performance under variable workloads, improve resource utilization efficiency, and shorten the time for dynamically adapting to new task types. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the overall flowchart of the method for optimizing data stream scheduling based on reinforcement learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0019] In at least one embodiment of the present invention, an optimization method for data flow scheduling based on reinforcement learning is disclosed. As Figure 1 shown, it includes the following steps: Step 1: Construct a federated reinforcement learning environment. Each participant locally maintains a data flow state representation and a scheduling policy model, processes sensitive data using the differential privacy algorithm, and uploads encrypted model parameters. In this step, the differential privacy algorithm is used to process sensitive data, and a data flow state representation and a scheduling policy model are locally maintained to implement a privacy-preserving reinforcement learning environment for multi-party collaboration. Specifically, it includes: Step 1.1: Construct a local state representation model. Use a multi-layer perceptron network to construct a local data flow state representation model, and map input data such as data flow task characteristics, resource utilization, and queue status to a low-dimensional state representation vector. The state representation vector is calculated from the local data of the participant : ; where represents the local original state data of the participant , including task queue length, processor utilization, memory usage, etc.; represents the state representation model of the participant .

[0020] The multi-layer perceptron network in this embodiment consists of an input layer, three hidden layers, and an output layer. The input layer receives the original data feature vector, and the dimension is the same as the number of original features; the three hidden layers respectively contain 128, 64, and 32 neurons, and each layer uses the ReLU activation function to increase the non-linear expression ability of the network; the output layer outputs a low-dimensional state representation vector, and the dimension is 16. The network structure is as follows: ; ; ; ; where is the activation function, is the original feature input vector of the participant , is the finally generated low-dimensional state representation vector, respectively represent the first, second, third, and fourth weight matrices, respectively represent the first, second, third, and fourth bias vectors, respectively represent the first, second, and third output vectors.

[0021] Step 1.2: Process sensitive data using the differential privacy algorithm; Apply the Laplace noise mechanism to the local state representation vector to achieve differential privacy protection and ensure that sensitive information is not leaked during the sharing process: ; where, is the state representation vector after adding noise, represents the Laplace distribution noise, and the scale parameter is determined by the privacy budget and the sensitivity of the state vector: , is the finally generated low-dimensional state representation vector.

[0022] Step 1.3: Construct a local scheduling policy model; Construct a local scheduling policy model based on the Double Deep Q-Network (Double DQN) algorithm. The input is the state representation vector, and the output is the data flow scheduling decision. The model includes an evaluation network and a target network , and the parameters are updated through temporal difference learning: ; where, is the learning rate, is the reward obtained for the current action, is the discount factor, is the next state, is the current state, is the current action being executed, are the evaluation network parameters, are the target network parameters, represents the predicted Q-value of executing action in state under parameter , represents the target Q-value of executing action in state under parameter , represents the gradient of the Q-value with respect to parameter , represents the action that maximizes the Q-value in state .

[0023] The double deep Q-network in this embodiment consists of an evaluation network and a target network. The two networks have the same structure and both adopt a fully connected neural network architecture, including an input layer, three hidden layers, and an output layer. The input layer receives a state representation vector with a dimension of 16. The three hidden layers contain 128, 96, and 64 neurons respectively, and the LeakyReLU activation function is used to enhance the network's learning ability for small gradients. The dimension of the output layer is equal to the number of possible actions, and no activation function is used. The network structure is as follows: ; ; ; ; Among them, , , , respectively represent the fifth, sixth, seventh, and eighth weight matrices, , , , respectively represent the fifth, sixth, seventh, and eighth bias vectors, is the activation function, , , respectively represent the fourth, fifth, and sixth output vectors, represents the predicted Q value of executing action under parameter for state . At the same time, the parameters of the evaluation network are softly updated to the target network every certain number of steps: ; Among them, is the soft update coefficient, and its value range is . In this embodiment, it takes 0.01, is the soft update coefficient, are the parameters of the evaluation network, are the parameters of the target network.

[0024] Step 1.4: Implement an encrypted model parameter transmission mechanism; The local model parameters are encrypted using a homomorphic encryption scheme and then transmitted: ; Among them, are the encrypted model parameters, is the homomorphic encryption function, is the public key, are model parameters. Using the properties of homomorphic encryption, the aggregation server can directly perform aggregation operations on the encrypted parameters: ; where, is the aggregated model parameter, is the weight coefficient of participant , is the total number of participants is the encrypted model parameter, are model parameters.

[0025] Step 2: Implement a dynamic coalition formation system. According to the task characteristics and resource status, evaluate the value of each participant through a contribution metric based on the Shapley value, and prompt the participants to autonomously form a temporary coalition; Specifically, it includes: Step 2.1: Construct a coalition value evaluation model; Construct a coalition value evaluation model to quantify the value that can be generated by different combinations of participants forming a coalition: ; where, represents a subset of the participant set , that is, a possible coalition; represents the task set; represents the task 's weight; represents the performance metric of the coalition processing the task (such as throughput or the reciprocal of the completion time); represents the coordination cost of the coalition ; represents the total value of the coalition , that is, the weighted performance gain obtained by the coalition processing all tasks minus the coordination cost.

[0026] Step 2.2: Apply the Shapley value algorithm to calculate the participant contribution degree; Based on the coalition value evaluation model, apply the Shapley value algorithm to calculate the marginal contribution of each participant : ; where, is the set of all participants, is the value function of the coalition , represents the Shapley value of participant , representing the set without participant , represents a subset of Denote the set the number of elements in denote the number of all participants, denote the participant joining the coalition the value after that. In the formula, denotes the factorial operation. Through this formula, calculate the average marginal contribution of participants for all possible coalition permutations .

[0027] Step 2.3: Implement the Monte Carlo sampling approximation calculation method; Since the complexity of accurately calculating the Shapley value is , use the Monte Carlo sampling method for approximate calculation: ; wherein, denote the Shapley value of participant , is the number of samplings, is the th sampling, a random subset that does not include participant , denote adding participant to the random subset to form the value of the new coalition, that is, evaluate the marginal contribution of participant to the coalition.

[0028] Step 2.4: Construct a coalition formation decision model; Based on the Shapley value of the participants and the current task characteristics, construct a coalition formation decision model to determine the optimal coalition combination: ; wherein, is the optimal coalition, is the cost function for participant to join the coalition , related to its Shapley value , denote summing over all participants in the coalition , denote the variable value when the objective function reaches the maximum value, is the value function of the coalition , denote the Shapley value of participant . By solving this optimization problem, participants can autonomously decide whether to join the coalition and with which participants to form a coalition according to the current task characteristics and resource status.

[0029] Step 3: Establish a multi-level incentive agreement to dynamically allocate benefits based on the data quality, computing contributions, and model improvement degree of the participating parties; In this step, a contribution evaluation model is constructed, a benefit distribution algorithm is constructed, and contract constraint rules are formulated to achieve fair benefit distribution based on the contributions of the participating parties. Specifically, it includes: Step 3.1: Construct a multi-dimensional contribution evaluation model; Construct a multi-dimensional contribution evaluation model to quantify the contributions of the participating parties from three dimensions: data quality, computing contributions, and model improvement: ; Among them, represents the comprehensive contribution degree of the participating party , represents the data quality score, represents the computing resource contribution score, represents the model improvement contribution score, , , respectively represent the weight coefficients of the data quality dimension, computing resource contribution dimension, and model improvement contribution dimension, and satisfy .

[0030] For the data quality score, an index combination evaluation is adopted: ; Among them, represents the data diversity score, represents the data quality score, represents the data validity score, represents the data freshness score, , , respectively represent the weight coefficients of the data diversity index, data validity index, and data freshness index.

[0031] Step 3.2: Construct a benefit distribution algorithm based on the contribution ratio; Based on the comprehensive contribution degree of the participating parties, construct a benefit distribution algorithm: ; Among them, represents the benefit obtained by the participating party , represents the total benefit generated by the alliance, represents the comprehensive contribution degree of the participating party , represents the additional reward of the participating party , Represents the total comprehensive contribution of all participating parties, used to incentivize specific behaviors, such as providing high-quality data or scarce computing resources.

[0032] Step 3.3: Construct a differential incentive hierarchy system; Construct a differential incentive hierarchy system, and set different incentive levels according to the historical contributions and current values of the participating parties: ; Among them, represents the incentive level of the participating party , represents the historical cumulative contribution, represents the current contribution, , , , , , are respectively the historical contribution threshold of level 1, the current contribution threshold of level 1, the historical contribution threshold of level 2, the current contribution threshold of level 2, the historical contribution threshold of level , the current contribution threshold of level , represents the different incentive levels from high to low, represents the total number of incentive levels, represents omitting the intermediate levels. Different levels correspond to different benefit coefficients, priorities, and resource allocation amounts.

[0033] Step 3.4: Formulate smart contract constraint rules; Formulate smart contract constraint rules based on the blockchain to clarify the rights and responsibilities of each participating party: ; Among them, represents the contract for the participating party to join the alliance , defines the rights of the participating party (such as data access scope, decision-making weight), defines the responsibilities (such as the computing resources that must be provided, response time), defines the default penalty, defines the additional reward for achieving the goal. The contract terms are automatically executed to ensure a fair and transparent collaboration environment.

[0034] Step 4: Introduce a reputation evaluation system, collect and record the historical collaboration behaviors of each participating party, calculate the reputation score, and use the reputation information as the state input of the reinforcement learning environment; This step constructs a behavior monitoring model, constructs a reputation calculation algorithm, applies a reputation reinforcement feedback system, and realizes the establishment and maintenance of long-term cooperative relationships among participants. Specifically, it includes: Step 4.1: Construct a multi-dimensional behavior monitoring model; Construct a multi-dimensional behavior monitoring model and collect and record the historical cooperation behaviors of participants: ; Among them, represents the set of behavior records of participant , , , respectively represent the 1st, 2nd, and mth behavior indicators, where m represents the total number of behavior indicators, that is, the total number of behavior dimensions monitored by the system, including data sharing accuracy, computing resource provision stability, protocol compliance, etc.

[0035] Each behavior indicator adopts a quantitative scoring method: ; Among them, represents the actual performance of participant on indicator , represents the expected performance, is the scoring function, represents the participant number, represents the behavior indicator number.

[0036] Step 4.2: Construct a reputation calculation algorithm based on time decay; Construct a reputation calculation algorithm based on time decay, making the impact of recent behaviors on reputation greater: ; Among them, represents the reputation score of participant at time , represents the weight of behavior indicator , is the time decay factor, is the time point when the historical behavior occurred, represents participant at time on behavior indicator score, is the total number of behavior indicators, is the total number of historical behavior records. Among them represents the current time point, represents the time interval from the occurrence of the historical behavior to the current time, Represents a weight coefficient that decreases with the increase of the time interval, ensuring that the impact of recent behavior on reputation is greater.

[0037] Step 4.3: Construct a reputation grading model; Construct a reputation grading model and divide the participating parties into different levels according to their reputation scores: ; Represents the reputation level of the participating party and Represents the reputation score of the participating party . , , Are the reputation thresholds for gold, silver, and bronze levels respectively. Gold represents the gold level, Silver represents the silver level, Bronze represents the bronze level, and Probation represents the observation period level. Different reputation levels correspond to different permissions and advantages.

[0038] Step 4.4: Apply the reputation reinforcement feedback system Apply the reputation reinforcement feedback system and input the reputation information as the state of the reinforcement learning environment: ; Among them, Is the extended state vector, which includes the original state of the participating party and the reputation information of all participating parties. And , , Represent the reputation scores of the 1st, 2nd, and Nth participating parties in the alliance respectively, where N is the total number of participating parties in the alliance.

[0039] Based on the extended state vector, adjust the reward function and add a reputation-related reward term: ; Among them, Is the adjusted reward, Is the original reward, Is the reputation change of the participating party (the difference between the current reputation and the reputation at the previous time point), Is the reputation reward weight coefficient, which is used to balance the influence degree of the original reward and the reputation change. Through this system, guide the participating parties to establish long-term cooperative relationships and reduce short-sighted behaviors.

[0040] Step 5: Execute decentralized policy aggregation. The participating parties train the scheduling model based on local data, and the participating parties with high reputation in the alliance execute model aggregation to optimize the global data flow scheduling strategy; This step constructs a local scheduling policy training model, constructs a model parameter aggregation algorithm within the alliance, and realizes the optimization of the global scheduling policy for task adaptation. Specifically, it includes: Step 5.1: Construct a local scheduling policy training model; Based on local data, construct and train a scheduling policy model to optimize the local data stream processing performance: ; Among them, represents the model parameters of participant at the iteration round , represents the model parameters of participant at the iteration round , represents the gradient operator with respect to the parameter , is the loss function, is the local training data batch. The loss function is designed based on the reinforcement learning objective: ; Among them, represents the Q-value function network of participant , is the transition sample of state-action-reward-next state, represents the current state, represents the current action, represents the obtained reward, represents the next state, represents choosing the action with the maximum Q-value among all optional actions in the next state , represents the mathematical expectation, is the discount factor.

[0041] Step 5.2: Construct a model parameter aggregation algorithm based on reputation weights; Construct a model parameter aggregation algorithm based on reputation weights, and let the participants with high reputation in the alliance perform the aggregation: ; Among them, is the aggregated global model parameter, is the number of participants participating in the aggregation, represents the model parameters of participant at the iteration round , is the aggregation weight of participant , which is positively correlated with its reputation score: ; Among them, is the reputation score of the participant , is its relevance score in the current task, represents the participant The product of the reputation score and the task relevance score of represents the comprehensive weight contribution of this participant in the aggregation process. This design ensures that the aggregation process takes into account both the credibility and professional ability of the participants, making the aggregation of model parameters biased towards more credible and relevant participants.

[0042] Step 5.3: Construct a differential parameter update model; Construct a differential parameter update model and adjust the parameter update amplitude according to the characteristics of the participants and the task requirements: ; Among them, is the local retention coefficient of the participant , which determines how much local characteristics are retained, is the global model parameter after aggregation, represents the participant at the iteration round of the model parameter. Adaptive adjustment according to task similarity and participant professionalism: ; Among them, represents the similarity between the task of the participant and the overall task of the alliance, represents the professionalism of the participant in related tasks, represents the current task characteristics (original ), indicating the specific attributes and requirements of the task, is the mapping function (original ), which is used to calculate the local retention coefficient.

[0043] Step 5.4: Implement a task-adaptive policy optimization model; Implement a task-adaptive policy optimization model and dynamically adjust the learning strategy according to the task type and resource status: ; Among them, represents the probability of selecting the action at the state and the task type , is the global Q-value function, is the temperature parameter, which controls the exploration-exploitation balance of the policy, represents all the optional actions.

[0044] According to the task characteristics , dynamically adjust the reward weights and exploration parameters: ; ; Among them, represents the total reward obtained when executing action in state and the task type is , represents the th reward component, is the reward weight for task , is the temperature adjustment function related to the task, is the base temperature parameter, is the number of reward components. is the temperature parameter, which is used to control the exploration-exploitation balance of the policy. A higher temperature value will increase the randomness and exploration of the policy, and a lower temperature value will make the policy more deterministic and exploitative. By dynamically adjusting the temperature parameter according to the task type , the system can adaptively balance exploration and exploitation for different task characteristics. Through this system, the policy can adapt to the requirements of different task scenarios.

[0045] Application example of this embodiment: This section will show the practical application example of this embodiment in a cloud-edge collaborative computing environment, and specifically illustrate how to improve the multi-party collaboration efficiency and protect privacy through a reinforcement learning-based data flow scheduling optimization method.

[0046] Application scenario: In this application example, we consider a hybrid computing environment that includes 3 cloud service providers (denoted as C1, C2, and C3 respectively) and 5 edge computing node owners (denoted as E1 to E5 respectively). These parties each have different scales and types of computing resources, as well as their own data flow processing tasks. The main tasks include real-time image recognition, intelligent transportation data analysis, and industrial sensor data processing, etc. Each party hopes to improve resource utilization and task processing efficiency through collaboration, while protecting sensitive business data and system state information.

[0047] Implementation process example: Build a federated reinforcement learning environment: Each participating party first constructs a local data flow state representation model. Taking cloud service provider C1 as an example, its original state data includes 40 features (including CPU utilization, memory usage, current task queue length, etc.), which are mapped into a 16-dimensional state representation vector through a multi-layer perceptron network. Table 1 shows some examples of C1's original state data: Table 1: Some examples of the original state data of cloud service provider C1

[0048] After the state representation model maps these original data into 16-dimensional vectors, the Laplace noise mechanism is applied for differential privacy protection, and the privacy budget is set to 0.1. Table 2 shows the comparison of some state representation vector values before and after adding noise: Table 2: Comparison of state representation vectors before and after adding differential privacy noise (partial dimensions)

[0049] Implement a dynamic coalition formation system: Based on the coalition value evaluation model, calculate the values of different combinations of participating parties. Suppose there is a batch of image recognition tasks to be processed at a certain moment. Table 3 shows the value evaluation results of different coalition combinations: Table 3: Value evaluation of different coalition combinations

[0050] Calculate the Shapley values of each participating party through the Monte Carlo sampling method, and the number of sampling times M = 1000. Table 4 shows the calculation results of the Shapley values of each participating party: Table 4: Shapley values of each participating party

[0051] Based on the Shapley values and the current task characteristics, the coalition formation decision model determines the optimal coalition as {C1, C2, C3, E1, E3}.

[0052] Establish a multi-level incentive protocol: Based on the multi-dimensional contribution evaluation model, calculate the contributions of all parties participating in the coalition. Table 5 shows the multi-dimensional contribution evaluation results of each participating party: Table 5: Multi-dimensional contribution evaluation of each participating party in the coalition

[0053] According to the comprehensive contribution degree, the revenue distribution algorithm calculates the revenue ratios that each participating party should obtain, and determines the incentive level in combination with historical contributions, as shown in Table 6: Table 6: Revenue distribution and incentive levels of each participating party in the coalition

[0054] Alliance operation and effect verification: After the alliance is formed, each participant trains a scheduling model based on local data and aggregates and optimizes the global scheduling strategy through a decentralized strategy. After 30 rounds of iterative training, the system performance has been significantly improved. Table 7 compares the performance differences between independent scheduling and collaborative scheduling: Table 7: Performance comparison between independent scheduling and collaborative scheduling

[0055] This embodiment has achieved remarkable technical effects in practical applications, mainly reflected in two aspects: privacy protection and collaboration efficiency.

[0056] Privacy protection effect verification: Using differential privacy and homomorphic encryption technologies to protect the sensitive data of participants, the privacy protection effect is obvious. Through simulated attack tests, the risk of original data leakage is evaluated, and the results are shown in Table 8: Table 8: Privacy protection effect evaluation

[0057] Collaboration stability effect verification: Through the reputation evaluation system and multi-level incentive protocol, the "free-rider" problem in traditional systems is effectively solved, and the collaboration stability is improved. The long-term operation results show that as the system operation time increases, stable collaboration relationships are established among participants. In the initial stage (1 - 10 rounds), the participant retention rate is 78.5%, the proportion of malicious behavior is 12.3%, and the alliance reorganization frequency is 3.2 times per day; in the middle stage (11 - 30 rounds), the participant retention rate increases to 92.1%, the proportion of malicious behavior drops to 5.8%, and the alliance reorganization frequency drops to 1.5 times per day; in the stable stage (30+ rounds), the participant retention rate is as high as 96.7%, the proportion of malicious behavior is only 1.2%, and the alliance reorganization frequency is further reduced to 0.7 times per day.

[0058] The experimental results prove the effectiveness of this solution in promoting long-term cooperation, which can gradually establish stable and reliable collaboration relationships as the operation time increases, significantly reduce malicious behavior, and improve the overall efficiency of the system.

Claims

1. A method for optimizing data flow scheduling based on reinforcement learning, characterized in that, It includes the following steps: Construct a federated reinforcement learning environment. Each participant maintains the data stream state representation and scheduling policy model locally, processes sensitive data using the differential privacy algorithm, and uploads the encrypted model parameters; Implement a dynamic coalition formation system. According to the task characteristics and resource status, evaluate the value of each participant through the contribution metric based on the Shapley value, and prompt the participants to autonomously form a temporary coalition; Establish a multi-level incentive protocol, and dynamically allocate benefits based on the data quality, computing contribution, and model improvement degree of the participants; Introduce a reputation evaluation system, collect and record the historical cooperation behaviors of each participant, calculate the reputation score, and use the reputation information as the state input of the reinforcement learning environment; Execute decentralized policy aggregation. The participants train the scheduling model based on local data, and the participant with a high reputation in the coalition executes model aggregation to optimize the global data stream scheduling policy.

2. The optimization method for data flow scheduling based on reinforcement learning according to claim 1, wherein The construction of the federated reinforcement learning environment includes: Use a multi-layer perceptron network to construct a local data stream state representation model, and map the data stream task characteristics, resource utilization, and queue state input data into a low-dimensional state representation vector; Apply the Laplace noise mechanism to the local state representation vector to achieve differential privacy protection; Based on the double deep Q-network algorithm, construct a local scheduling policy model, with the state representation vector as the input and the data stream scheduling decision as the output; Use a homomorphic encryption scheme to encrypt the local model parameters before transmission.

3. The optimization method for data flow scheduling based on reinforcement learning according to claim 1, wherein The implementation of the dynamic coalition formation system includes: Construct a coalition value evaluation model to quantify the value that can be generated by coalitions formed by different participant combinations; Based on the coalition value evaluation model, apply the Shapley value algorithm to calculate the marginal contribution of each participant; Use the Monte Carlo sampling method to approximately calculate the Shapley value; Based on the Shapley value of the participants and the current task characteristics, construct a coalition formation decision model to determine the optimal coalition combination.

4. The optimization method for data flow scheduling based on reinforcement learning according to claim 1, wherein The establishment of the multi-level incentive protocol includes: Construct a multi-dimensional contribution evaluation model to quantify the contributions of participants from three dimensions: data quality, computing contribution, and model improvement; Based on the comprehensive contribution degree of the participants, construct a revenue distribution algorithm; Set different incentive levels according to the historical contributions and current values of the participants; Formulate smart contract constraint rules based on blockchain to clarify the rights and responsibilities of each participant.

5. The optimization method for data flow scheduling based on reinforcement learning according to claim 1, characterized in that The introduction of the reputation evaluation system includes: Construct a multi-dimensional behavior monitoring model to collect and record the historical cooperation behaviors of participants; Construct a reputation calculation algorithm based on time decay, making the impact of recent behaviors on reputation greater; Divide the participants into different levels according to the reputation score; Use the reputation information as the state input of the reinforcement learning environment, and add a reputation-related reward item to the reward function.

6. The optimization method for data stream scheduling based on reinforcement learning according to claim 1, wherein The execution of decentralized policy aggregation includes: Based on local data, construct and train a scheduling policy model to optimize the local data stream processing performance; Construct a model parameter aggregation algorithm based on reputation weights, and the participant with a high reputation in the coalition executes the aggregation; Adjust the parameter update amplitude according to the characteristics of the participants and the task requirements; Dynamically adjust the learning strategy according to the task type and resource status.

7. The data flow scheduling optimization method based on reinforcement learning according to claim 1, wherein The multi-layer perceptron network in the federated reinforcement learning environment includes: The input layer receives the original data feature vector, with the dimension being the same as the number of original features; Three hidden layers, which respectively contain 128, 64, and 32 neurons, and each layer uses the ReLU activation function; The output layer outputs a low-dimensional state representation vector with a dimension of 16.

8. The method for optimizing data stream scheduling based on reinforcement learning according to claim 1, wherein The double deep Q-network in the federated reinforcement learning environment includes: The evaluation network, which is composed of an input layer, three hidden layers, and an output layer. The three hidden layers respectively contain 128, 96, and 64 neurons, and use the LeakyReLU activation function; The target network, with the same structure as the evaluation network; The parameters of the target network are obtained from the evaluation network through soft update, and the update formula is: ; Among them, is the soft update coefficient, is the evaluation network parameter, is the target network parameter.

9. The method for optimizing data flow scheduling based on reinforcement learning according to claim 1, wherein The time decay-based reputation calculation algorithm in the reputation evaluation system adopts the following formula: ; Among them, represents the reputation score of the participant at time ; represents the weight of the behavior indicator ; is the time decay factor, is the time point when the historical behavior occurred, represents the behavior indicator of the participant at time ; is the score of represents the time interval from the occurrence of the historical behavior to the present, represents the weight coefficient that decreases as the time interval increases, ensuring that the most recent behavior has a greater impact on the reputation.

10. A data flow scheduling optimization system based on reinforcement learning, characterized in that, For performing the optimization method of data flow scheduling based on reinforcement learning as described in any one of claims 1-9, it includes: The federated reinforcement learning environment module is used for each participant to locally maintain the data flow state representation and scheduling policy model, process sensitive data, and upload encrypted model parameters; The dynamic coalition formation module is used to evaluate the value of each participant according to the task characteristics and resource status and prompt them to autonomously form a temporary coalition; The multi-level incentive protocol module is used to dynamically allocate benefits based on the multi-dimensional contributions of the participants; The reputation evaluation module is used to collect and record historical cooperation behaviors and calculate the reputation score; The decentralized policy aggregation module is used to train the scheduling model based on local data and perform model aggregation to optimize the global data flow scheduling policy.

Citation Information

Cited By

  • Internet of vehicles federal learning excitation method and system based on Nash game and coalition game

    CN120751410A

  • A Federated Learning Incentive Method and System for Vehicle Networks Based on Nash Game Theory and Coalition Game Theory

    CN120751410B