Cooperative computing resource optimal configuration method based on MapReduce framework

By combining the MapReduce framework and collaborative computing system, using reinforcement learning and convex optimization algorithms to optimize equipment resources, the communication overhead and calculation accuracy problems of multi-client collaboration in complex wireless environments are solved, and efficient computing and energy saving and emission reduction under device heterogeneity are achieved.

CN120417047APending Publication Date: 2025-08-01SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510490625.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art cannot effectively handle the communication overhead and calculation accuracy balance during multi-client collaboration in complex wireless environments, and lacks an adaptive optimization mechanism for device heterogeneity such as computing power and power thresholds, making it difficult to meet the efficient computing needs of smart cities and other scenarios.

Method used

Combining the MapReduce framework and collaborative computing system, using reinforcement learning DDPG algorithm and convex optimization algorithm, an edge computing system model is established, equipment energy planning and resource optimization configuration is carried out, equipment energy consumption and resource allocation are optimized through reinforcement learning DDPG algorithm, and computing resources are optimized, and computing resources are optimized by combining convex optimization solver.

Benefits of technology

It significantly improves the system's data processing efficiency and energy efficiency, improves the processing capacity of equipment collaborative computing, and realizes the rational utilization of energy and the efficient operation of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120417047A_ABST
    Figure CN120417047A_ABST
Patent Text Reader

Abstract

The invention discloses a cooperative computing resource optimal configuration method based on a MapReduce framework. The method comprises the following steps: establishing an implementation scene and initializing the battery electric quantity and channel gain of equipment; at the beginning of each time slot, a reinforcement learning algorithm is utilized to solve an advance time energy planning problem according to the current state, and a computing resource optimization problem, an unloading decision problem and a transmission rate control problem are solved at an AP end at each time slot while the optimal equipment energy consumption is obtained; the optimal time allocation and the optimal transmission power decision are obtained by the equipment end, and the equipment energy and the channel information are updated by each piece of equipment. According to the method, the MapReduce framework is combined with the cooperative computing system, a system model of the edge computing system and the battery on the long-time scale is established, the used electric quantity can be planned in advance, and the purposes of saving energy, reducing emission and reasonably utilizing energy are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of wireless resource allocation, specifically a method for optimizing the allocation of collaborative computing resources based on the MapReduce framework. Background Art

[0002] With the rapid development of emerging fields such as the Internet of Things (IoT) and the industrial Internet, computing requirements exhibit three characteristics: cross-domain, heterogeneous, and real-time. Taking a smart city as an example, in scenarios such as traffic monitoring, environmental perception, and emergency dispatch, it is necessary to integrate multimodal data (video streams, sensor signals, geographical information, etc.) generated by millions of terminal devices and complete collaborative decision-making within millisecond-level latency, which poses a huge pressure on the computing capabilities of devices and networks. Traditional centralized cloud computing architectures are difficult to meet the stringent requirements of such scenarios due to problems such as long transmission latency, bandwidth bottlenecks, and data privacy risks. Summary of the Invention

[0003] In view of the deficiencies of the prior art in dealing with the communication overhead and the balance of computing accuracy during multi-client collaboration in a complex wireless environment, and lacking an adaptive optimization mechanism for device heterogeneity (such as computing power and battery threshold), the present invention proposes a method for optimizing the allocation of collaborative computing resources based on the MapReduce framework. By combining the MapReduce framework with a collaborative computing system, a system model of the edge computing system and the battery on a long time scale is established, which can plan the electricity consumption in advance to achieve the purpose of energy conservation and emission reduction and rational use of energy.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to a method for optimizing the allocation of collaborative computing resources based on the MapReduce framework, including:

[0006] Step 1: Establish an implementation scenario including N heterogeneous wireless devices and a shared common access point (AP) and initialize the battery power E bat,n,0 and the channel gain h n,0 .

[0007] Step 2: At the beginning of each time slot T, use the deep deterministic policy gradient (DDPG) algorithm of reinforcement learning to solve the problem of advance-time energy planning according to the current state s t to obtain the optimal device energy consumption {E cos,n,t}.

[0008] Step 3: According to the device energy consumption {E cos,n,t} obtained in Step 2, at each time slot T, solve the computing resource optimization problem, offloading decision problem, and transmission rate control problem at the AP side to obtain the optimal time allocation obtained by the device side and the optimal transmission power decision {p n,t}.

[0009] Step 4: At each time slot T, each device updates its device energy and channel information.

[0010] The present invention relates to a system for implementing the above method, including: a reinforcement learning end, a collaborative optimization end, heterogeneous devices, and a charging and discharging battery, wherein: the reinforcement learning end makes actions and sends data to the collaborative optimization end according to the perception of device information and environmental information, and also obtains corresponding rewards from the collaborative optimization end; the collaborative optimization end allocates resources for the devices according to the data received from the reinforcement learning end, sends instructions to the heterogeneous devices, and sends the calculated real-time rewards back to the reinforcement learning end; the device end makes reasonable resource allocation according to the instructions of the collaborative optimization end to obtain the final calculation result; the charging and discharging battery continuously collects renewable energy and transmits the battery state to the AP end in real time. Technical effects

[0011] The present invention aims at the scenario of combining MapReduce and collaborative computing, and uses an algorithm that combines reinforcement learning and convex optimization to solve resource allocation in collaborative computing; compared with other methods, it effectively improves the data processing efficiency of the system, and significantly improves energy efficiency and processing capacity through the collaborative computing of heterogeneous devices. Brief description of the drawings

[0012] Figure 1 is a flow chart of the present invention;

[0013] Figure 2 is a schematic diagram of the system scenario based on MapReduce collaborative computing;

[0014] Figure 3 is a schematic diagram of the DDPG-cvx algorithm structure;

[0015] Figure 4 is a flow chart for solving the optimization problem. Detailed implementation manners

[0016] As Figure 1 shown, this embodiment relates to a method for optimizing the configuration of collaborative computing resources based on the MapReduce framework, including:

[0017] Step 1: Establish a system scenario: As Figure 2 shown, it includes N heterogeneous wireless devices and a shared common access point (AP), and the set of collaborative computing tasks of N devices under the MapReduce framework where: the index of the time slot is t ∈ T := {1, 2,...., T}, the duration of each time slot is τ, the length of the common data ω is L bits, and the local data shared by each device through the access point is Its length is D t bits.

[0018] In the described system scenario, each device works collaboratively, resulting in a significant improvement in the overall computing power. That is, the L-bit common data ω is split into l-bit data ω for each device n ∈ N during time slot t n,t for each device n ∈ N during time slot t n,t , that is This collaborative work specifically includes:

[0019] 1) Map phase: Each device locally calculates intermediate key-value pairs. That is, device n calculates the intermediate value m based on the machine learning model ω and the l-bit shard data n (d1, ω n ), m n (d2, ω n ),..., m n (d N , ω n )), where: m n is the function executed by device n during the Map phase. This function calculates all intermediate results m n (d x , ω y ) for other devices where x ≠ y.

[0020] 2) Shuffle phase: The access point coordinates the exchange of intermediate key-value pairs between devices. That is, when the intermediate calculation results generated by device n during the Map phase need to be transmitted to device m, and the data volume βl n,t is proportional to the computing load ω n of device n, then device n transmits a total of (N - 1)βl n,t to other devices through the AP during the Shuffle phase, where: (N - 1)β = α, and α is generally much less than 1, that is, the data volume transmitted in this phase is less than the data volume calculated in the Map phase.

[0021] Preferably, each device uses a minimum bandwidth code for encoding during the Shuffle phase. When the device has performed q times the computing task, the communication load can be reduced to 1 / q of the original.

[0022] 3) Reduce phase: Each device aggregates the intermediate key-value pairs received during the Shuffle phase and then calculates the final result r n (m1(d n , ω1), m2(d n , ω2),..., m n (d n , ω N ))), where: r nIt is a function that device n executes on all intermediate results received during the Reduce stage.

[0023] Since the data processed in the Reduce stage comes from all devices, this stage must wait for all devices to complete the data transmission in the Shuffle stage before it can be started. Therefore, the reduction stage latency of all devices is the same, that is, it satisfies And increasing the computing time is beneficial to reducing energy consumption.

[0024] The time required for the Map, Shuffle, and Reduce stages of the described system scenario satisfies: where: τ is the length of a single time slot.

[0025] Step 2, establish an optimization problem based on the system computing model and transmission model, including:

[0026] The computing energy consumption of device n in the Map stage Similarly, the computing energy consumption of device n in the Reduce stage is where: c n is the number of CPU cycles required for device n to process one data, and k n is the capacitance coefficient depending on the chip architecture.

[0027] The constraint condition of the described computing energy consumption is: where: is the maximum CPU frequency of the device, which means there is an upper limit to the amount of data processed by the device, that is, the device has the maximum working ability when the CPU frequency is the highest. In addition, in the Map stage, each device performs inference with the l n,t -bit model ω n and local data Assuming L >> D, then only the data volume of the model ω n is considered, and since there are N data that need to be inferred N times, that is, the total data volume should be N times.

[0028] The transmission rate at which device n exchanges intermediate files through the access point AP in the Shuffle stage where: is the transmission power of device n, N0 is the receiver noise power of the base station, and B is the communication bandwidth. Since different edge servers (more specifically, the base stations they are associated with) operate on different and non-overlapping channels, and the selected clients in each channel use the orthogonal frequency division multiple access (OFDMA) protocol, there is no mutual interference between the base station and the device in the uplink.

[0029] Total energy consumption of all devices during the Shuffle stage Among them: is the communication circuit energy consumption of device n, is the maximum radio frequency transmission power of device n.

[0030] The constraint conditions of the total energy consumption model include:

[0031] Battery energy model, that is, the battery power queue set according to the battery power is E bat,n,t+1 = E bat,n,t + E get,n,t - E tot,n,t , where: E get,n,t is the renewable energy collected by device n within time slot t, E bat,n,t is the battery power of device n at the start of time slot t, and the total energy consumed by device n at time t

[0032] The constraint conditions of the battery energy model include: E bat,n,t+1 = min{E bat,n,t + E get,n,t - E tot,n,t , E max}, E bat,n,t + E get,n,t ≥ E tot,n,t , where: E max is the upper limit of the battery capacity.

[0033] The optimization problem mentioned above refers to: taking the total throughput of device collaboration computing as the optimization goal, and using battery energy, device power, computing power and time slot length as constraints, and pursuing the maximum average throughput on a long time scale of the system. Specifically: Among them: M is the total number of time slots for the system to work, n is the number of devices in the scenario, l is the amount of information processed by the device, t MAP , t SHu , t RED are the times consumed by the device in the Map, Shuffle, and Reduce stages respectively, and p is the power transmitted by the device in the Shuffle stage.

[0034] The constraint conditions of the optimization problem include: E bat,n,t+1 = min{E bat,n,t + E get,n,t - E tot,n,t , E max}, E bat,n,t + E get,n,t ≥ E tot,n,t ,

[0035] Step 3: Figure 4 As shown in Figure 2, the DDPG algorithm is used to transform the optimization problem proposed in step 2 into a single-time-slot subproblem, and the corresponding action is given according to the current system state, that is, the energy consumed by each device in the time slot, including:

[0036] 3.1 Construct a Markov decision process (MDP) framework suitable for reinforcement learning. As a Markov decision process (MDP), the optimization problem is represented by the triple (S, A, r) of the MDP state, action and reward.

[0037] The state S is: In time slot t, the system state can be defined as That is, is the battery power of the device at this time, is the channel gain, and together they can be used to determine the current overall system status.

[0038] The action A is: To reduce the action space and accelerate convergence, the action of the DDPG model in time slot t is designed to be

[0039] The reward is a function used to evaluate the quality of the action performed. n,t Set as instant reward, specifically in: It is a penalty term introduced when the constraint conditions are not met, used to reduce the reward r t , the input of the immediate reward is the action taken by the DDPG agent.

[0040] 3.2 Deploy the DDPG agent in the MapReduce framework to perceive the state s at any time slot t t And perform action a t After executing the action, the environment will feedback a scalar reward r t , and change the state from s t Transfer to s t+1 .

[0041] The action-value function of the agent is the mathematical expectation of the expected cumulative discounted reward in infinite time with the initial state s and the initial action a as the starting point under the strategy π: a], where: π(s t ) is to change the state s t Mapped to action a t strategy, γ is a discount factor used to weigh the importance of immediate rewards and future rewards.

[0042] The goal of the agent is to learn the optimal strategy Among them: the optimal action-value function represents the maximum expected cumulative reward under all policies.

[0043] 3.3 Use the DDPG model solving steps based on the actor-critic framework as Figure 3 shown to solve the optimization problem in step 2, that is, approximate the policy π through the actor network containing the target network and the online network, and approximate the action-value function Q π (s,a). Specifically: in each round of edge aggregation process, the DDPG agent collects the necessary state information s t from the environment, and the actor online network outputs the action a t . After that, based on a t optimize the remaining decision variables to calculate the reward value, and the agent obtains the reward r t at the end of the edge aggregation round t, and observes the new state s t+1 .

[0044] The actor online network π(s t |θ π ) and the critic online network Q(s t ,a t |θ Q ), where: θ π and θ Q are the model parameters of two deep neural networks (DNNs); the actor target network π ′ (s t |θ π′ ) and the critic target network Q ′ (s t ,a t |θ Q′ ), where: θ π′ and θ Q′ are the model parameters of two DNNs.

[0045] For the critic online network described above, its parameter update uses the minimization of the loss function where: y t = r t + γQ ′ (s t ,π ′ (s t |θ π′ )|θ Q′ ), M ′ is the mini-batch sample size. Optimize the actor online network parameters along the gradient direction to maximize the policy objective function J(θπ ) = E[Q(s t , π(s t |θ π )|θ Q , is the gradient of θ π .

[0046] For the actuator target network and the critic target network described above, their parameter updates adopt the soft update method, specifically: Where: is the parameter of the learning rate.

[0047] In the DDPG model described above, an experience replay buffer pool is preferably provided to store the quadruple experience {s t , a t , r t , s t+1} generated in each round of iteration. When the data in the buffer pool reaches the capacity threshold, the learning process is officially started: M ′ experiences are obtained through random sampling to form a mini-batch sample for training the DDPG network. Then, the critic online network is updated by minimizing the loss function, and at the same time, the actuator online network is updated by maximizing the policy objective function; after the optimization process is completed, the target network is further updated. By effectively suppressing the correlation between observation data and actively exploring different environmental states, the model can achieve convergence after hundreds of training cycles.

[0048] Step 4. The action a t output by the DDPG model can simplify the optimization problem in Step 2 to a situation without battery constraints. The sub-problem model for a single time slot is specifically: Its constraints include: Furthermore, a convex optimization solver is used to complete the subsequent solution, so as to obtain the time allocation and power allocation.

[0049] After specific actual experiments, considering a MapReduce system containing 5 heterogeneous wireless devices, in the specific environmental setting where these devices communicate through a central access point (AP), with the number of clients (N): 5, the edge server bandwidth (B): 15 KHz; the base station received noise power (N0): 1 nW / Hz; the channel gain from the device to the base station (h): CN(0, 10^-3); the effective capacitance coefficient (k n ): [10^-28, 10^-27]; the maximum transmission power [10, 25] mW; the constant energy consumption of the communication circuit [10, 25] mW; the maximum CPU frequency [1, 3] GHz; the number of CPU cycles required for each bit of data calculation (cn ):[500, 1500]; Batch size (M and M'): 32; Discount factor (γ): 0.95; DDPG experience replay memory size: 4000; DDPG actor network learning rate: 0.001; DDPG critic network learning rate: 0.002.

[0050] As Figure 4 shown, based on the above settings, the simulation is implemented in the following way:

[0051] Initialization: Initialize the system state, including the initial power and the initial channel state;

[0052] Advance energy planning: At the beginning of each time slot nT, determine the optimal planned energy amount E based on the current power and channel conditions cos,n,t ;

[0053] Real-time energy balance: Given E cos,n,t , in each time slot t ∈ [nT, (n + 1)T - 1], by solving the simplified optimization problem, the optimal time allocation obtained by the device side and the optimal transmission power decision {p n,t};

[0054] State update: Update the system state according to the channel conditions and the collected energy.

[0055] Finally, the experimental results are shown in Table 1.

[0056] Table 1 Experimental method Experimental result / bit Reinforcement learning + convex optimization 186687 Reinforcement learning 166684 Greedy algorithm 171830 Equal distribution 121086

[0057] Compared with the prior art, the present method uses an algorithm that combines reinforcement learning and convex optimization, achieving effective management of energy, and enabling better experimental results than other algorithms on a long time scale.

[0058] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.

Claims

1. A collaborative computing resource optimization configuration method based on the MapReduce framework, characterized in that Establish an implementation scenario including N heterogeneous wireless devices and a shared common access point (AP), and initialize the battery power and channel gain of the devices; at the beginning of each time slot: use the deep deterministic policy gradient (DDPG) algorithm of reinforcement learning to solve the advance-time energy planning problem according to the current state. After obtaining the optimal device energy consumption, solve the computing resource optimization problem, offloading decision problem, and transmission rate control problem at the AP side, obtain the optimal time allocation and optimal transmission power decision for the three stages of collaborative work at the device side, and then update the device energy and channel information through each device.

2. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 1, characterized in that The three stages of the collaborative work include: 1) Map Phase: Each device locally computes intermediate key-value pairs, i.e., device n computes intermediate value m based on the machine learning model ω and l-bit shard data n (d1, ω n ), m n (d2, ω n ),..., m n (d N , ω n ), where: m n is the function executed by device n in the Map phase, and this function computes all intermediate results m n (d x , ω y ) for which x ≠ y for other devices. 2) Shuffle Phase: The access point coordinates the exchange of intermediate key-value pairs between devices. That is, when the intermediate calculation result generated by device n in the Map phase needs to be transmitted to device m, the data volume βl n,t is proportional to the computing load ω of device n n , then device n transmits a total of (N - 1)βl to other devices through the AP in the Shuffle phase n,t , where: (N - 1)β = α, and α is generally much less than 1, that is, the data volume transmitted in this phase is less than the data volume calculated in the Map phase. 3) Reduce Phase: After aggregating the intermediate key-value pairs received in the Shuffle phase, each device calculates the final result r n (m1(d n , ω1), m2(d n , ω2),..., m n (d n , ω N ))), where: r n is the function that device n executes on all the intermediate results received during the Reduce phase. The time required for the Map, Shuffle, and Reduce stages of the described system scenario Satisfy: Where: τ is the length of a single time slot.

3. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 2, characterized in that, The aforementioned early - time energy planning problem, that is, taking the total throughput calculated by device cooperation as the optimization objective, with battery energy, device power, computing ability, and time - slot length as constraints, and pursuing the maximum average throughput on a long - time scale of the system, specifically: where: M is the total number of time - slots for the system to work, n is the number of devices in the scenario, l is the amount of information processed by the device, t MAP t SHU t RED is the time consumed by the device in the Map, Shuffle, and Reduce stages, and p is the power emitted by the device in the Shuffle stage; The constraint conditions of the optimization problem include: E bat,n,t+1 = min{E bat,n,t + E get,n,t - E tot,n,t , E max}, E bat,n,t + E get,n,t ≥ E tot,n,t , where: the computing energy consumption of device n in the Map stage the computing energy consumption of device n in the Reduce stage where: c n is the number of CPU cycles required for device n to process one data, and k n is the capacitance coefficient dependent on the chip architecture; The transmission rate of the intermediate file exchanged by the device n through the access point AP during the Shuffle phase Where: is the transmission power of the device n, N0 is the receiver noise power of the base station, and B is the communication bandwidth. Since different edge servers (more specifically, the base stations associated with them) operate on different and non-overlapping channels, and the selected clients in each channel adopt the orthogonal frequency division multiple access (OFDMA) protocol, there is no mutual interference between the base station and the device in the uplink; Total energy consumption of all devices during the Shuffle phase Among them: is the communication circuit energy consumption of device n, is the maximum radio frequency transmission power of device n; The battery energy model, i.e., the power queue set according to the battery power, is E bat,n,t+1 = E bat,n,t + E get,n,t - E tot,n,t , where: E get,n,t is the renewable energy collected by device n within time slot t, and E bat,n,t is the battery power of device n at the start of time slot t, and the total energy consumed by device n at time t The constraint conditions for calculating energy consumption are as follows: Among them: is the maximum CPU frequency of the device, which means there is an upper limit to the amount of data processed by the device. That is, when the device's CPU frequency is the highest, it has the maximum working ability. In addition, during the Map stage, each device has a model ω of l n,t bits n and local data for inference. Assuming L >> D, only the data volume of the model ω n is considered, and since there are N data that need to be inferred N times, that is, the total data volume should be N times; The constraint conditions of the total energy consumption model described above include: The constraint conditions of the battery energy model include: E bat,n,t+1 = min{E bat,n,t + E get,n,t - E tot,n,t , E max}, E bat,n,t + E get,n,t ≥ E tot,n,t , where: E max is the upper limit of the battery capacity.

4. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 1 or 3, characterized in that, The solution of the advance-time energy planning problem includes: using the DDPG algorithm to transform the optimization problem into a sub-problem of a single time slot, and giving corresponding actions according to the current system state, that is, the energy consumed by each device in this time slot.

5. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 1 or 3, characterized in that The solution of the advance-time energy planning problem specifically includes: 3.1 Construct a Markov decision process (MDP) framework suitable for reinforcement learning. As a Markov decision process (MDP), represent the optimization problem through the triple (S, A, r) of the state, action, and reward of the MDP; 3.2 Deploy the DDPG agent in the MapReduce framework, sense the state s at any time slot t t and execute the action a t ; After executing the action, the environment will feedback a scalar reward r t , and transfer the state from s t to s t+1 ; 3.3 The DDPG model based on the actor-critic framework is used to solve the optimization problem, that is, the policy π is fitted through the actor network including the target network and the online network, and the action value function Q(s,a) is approximated through the critic network also including the target network and the online network. Specifically, in each round of edge aggregation, the DDPG agent collects the necessary state information s from the environment. π After the actor online network outputs the action a, based on a, the remaining decision variables are optimized to calculate the reward value. The agent obtains the reward r at the end of the edge aggregation round t t and observes the new state s'. t t t t+1 ​ 6. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 5, characterized in that The state S mentioned above means that within time slot t, the system state can be defined as That is, the battery power of the device at this time, which is the channel gain. Overall, the current overall state of the system can be obtained. The described Action A refers to: To reduce the action space and accelerate convergence, the action of the DDPG model in time slot t is designed as The described reward means that the reward function is used to evaluate the quality of the executed action; preferably, the instantaneous objective function l in the optimization problem is set as the immediate reward, specifically n,t Let it be the immediate reward, specifically where: is the penalty term introduced when the constraint condition is not satisfied, used to reduce the reward r t , and the input of the immediate reward is the action taken by the DDPG agent; The action-value function of the agent, that is, under the policy π, starting from the initial state s and the initial action a, the mathematical expectation of the expected cumulative discounted reward within an infinite time is specifically: Where: π(s t ) is the policy that maps the state s t to the action a t , γ is the discount factor used to balance the importance of immediate rewards and future rewards; The goal of the agent, which is to learn the optimal policy where: the optimal action-value function represents the maximum expected cumulative reward under all policies; The online network of the actuator π(s t |θ π ) and the online network of the critic Q(s t ,a t |θ Q ), where: θ π and θ Q are the model parameters of two deep neural networks (DNNs); the target network of the actuator π ′ (s t |θ π′ ) and the target network of the critic Q ′ (s t ,a t |θ Q′ ), where: θ π′ and θ Q′ are the model parameters of two DNNs; For the described online network of the evaluator, its parameter update uses minimizing the loss function where: y t = r t + γQ ′ (s t , π ′ (s t | θ π′ ) | θ Q′ ) , M ′ is the mini - batch sample size; by optimizing the online network parameters of the actuator along the gradient direction to maximize the policy objective function J(θ π ) = E[Q(s t , π(s t | θ π ) | θ Q , is the gradient of θ π ; For the described actuator target network and discriminator target network, their parameter updates adopt a soft update method, specifically: Where: is the parameter of the learning rate; The preferred DDPG model is provided with an experience replay buffer pool to store the quadruple experience {s t , a t , r t , s t+1} generated in each round of iteration; when the data in the buffer pool reaches the capacity threshold, the learning process is officially started: obtain M ′ experiences through random sampling to form a mini-batch sample for training the DDPG network, then update the online critic network by minimizing the loss function, and at the same time update the online actor network by maximizing the policy objective function; further update the target network after the optimization process is completed; by effectively suppressing the correlation between observation data and actively exploring different environmental states, the model can achieve convergence after hundreds of training cycles.

7. The collaborative computing resource optimization configuration method based on the MapReduce framework according to claim 4, characterized in that The AP-side solution mentioned above refers to: the action a output by the DDPG model t The optimization problem is simplified to a case without battery constraints. The sub-problem model for a single time slot is specifically as follows: Its constraints include: Subsequently, a convex optimization solver is used to complete the subsequent solution, thereby obtaining the time allocation and power allocation.

8. A collaborative computing resource optimization configuration system for implementing the method according to any one of claims 1-7, characterized in that including: A reinforcement learning end, a collaborative optimization end, heterogeneous devices, and a charging and discharging battery. Among them: the reinforcement learning end makes actions and sends data to the collaborative optimization end according to the perception of device information and environmental information, and also obtains corresponding rewards from the collaborative optimization end; the collaborative optimization end will perform resource allocation for the devices according to the data received from the reinforcement learning end, send instructions to the heterogeneous devices, and send the calculated real-time rewards back to the reinforcement learning end; The device side performs reasonable resource allocation according to the instructions of the collaborative optimization end to obtain the final calculation result; the charging and discharging battery continuously collects renewable energy and transmits the battery state to the AP side in real time.