Multi-agent collaborative learning heterogeneous fusion network resource scheduling method and device

The heterogeneous fusion network resource scheduling method based on multi-agent collaborative learning utilizes state observation and task execution action information to rationally schedule and allocate network resources, solving the problem of insufficient resource allocation in existing technologies and improving task completion rate and resource utilization.

CN116647564BActive Publication Date: 2025-10-21BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310564610.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-10-21
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

Existing network resource scheduling methods cannot effectively schedule and allocate network resources for multi-agent systems, resulting in insufficient computing service capabilities to meet the needs of latency-sensitive and resource-intensive tasks.

Method used

The heterogeneous fusion network resource scheduling method based on multi-agent collaborative learning utilizes the state observation information and task execution action information of each agent to determine the value information of task execution actions, and generates task execution instructions based on contribution information, so as to rationally schedule and allocate network resources.

Benefits of technology

It improves the task completion rate of multi-agent systems, maximizes the utilization of network resources and optimizes agent performance, reduces task latency, and improves server resource utilization and task completion rate of network nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116647564B_ABST
    Figure CN116647564B_ABST
Patent Text Reader

Abstract

The application provides a multi-agent collaborative learning heterogeneous fusion network resource scheduling method and device, which comprises the following steps: determining the task execution action value information corresponding to each agent according to the state observation information and task execution action information corresponding to each agent; determining the contribution information corresponding to each task in the multiple tasks executed by the agent corresponding to the task execution action value information according to the task execution action value information; generating a task execution instruction according to the target task corresponding to the target contribution information in all contribution information, wherein the task execution instruction is used for instructing the agent to execute the target task; and sending multiple task execution instructions to the corresponding agents. The method can effectively and reasonably schedule and allocate the tasks corresponding to each agent according to the contribution information corresponding to each task in the multiple tasks executed by the agent, so as to improve the task completion rate of the multi-agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a heterogeneous fusion network resource scheduling method and device for multi-agent collaborative learning. Background Art

[0002] In recent years, with the rapid development of advanced communication technologies and the continuous improvement of the mobile application ecosystem, intelligent agents have become increasingly latency-sensitive, resource-intensive, and data-intensive. Furthermore, the rapid growth of the Internet of Things has led to the large-scale deployment of intelligent agents in various life scenarios, continuously performing tasks such as collecting environmental information such as images and / or videos. However, due to the limited battery storage, computing power, and transmission capabilities of intelligent agents, the computing services they can provide are often unable to meet the large-scale computing needs and timely response requirements. In this way, service devices can rationally schedule and allocate various network resources to maximize network resource utilization and optimize intelligent agent performance.

[0003] Existing network resource scheduling methods often use deep reinforcement learning algorithms or reinforcement learning models. However, due to certain limitations of the deep reinforcement learning algorithms and reinforcement learning models, service equipment is unable to effectively schedule and allocate network resources corresponding to multiple agents. Summary of the Invention

[0004] The present invention provides a heterogeneous fusion network resource scheduling method and device for multi-agent collaborative learning, which can utilize the contribution information corresponding to each task in the multiple tasks performed by each agent to effectively and reasonably schedule and allocate the tasks corresponding to each agent, and thus effectively and reasonably schedule and allocate the heterogeneous fusion network resources corresponding to the multiple agents, so as to improve the task completion rate of the multiple agents.

[0005] The present invention provides a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning, comprising:

[0006] Determine the task execution action value information corresponding to each intelligent agent based on the state observation information and task execution action information corresponding to each intelligent agent;

[0007] For each task execution action value information, determining, based on the task execution action value information, contribution information corresponding to each of the multiple tasks executed by the agent corresponding to the task execution action value information;

[0008] Generate a task execution instruction based on the target task corresponding to the target contribution information in all contribution information, and the task execution instruction is used to instruct the agent to perform the target task;

[0009] Send multiple task execution instructions to the corresponding agents

[0010] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, the task execution action value information corresponding to each intelligent agent is determined based on the state observation information and task execution action information corresponding to each of the multiple intelligent agents, including: for each intelligent agent, obtaining the indicator attributes corresponding to the intelligent agent when executing the multiple tasks, the relevant state information corresponding to the intelligent agent, and the task execution action information; determining the state observation information corresponding to the intelligent agent based on the network space state information corresponding to the indicator attributes and the relevant state information; and determining the task execution action value information corresponding to the intelligent agent based on the state observation information and the task execution action information.

[0011] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, the task execution action value information corresponding to the agent is determined based on the state observation information and the task execution action information, including: determining the first convolution embedding result corresponding to the agent based on the state observation information and the task execution action information; determining the influence parameters generated by other agents among the multiple agents except the agent when the agent executes the multiple tasks; and determining the task execution action value information corresponding to the agent based on the first convolution embedding result and the influence parameters.

[0012] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, the determining of the influence parameters generated by other agents among the multiple agents except the agent when the agent performs the multiple tasks includes: determining the second convolution embedding results corresponding to each of the other agents among the multiple agents except the agent; for each other agent, using an activation function, performing a linear transformation on the second convolution embedding results corresponding to the other agents to obtain a target embedding result; and determining the influence parameters generated by all other agents when the agent performs the multiple tasks based on all target embedding results and the weight matrices corresponding to all target embedding results.

[0013] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, the contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information is determined based on the task execution action value information, including: for each task, determining the selection probability of the agent for the task; based on the selection probability and the task execution action value information, determining the contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information.

[0014] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, determining the selection probability of the agent for the task includes: determining the third convolution embedding results corresponding to the agent and the other agents based on the influence relationship between the agent and the other agents; determining the interaction information according to the first convolution embedding result and the third convolution embedding result corresponding to the agent; normalizing the distribution of the state observation information to obtain allocation information; and determining the selection probability of the agent for the task according to the interaction information and the allocation information.

[0015] According to a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, a task execution instruction is generated based on a target task corresponding to a target contribution information in all contribution information, including: determining the maximum contribution information among all the contribution information as the target contribution information; or determining the contribution information among all the contribution information that is greater than a preset contribution threshold as the target contribution information; and generating the task execution instruction based on the target task corresponding to the target contribution information.

[0016] The present invention also provides a heterogeneous fusion network resource scheduling device for multi-agent collaborative learning, comprising:

[0017] a processing module configured to determine, based on state observation information and task execution action information corresponding to each of the plurality of agents, task execution action value information corresponding to each agent; determine, for each task execution action value information, contribution information corresponding to each of the plurality of tasks executed by the agent corresponding to the task execution action value information; and generate a task execution instruction based on a target task corresponding to target contribution information in all contribution information, the task execution instruction being used to instruct the agent to execute the target task;

[0018] The transceiver module is used to send multiple task execution instructions to the corresponding intelligent agent.

[0019] The present invention also provides a service device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning as described in any one of the above.

[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning as described in any one of the above.

[0021] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the heterogeneous fusion network resource scheduling methods for multi-agent collaborative learning as described above.

[0022] The present invention provides a heterogeneous converged network resource scheduling method and device for multi-agent collaborative learning. The method and device determine the task execution action value information corresponding to each agent based on the state observation information and task execution action information corresponding to each of the multiple agents. For each task execution action value information, the method determines the contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information. Based on the target task corresponding to the target contribution information in all the contribution information, the method generates a task execution instruction, which instructs the agent to perform the target task. The multiple task execution instructions are then sent to the corresponding agents. This method utilizes the contribution information corresponding to each of the multiple tasks performed by each agent to effectively and reasonably schedule and allocate the tasks corresponding to each agent. This method can effectively and reasonably schedule and allocate heterogeneous converged network resources corresponding to multiple agents, thereby improving the task completion rate of multiple agents, maximizing the utilization of network resources, and optimizing agent performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 This is a scenario diagram of the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention;

[0025] Figure 2 This is a flow chart of a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of the heterogeneous fusion network resource scheduling device for multi-agent collaborative learning provided by the present invention;

[0027] Figure 4 It is a structural diagram of the service equipment provided by the present invention. DETAILED DESCRIPTION

[0028] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0029] like Figure 1 The figure is a schematic diagram of the scenario of the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention. Figure 1 In the embodiment, the service device 10 and the agent 20 are connected via wireless communication technology. When the service device 10 is a cloud server, the agent 20 is an edge server and / or an electronic device; when the service device 10 is an edge server, the agent 20 is an electronic device.

[0030] The service device 10 refers to a radio transceiver station that transmits data to the intelligent agent 20 via wireless communication technology within a certain radio coverage area.

[0031] Optionally, the number of service devices 10 is not limited. Figure 1 In the example, the number of service devices 10 is 1.

[0032] The agent 20 refers to a device that can perform data transmission with the service device 10 via wireless communication technology.

[0033] Optionally, the number of agents 20 is unlimited. Figure 1 In the figure, the number of agents 20 is 3, namely agent c1, agent c2 and agent c3.

[0034] Optionally, when the agent 20 is an electronic device, the electronic device may include but is not limited to: a computer, a mobile terminal, and a wearable device.

[0035] Optionally, the wireless communication technology may include but is not limited to one of the following: the 4th Generation mobile communication technology (4G), the 5th Generation mobile communication technology (5G) and Wireless Fidelity (WiFi), etc.

[0036] It should be noted that the execution subject involved in the embodiment of the present invention can be a heterogeneous fusion network resource scheduling device for multi-agent collaborative learning, or it can be a service device. The embodiment of the present invention is further explained below using the service device as an example.

[0037] like Figure 2 FIG. 1 is a flow chart of a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the present invention, which may include:

[0038] 201. Determine the task execution action value information corresponding to each intelligent agent based on the state observation information and task execution action information corresponding to each of the multiple intelligent agents.

[0039] Among them, the number of agents is N, N≥2.

[0040] State observation information refers to the information composed of the cyberspace state information and related state information corresponding to the agent, which can be represented by O, where O = (o1,…o i ,…,o N ), o i Represents the state observation information corresponding to the i-th agent among multiple agents.

[0041] Cyberspace status information refers to the network operation status that can reflect the intelligent agent when performing tasks; related status information refers to the hardware operation status corresponding to the intelligent agent itself excluding the network operation status. Optionally, the related status information may include but is not limited to: posture information and filter gain information, etc.

[0042] Task execution action information refers to the information related to the next execution task performed by the agent, which can be represented by A, where A = (a1,…a i ,…,a N ), a i Represents the task execution action information corresponding to the above-mentioned i-th agent.

[0043] Task execution action value information refers to the information used to measure the execution value corresponding to the task execution action information, which can be expressed as Q(O,A), where Q i (O, A) represents the task execution action value information corresponding to the above-mentioned i-th agent.

[0044] For the i-th agent among multiple agents, the agent can first obtain its corresponding state observation information and task execution action information; then, the agent sends this state observation information and task execution action information to the service device. After receiving the state observation information and task execution action information sent by the agent, the service device can determine the corresponding task execution action value information for the agent based on this state observation information and task execution action information. Based on this, the service device can ultimately determine the same amount of task execution action value information as the number of agents, allowing for the subsequent accurate determination of the corresponding contribution information of each agent when performing multiple tasks.

[0045] In some embodiments, the service device determines the task execution action value information corresponding to each agent based on the state observation information and task execution action information corresponding to each of the multiple agents, which may include but is not limited to at least one of the following implementation methods:

[0046] Implementation method 1: For each intelligent agent, the service device obtains the indicator attributes corresponding to the intelligent agent when performing multiple tasks, the relevant status information corresponding to the intelligent agent, and the task execution action information; the service device determines the state observation information corresponding to the intelligent agent based on the cyberspace state information and relevant status information corresponding to the indicator attributes; the service device determines the task execution action value information corresponding to the intelligent agent based on the state observation information and the task execution action information.

[0047] Among them, the task refers to the next task execution action selected by the agent, and the number of tasks is indivual, Multiple tasks can form a task queue.

[0048] Optionally, tasks may include: face recognition, voice recognition, and data transmission.

[0049] Indicator attributes refer to parameters that measure the state changes generated when an agent performs multiple tasks.

[0050] Optionally, the indicator attributes may include but are not limited to at least one of the following: bandwidth resources occupied by the task, computing resources occupied by the task, and task latency, etc.

[0051] For the i-th intelligent agent among N intelligent agents, the service device can first obtain the indicator attributes corresponding to the intelligent agent when performing multiple tasks, the relevant state information corresponding to the intelligent agent, and the task execution action information corresponding to the intelligent agent; then, the service device determines the cyberspace state information corresponding to the intelligent agent based on the indicator attributes; then, the service device constructs the state observation information corresponding to the intelligent agent based on the cyberspace state information and the relevant state information, and then combines the task execution action information to accurately obtain the task execution action value information corresponding to the intelligent agent.

[0052] It should be noted that there is no limit to the timing of the service device obtaining the indicator attributes, related status information and task execution action information.

[0053] Optionally, when the indicator attribute is the bandwidth resources occupied by the task, the service device obtains the indicator attribute corresponding to the agent performing multiple tasks, which may include: the service device determines the bandwidth resources occupied by the corresponding tasks when the agent performs multiple tasks according to the first formula.

[0054] Among them, the first formula is:

[0055] B i (t) represents the execution of the i-th agent at time t When there are tasks, the corresponding tasks occupy bandwidth resources; l represents The lth task among the tasks; B(t) represents the bandwidth resources occupied by the corresponding task when the i-th agent performs the lth task at time t.

[0056] Based on the first formula, For each task, the service device can first determine the bandwidth resources occupied by the corresponding task when the i-th agent performs each task, so that The task occupies bandwidth resources; then, the service device The average value of bandwidth resources occupied by tasks is determined as the execution of the i-th agent. When a task is executed, the corresponding task occupies bandwidth resources.

[0057] Optionally, when the indicator attribute is the computing resources occupied by the task, the service device obtains the indicator attribute corresponding to the intelligent agent performing multiple tasks, which may include: the service device determines the computing resources occupied by the corresponding tasks when the intelligent agent performs multiple tasks according to the second formula.

[0058] Among them, the second formula is:

[0059] C i (t) represents the execution of the i-th agent at time t C(t) represents the computing resources occupied by the corresponding task when the i-th agent performs the l-th task at time t.

[0060] Based on the second formula, The service device can first determine the computing resources occupied by each task when the i-th agent performs each task, so that The task occupies computing resources; then, the service device The average value of computing resources occupied by tasks is determined as the execution of the i-th agent. When a task is executed, the corresponding task occupies computing resources.

[0061] Optionally, when the indicator attribute is task delay, the service device obtains the indicator attribute corresponding to the agent performing multiple tasks, which may include: the service device determines the task delay corresponding to the agent performing multiple tasks according to the third formula.

[0062] Among them, the third formula is:

[0063] D i (t) represents the execution of the i-th agent at time t D(t) represents the task delay corresponding to the i-th agent performing the l-th task at time t.

[0064] Based on the third formula, For each task, the service device can first determine the task delay corresponding to each task performed by the i-th agent, so that Then, the service device will The mean value of the task delay corresponding to the i-th agent execution is determined as The task delay corresponding to each task.

[0065] In summary, for the i-th agent among N agents, the indicator attributes include the task occupying bandwidth resources B i (t), the computing resources occupied by the task C i (t) and task delay D i (t) In the case of this indicator attribute, the cyberspace state information corresponding to this indicator attribute can be defined as At this time, for the network space status information corresponding to each of the N intelligent agents, the service device can obtain the network topology information corresponding to the N intelligent agents, which can be used It indicates that the network topology information can be used to implement effective network resource scheduling decisions in the future.

[0066] Implementation method 2: The service device determines the first convolution embedding result corresponding to the intelligent agent based on the state observation information and the task execution action information; the service device determines the impact parameters generated by other intelligent agents among the multiple intelligent agents except the intelligent agent when the intelligent agent performs multiple tasks; the service device determines the task execution action value information corresponding to the intelligent agent based on the first convolution embedding result and the impact parameters.

[0067] Among them, the influence parameter refers to the degree of influence of other agents on the agent when the agent performs multiple tasks, which can be expressed as z i express.

[0068] The first convolution embedding result refers to the embedding result obtained by the service device based on the single-layer graph convolutional neural network for the i-th agent. The first convolution embedding result corresponding to the i-th agent can be obtained by c i (o i ,a i )express.

[0069] For the i-th agent among N agents, the service device can first input the state observation information and task execution action information corresponding to the agent into a single-layer graph convolutional neural network to obtain the first convolution embedding result corresponding to the agent output by the single-layer graph convolutional neural network; then, the service device detects the influence of other agents on the agent's execution of multiple tasks one by one to obtain the influence parameters of all other agents on the agent's execution of the multiple tasks; finally, the service device accurately determines the task execution action value information corresponding to the agent based on the first convolution embedding result and the influence parameter

[0070] It should be noted that the service device is not limited to the timing of determining the first convolution embedding result and the influencing parameters.

[0071] In summary, it should be noted that the service device may pre-model the network resource scheduling process as a Markov decision process, in which the state space, action space and reward function may be defined in detail.

[0072] Among them, the state space can be the network topology information corresponding to all the above network space state information, that is,

[0073] The action space is determined by the service device according to the definition of resource scheduling, based on the request and recovery of computing resources and broadband resources by each agent, that is, based on the computing resource utilization and broadband resource utilization. The computing resource utilization can be expressed as R C Indicates that the broadband resource utilization can be expressed as R B express.

[0074] Since resource allocation is continuous, the resource utilization R is calculated C ∈[Q C1 ,Q C2 ], Q C1 <Q C2 , Q C1 represents the first computing resource threshold, Q C2 represents the second computing resource threshold; broadband resource utilization R B ∈[Q B1 ,Q B2 ], Q B1 <Q B2 , Q B1Indicates the first broadband resource threshold, Q B2 Indicates the second broadband resource threshold.

[0075] For example, R C ∈[-Q C2 ,Q C2 ], namely Q C1 =-Q C2 ; R B ∈[-Q B2 ,Q B2 ], namely Q B1 =-Q B2 .

[0076] In addition, the service device also needs to adjust the maximum task delay threshold and the task queue fluctuation length threshold. The maximum task delay threshold can be used MAX Indicates that the task queue fluctuation length threshold can be used U MAX express.

[0077] In summary, the action space can be a={R C ,Q C ,D MAX ,U MAX}express.

[0078] Since rewards are feedback information given to the agent during the network resource scheduling process, the service device determines the reward function based on the computing resource utilization, bandwidth resource utilization, average task latency, and task completion rate to obtain rewards. The reward function is: Indicates reward; D C represents the average task delay, σ represents the average task delay D C The corresponding first preset parameter; M C represents the task completion rate, α represents the task completion rate M C The corresponding second preset parameter.

[0079] The first preset parameter σ and the second preset parameter α may be factory-set for the service device or user-defined, and are not specifically limited here.

[0080] Optionally, the service device determines the task execution action value information corresponding to the intelligent agent based on the first convolution embedding result and the influencing parameter, which may include: the service device obtains the task execution action value information corresponding to the intelligent agent based on the action value function.

[0081] Among them, the action value function is: Q i (O,A)=M i (c i (o i ,a i),z i );

[0082] M i (·) represents the embedding function obtained by the service device based on two multi-layer perceptrons based on the i-th agent.

[0083] The service device inputs the first convolution embedding result and the influencing parameters into two multi-layer perceptrons for calculation, and can accurately obtain the task execution action value information Q corresponding to the i-th agent. i (O,A).

[0084] In some embodiments, the service device determines the impact parameters of other intelligent agents among multiple intelligent agents except the intelligent agent when the intelligent agent performs multiple tasks, which may include: the service device determines the second convolution embedding results corresponding to each of the other intelligent agents except the intelligent agent among the multiple intelligent agents; the service device uses the activation function to perform a linear transformation on the second convolution embedding results corresponding to the other intelligent agents for each other intelligent agent to obtain a target embedding result; the service device determines the impact parameters of all other intelligent agents when the intelligent agent performs multiple tasks based on all target embedding results and the weight matrices corresponding to all target embedding results.

[0085] It should be noted that, for the i-th intelligent agent among N intelligent agents, since j≠i, the j-th intelligent agent among the N intelligent agents is the other intelligent agent corresponding to the i-th intelligent agent.

[0086] The second convolution embedding result refers to the embedding result obtained by the service device based on the single-layer graph convolutional neural network for the j-th agent. The second convolution embedding result corresponding to the j-th agent can be obtained by c j (o j ,a j ) means, j Represents the state observation information corresponding to the j-th agent, a j Represents the task execution action information corresponding to the j-th agent; the target embedding result corresponding to the j-th agent can be used j express.

[0087] The activation function refers to the function that runs on the neurons of the artificial neural network, which is responsible for linearly transforming the input results of the neurons and mapping them to the output end.

[0088] Optionally, the activation function may include: a ReLU function.

[0089] The weight matrix refers to the proportion of the embedding results. The weight matrices corresponding to all target embedding results can be the same or different, and there is no specific limitation here.

[0090] For the i-th agent among multiple agents, that is, for the j-th agent, the service device can first determine the second convolutional embedding result corresponding to the j-th agent. Then, the service device inputs this second convolutional embedding result into the activation layer, and in this activation layer, uses the activation function to perform a linear transformation on this second convolutional embedding result to obtain the target embedding result corresponding to the j-th agent. In this way, the service device will determine as many target embedding results as there are other agents. Next, the service device determines the weight matrix corresponding to each target embedding result, and combines all the target embedding results to determine the impact parameters of all other agents on the i-th agent's execution of multiple tasks.

[0091] Optionally, the service device determines the influence parameters of all other intelligent agents on the intelligent agent's execution of multiple tasks based on all target embedding results and the weight matrices corresponding to all target embedding results, which may include: the service device obtains the influence parameters of all other intelligent agents on the intelligent agent's execution of multiple tasks based on the influence parameter formula.

[0092] Among them, the influencing parameter formula is: i =∑ j≠i α j v j =∑ j≠i α j r(Sc j (o j ,a j ));

[0093] α j Indicates the target embedding result v corresponding to the jth agent j The corresponding weight matrix; S represents the preset shared matrix; r(·) represents the ReLU activation function.

[0094] Based on the influence parameter formula, the service device uses the activation function and the shared matrix to accurately obtain the influence parameters of all other agents on the i-th agent when performing multiple tasks.

[0095] Optionally, the target embedding result v corresponding to the jth agent j The corresponding weight matrix α j It can be determined according to the weight formula, where the weight formula is:

[0096] The first convolution embedding result e i =c i (o i ,a i ), the second convolution embedding result e j =c j (o j,a j ); B1 represents the second convolution embedding result e j The corresponding first preset weight parameter; B2 represents the first convolution embedding result e i The corresponding second preset weight parameter.

[0097] Based on the weight formula, the service device can use the first preset weight parameter B1 to embed the second convolution into the result e j Convert it into query value (query), and use the second preset weight parameter B2 to embed the first convolution into the result e i Converted into a key value (key); then, the service device matches and scales the dimensions of the query value and key value matrices to obtain the weight matrix α j , to prevent the gradient from disappearing.

[0098] In summary, the service device can adopt a multi-head attention mechanism. Based on the influence parameter formula and the weight formula, each head uses a separate set of parameters (such as the convolution embedding result) to generate the aggregate similarity from all other intelligent agents. Then, the output results of all heads are simply spliced ​​together to generate a single vector as the aggregate similarity, and then a weighted sum is performed on the data of different intelligent agents to obtain an influence parameter with higher accuracy.

[0099] That is to say, the service device adopts a multi-head attention mechanism to update the value network in the service device (the value network may include a single-layer graph convolutional neural network and two multi-layer perceptrons), so that the service device can take into account the task execution action information of all intelligent agents, realize centralized training and decentralized execution, and at the same time, also consider the interaction between multiple agents in resource scheduling, and can dynamically select which agents to pay attention to at each time point in the training process, thereby improving the performance of multi-agent resource scheduling with complex interactions.

[0100] 202. For each task execution action value information, determine the contribution information corresponding to each task among the multiple tasks executed by the intelligent agent corresponding to the task execution action value information based on the task execution action value information.

[0101] Among them, the contribution information refers to the contribution of the task execution action of the i-th agent to the global reward based on the centralized evaluator as the baseline of the i-th agent. The contribution information can be expressed as A i (o i ,a i )express.

[0102] Based on the action value information for each agent's task execution, the service device can determine the contribution information corresponding to each of the multiple tasks performed by each agent. In this way, the contribution information obtained through multi-agent collaboration can maximize the global reward.

[0103] In some embodiments, the service device determines the contribution information corresponding to each of the multiple tasks performed by the intelligent agent corresponding to the task execution action value information based on the task execution action value information, which may include: the service device determines the corresponding selection probability when the intelligent agent performs the task for each task; the service device determines the contribution information corresponding to each of the multiple tasks performed by the intelligent agent based on the selection probability and the task execution action value information.

[0104] The selection probability refers to the probability that the agent chooses to perform the task, which can be expressed as π(a i,j ∣o i )express.

[0105] For any of the multiple tasks, the service device can first determine the agent's corresponding selection probability when performing the task. Then, combined with the previously determined task execution action value information, it can determine the agent's corresponding contribution information when performing the task. In other words, the service device can determine as many selection probabilities, and thus contribution information, as there are tasks.

[0106] Optionally, the service device determines the contribution information corresponding to each of the multiple tasks performed by the intelligent agent corresponding to the task execution action value information based on the selection probability and the task execution action value information, which may include: the service device obtains the contribution information corresponding to each of the multiple tasks performed by the intelligent agent corresponding to the task execution action value information based on the contribution formula.

[0107] The contribution formula is:

[0108] R i Represents the task execution actions chosen by all other agents.

[0109] Based on the above contribution formula, the service device can determine the contribution information corresponding to the i-th agent performing any of the multiple tasks.

[0110] In some embodiments, the service device determines the corresponding selection probability when the intelligent agent performs a task, which may include: the service device determines the third convolution embedding result corresponding to the intelligent agent and other intelligent agents based on the influence relationship between the intelligent agent and other intelligent agents; the service device determines the interaction information based on the first convolution embedding result and the third convolution embedding result corresponding to the intelligent agent; the service device normalizes the distribution of the state observation information to obtain the allocation information; the service device determines the corresponding selection probability when the intelligent agent performs the task based on the interaction information and the allocation information.

[0111] The influence relationship refers to the impact of other agents on the computing resource utilization and / or bandwidth resources of an agent when the agent performs a task.

[0112] The third convolution embedding result refers to the embedding result obtained by the service device based on the single-layer graph convolutional neural network for the i-th agent and the j-th agent, which can be used i,j =c i,j (o i,j ,a i,j ) means; o i,j represents the state observation information corresponding to the impact of the j-th agent on the i-th agent, a j Represents the task execution action information corresponding to when the j-th agent affects the i-th agent;.

[0113] Interaction information refers to the information obtained after the first convolution embedding result and the third convolution embedding result influence each other.

[0114] Distribution information refers to the information obtained after normalizing the distribution of state observation information.

[0115] After obtaining the interaction information and allocation information, the service device can accurately determine the corresponding selection probability when the agent performs the task.

[0116] Optionally, the service device determines the corresponding selection probability when the agent performs the task based on the interaction information and the allocation information, which may include: the service device obtains the corresponding selection probability when the agent performs the task based on a probability formula.

[0117] The probability formula is:

[0118] Represents interaction information; Indicates allocation information.

[0119] Since the action selected by the i-th agent to perform the task is proportional to the output result, that is, Therefore, based on this proportional relationship, the service device can derive the above probability formula to accurately obtain the corresponding selection probability when the i-th intelligent agent performs the task.

[0120] In addition, for each agent, the policy network update corresponding to each agent is based on the baseline function, where the updated policy gradient corresponding to the policy network is: Represents the initial policy gradient update formula, θ represents the network parameters corresponding to the policy network. The i-th agent updates the policy gradient according to the above Update the policy network corresponding to the agent. The update process is as follows: represents the network parameters corresponding to the policy network after the k+1th update, represents the network parameters corresponding to the policy network after the kth update, and γ represents the preset update parameters; represents the loss function.

[0121] That is to say, in the process of the agent using the baseline function and the action semantic network to update the policy network corresponding to the agent, the action semantics in the task execution action information between the agents are explicitly considered to improve the estimation accuracy of the execution actions of different tasks, which can effectively avoid the negative impact of irrelevant information, thereby providing a more accurate estimate of the execution of each action.

[0122] 203. Generate a task execution instruction according to the target task corresponding to the target contribution information in all contribution information.

[0123] Among them, the task execution instruction is used to instruct the agent to perform the target task.

[0124] After obtaining all contribution information, the service device can determine the target contribution information among all the contribution information. Since each contribution information has a one-to-one correspondence with the task, the service device can obtain the target task corresponding to the target contribution information and generate corresponding task execution instructions.

[0125] In some embodiments, the service device generates a task execution instruction based on the target task corresponding to the target contribution information in all the contribution information, which may include but is not limited to one of the following implementation methods:

[0126] Implementation method 1: The service device determines the maximum contribution information among all contribution information as the target contribution information; the service device generates a task execution instruction according to the target task corresponding to the target contribution information.

[0127] After obtaining all contribution information, the service device can compare the contribution information pairwise and ultimately determine the maximum contribution information, and then generate a task execution instruction based on the target task corresponding to the target contribution information.

[0128] If there are multiple maximum contribution information, the service device may select any number of maximum contribution information from the multiple maximum contribution information, and generate a task execution instruction according to the target tasks corresponding to the respective maximum contribution information.

[0129] Implementation method 2: The service device determines the contribution information greater than a preset contribution threshold among all the contribution information as target contribution information; the service device generates a task execution instruction according to the target task corresponding to the target contribution information.

[0130] The preset contribution threshold may be set before the service device leaves the factory, or may be user-defined, and is not specifically limited here.

[0131] After obtaining all contribution information, the service device can compare the contribution information with the preset contribution threshold one by one, and finally determine the contribution information greater than the preset contribution threshold, and determine the contribution information greater than the preset contribution threshold as the target contribution information, and then generate a task execution instruction.

[0132] For example, the above steps 201 to 203 are simulated, and the simulation process is as follows: Figure 1 The service device determines the cyberspace state information corresponding to agent c1 at time t based on the first, second, and third formulas and the definition of cyberspace state information: The cyberspace state information corresponding to agent c2 is And the cyberspace state information corresponding to agent c3 is Then the network topology information corresponding to these three network space status information is obtained as follows: Then, in R C ∈[-Q C ,Q C ], R B ∈[-Q B ,Q B ], the service device sets Q C =2000Mips, set Q B =2100Mbps, and set the corresponding reward function.

[0133] Then, the service device observes the status information obtained and task execution action information The task execution action value information corresponding to the i-th agent among the three agents is obtained as Q i(O, A) = [19.82, 20.78, …, 26.23]. This task execution action value information can subsequently enable the i-th agent to perform the task execution action corresponding to the target contribution information, thereby realizing action selection in reinforcement learning.

[0134] Afterwards, the service device executes the action value information Q according to the maximum task at time t i (O, A) = 28.32, the maximum contribution information A is obtained i (o i ,a i )=0.392.

[0135] In addition, the service device may use the maximum contribution information A i (o i ,a i )=0.392Substitute into the formula and To perform gradient updates on the policy network.

[0136] Then, the service equipment uses the formula and While considering other agents’ selection of task execution actions, calculate the corresponding selection probability π(a i,j ∣o i )=0.9782.

[0137] 204. Send multiple task execution instructions to corresponding intelligent agents.

[0138] It should be noted that the service device will eventually generate as many task execution instructions as there are intelligent agents; then, the service device will send each task execution instruction to the corresponding intelligent agent; after receiving the corresponding task execution instruction sent by the service device, each intelligent agent can execute the target task based on the task execution instruction to effectively improve the task execution efficiency.

[0139] Optionally, after step 204, the service device determines the evaluation index corresponding to the agent, which may include the corresponding amount of computing resource utilization Rc, bandwidth resource utilization Rb, task completion rate M C and the average task delay D C wait.

[0140] For example, based on the simulation process in step 203, the service device can determine the computing resource utilization R corresponding to the i-th agent. c =60.28%, bandwidth resource utilization R b =27.42%, average task delay D C = 0.58 seconds, task completion rate M C=94.38%.

[0141] In an embodiment of the present invention, the task execution action value information corresponding to each agent is determined based on the state observation information and task execution action information corresponding to each of the multiple agents. Contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information is determined based on the task execution action value information. Task execution instructions are generated based on the target task corresponding to the target contribution information in all the contribution information. The multiple task execution instructions are then sent to the corresponding agents. This method can utilize the contribution information corresponding to each of the multiple tasks performed by each agent to effectively and reasonably schedule and allocate the tasks corresponding to each agent. This method can effectively and reasonably schedule and allocate heterogeneous fusion network resources corresponding to multiple agents, thereby improving the task completion rate of multiple agents, maximizing the utilization of network resources, and optimizing agent performance.

[0142] Furthermore, the heterogeneous converged network resource scheduling method based on multi-agent collaborative learning can effectively improve the performance and service quality of mobile communication networks, supporting a wider range of business needs and application scenarios. While rationally scheduling and allocating network resources, it can significantly reduce task latency, improve overall server resource utilization and the task completion rate of network nodes, thereby enhancing the user experience.

[0143] The heterogeneous fusion network resource scheduling device for multi-agent collaborative learning provided by the present invention is described below. The heterogeneous fusion network resource scheduling device for multi-agent collaborative learning described below and the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning described above can be referenced to each other.

[0144] like Figure 3 , which is a structural diagram of a heterogeneous fusion network resource scheduling device for multi-agent collaborative learning provided by the present invention, and may include: a processing module 301, a transceiver module 302, and an acquisition module 303;

[0145] Processing module 301 is configured to determine task execution action value information corresponding to each agent based on state observation information and task execution action information corresponding to each of the multiple agents; determine contribution information corresponding to each of the multiple tasks executed by the agent corresponding to the task execution action value information based on the task execution action value information; and generate a task execution instruction based on a target task corresponding to a target contribution information in all the contribution information, wherein the task execution instruction is used to instruct the agent to execute the target task;

[0146] The transceiver module 302 is used to send multiple task execution instructions to the corresponding intelligent agents.

[0147] Optionally, the acquisition module 303 acquires, for each agent, the indicator attributes corresponding to the agent when executing the multiple tasks, the relevant state information corresponding to the agent, and the task execution action information;

[0148] The processing module 301 is specifically used to determine the state observation information corresponding to the intelligent agent based on the cyberspace state information corresponding to the indicator attribute and the related state information; and determine the task execution action value information corresponding to the intelligent agent based on the state observation information and the task execution action information.

[0149] Optionally, the processing module 301 is specifically used to determine the first convolution embedding result corresponding to the agent based on the state observation information and the task execution action information; determine the influence parameters generated by other agents among the multiple agents except the agent when the agent performs the multiple tasks; and determine the task execution action value information corresponding to the agent based on the first convolution embedding result and the influence parameters.

[0150] Optionally, the processing module 301 is specifically used to determine the second convolution embedding results corresponding to each of the multiple agents except the agent; for each other agent, use the activation function to perform a linear transformation on the second convolution embedding results corresponding to the other agent to obtain a target embedding result; based on all target embedding results and the weight matrices corresponding to all target embedding results, determine the influence parameters of all other agents on the agent when performing the multiple tasks.

[0151] Optionally, the processing module 301 is specifically used to determine the selection probability of the intelligent agent for each task; and determine the contribution information corresponding to each task among the multiple tasks performed by the intelligent agent corresponding to the task execution action value information based on the selection probability and the task execution action value information.

[0152] Optionally, the processing module 301 is specifically used to determine the third convolution embedding results corresponding to the intelligent agent and other intelligent agents based on the influence relationship between the intelligent agent and the other intelligent agents; determine the interaction information according to the first convolution embedding result and the third convolution embedding result corresponding to the intelligent agent; normalize the distribution of the state observation information to obtain allocation information; and determine the selection probability of the intelligent agent for the task according to the interaction information and the allocation information.

[0153] Optionally, the processing module 301 is specifically used to determine the maximum contribution information among all the contribution information as the target contribution information; or, to determine the contribution information among all the contribution information that is greater than a preset contribution threshold as the target contribution information; and to generate the task execution instruction based on the target task corresponding to the target contribution information.

[0154] like Figure 4 FIG. 4 is a schematic diagram of the structure of a service device provided by the present invention. The service device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call logic instructions in the memory 430 to execute a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning. The method includes: determining task execution action value information corresponding to each agent based on state observation information and task execution action information corresponding to each of the multiple agents; determining contribution information corresponding to each of the multiple tasks executed by the agent corresponding to the task execution action value information based on the task execution action value information; generating a task execution instruction based on the target task corresponding to the target contribution information in all the contribution information, the task execution instruction being used to instruct the agent to execute the target task; and sending the multiple task execution instructions to the corresponding agents.

[0155] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0156] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the above methods. The method includes: determining the task execution action value information corresponding to each agent based on the state observation information and task execution action information corresponding to each of the multiple agents; for each task execution action value information, determining the contribution information corresponding to each task in the multiple tasks performed by the agent corresponding to the task execution action value information based on the task execution action value information; generating a task execution instruction based on the target task corresponding to the target contribution information in all the contribution information, and the task execution instruction is used to instruct the agent to execute the target task; and sending multiple task execution instructions to the corresponding agents.

[0157] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a heterogeneous fusion network resource scheduling method for multi-agent collaborative learning provided by the above-mentioned methods, the method comprising: determining the task execution action value information corresponding to each agent based on the state observation information and task execution action information corresponding to each of the multiple agents; for each task execution action value information, determining the contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information based on the task execution action value information; generating a task execution instruction based on the target task corresponding to the target contribution information in all the contribution information, the task execution instruction being used to instruct the agent to execute the target task; and sending multiple task execution instructions to the corresponding agent.

[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A heterogeneous fusion network resource scheduling method based on multi-agent collaborative learning, characterized in that: include: Determine the task execution action value information corresponding to each intelligent agent based on the state observation information and task execution action information corresponding to each intelligent agent; For each task execution action value information, determining, based on the task execution action value information, contribution information corresponding to each of the multiple tasks executed by the agent corresponding to the task execution action value information; generating a task execution instruction based on a target task corresponding to target contribution information in all contribution information, wherein the task execution instruction is used to instruct the agent to execute the target task; Send multiple task execution instructions to the corresponding agents; The step of determining the task execution action value information corresponding to each intelligent agent based on the state observation information and task execution action information corresponding to each of the multiple intelligent agents includes: For each of the agents, obtaining the corresponding indicator attributes, relevant state information, and task execution action information of the agent when the agent performs the multiple tasks; Determining state observation information corresponding to the agent based on the cyberspace state information corresponding to the indicator attribute and the relevant state information; Determining a first convolutional embedding result corresponding to the agent according to the state observation information and the task execution action information; Determining impact parameters of other intelligent agents among the plurality of intelligent agents, excluding the intelligent agent, on the intelligent agent performing the plurality of tasks; Determine the task execution action value information corresponding to the agent based on the first convolution embedding result and the influencing parameter.

2. The method according to claim 1, characterized in that The determining of the impact parameters of other agents among the plurality of agents, except the agent, on the agent performing the plurality of tasks includes: Determine a second convolution embedding result corresponding to each of the other agents in the plurality of agents except the agent; For each other agent, use the activation function to perform a linear transformation on the second convolution embedding result corresponding to the other agent to obtain the target embedding result; According to all target embedding results and the weight matrices corresponding to all target embedding results, the influence parameters generated by all other agents on the agent when performing the multiple tasks are determined.

3. The method according to claim 1 or 2, characterized in that The determining, based on the task execution action value information, contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information includes: For each task, determining the probability of the agent selecting the task; Contribution information corresponding to each of the multiple tasks performed by the agent corresponding to the task execution action value information is determined based on the selection probability and the task execution action value information.

4. The method according to claim 3, characterized in that Determining the selection probability of the agent for the task includes: Determining third convolution embedding results corresponding to the agent and the other agents based on the influence relationship between the agent and the other agents; Determining interaction information according to the first convolution embedding result and the third convolution embedding result corresponding to the agent; Normalizing the distribution of the state observation information to obtain distribution information; The selection probability of the agent for the task is determined according to the interaction information and the allocation information.

5. The method according to claim 1 or 2, characterized in that The step of generating a task execution instruction according to the target task corresponding to the target contribution information in all the contribution information includes: Determine the maximum contribution information among all the contribution information as the target contribution information; or determine the contribution information greater than a preset contribution threshold among all the contribution information as the target contribution information; The task execution instruction is generated according to the target task corresponding to the target contribution information.

6. A heterogeneous fusion network resource scheduling device for multi-agent collaborative learning, characterized in that: include: a processing module configured to determine, based on state observation information and task execution action information corresponding to each of the plurality of agents, task execution action value information corresponding to each agent; determine, for each task execution action value information, contribution information corresponding to each of the plurality of tasks executed by the agent corresponding to the task execution action value information; and generate a task execution instruction based on a target task corresponding to a target contribution information in all the contribution information, the task execution instruction being used to instruct the agent to execute the target task; The transceiver module is used to send multiple task execution instructions to the corresponding intelligent agents; Processing module, specifically used for: For each of the agents, obtaining the corresponding indicator attributes, relevant state information, and task execution action information of the agent when the agent performs the multiple tasks; Determining state observation information corresponding to the agent based on the cyberspace state information corresponding to the indicator attribute and the relevant state information; Determining a first convolutional embedding result corresponding to the agent according to the state observation information and the task execution action information; Determining impact parameters of other intelligent agents among the plurality of intelligent agents, excluding the intelligent agent, on the intelligent agent performing the plurality of tasks; Determine the task execution action value information corresponding to the agent based on the first convolution embedding result and the influencing parameter.

7. A service device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the heterogeneous fusion network resource scheduling method for multi-agent collaborative learning as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Mobile edge network intelligent resource allocation method capable of dividing tasks

    CN113873022A