Construction method and system based on multi-agent game decision model

By building a multi-agent game decision model, using deep learning network and punishment mechanism to optimize transportation parameters, the problem of insufficient adaptability and global optimization capabilities of traditional PID controllers in material transportation is solved, and efficient and low-energy transportation control is achieved.

CN120257585AInactive Publication Date: 2025-07-04GONGSHEN ZHIHUI (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510283295.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional PID controllers lack adaptability and global optimization capabilities in material transportation, resulting in unstable transportation speed, low efficiency and high energy consumption.

Method used

Using a multi-agent game decision model, a historical transportation data matrix is ​​constructed, deep learning network is trained, and the decision-making agent training and punishment analysis is carried out, and the transportation parameters are optimized to achieve the reference transportation speed and minimum power loss.

Benefits of technology

It improves the efficiency and stability of material transportation and reduces energy consumption during transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257585A_ABST
    Figure CN120257585A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic transportation, in particular to a construction method and system based on a multi-agent game decision-making model, and the method comprises the steps: determining to-be-controlled equipment and a reference decision-making target, constructing a historical transportation data matrix of an original agent, obtaining a decision-making agent based on the historical transportation data matrix, and carrying out the game decision-making of the decision-making agent. In the material transportation step, transportation parameter detection is carried out to obtain a real-time transportation data set, decision punishment analysis is carried out on a plurality of decision intelligent agents according to a decision target formula, a reference decision target and the real-time transportation data set to obtain a decision transportation data set, and the decision transportation data set is input into a target intelligent agent; and obtaining target output power, taking the target output power set as an initial output power set, and repeating the steps until the to-be-controlled equipment receives a transportation stop instruction. The conveying efficiency of material conveying can be improved, and energy consumption in the conveying process can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automated transportation, and particularly to a construction method and system based on a multi-agent game decision model. Background Art

[0002] In modern industrial and logistics systems, the material transportation efficiency is directly related to production costs, resource utilization rates, and the stability of the overall supply chain. Among them, the precise control of transportation speed is a key factor to ensure that materials reach the designated location on time, reduce production delays, and waste of resources.

[0003] Traditional technologies usually use simple PID controllers to adjust the transportation speed. Although this method can achieve the control of transportation speed to a certain extent, they rely on preset parameters and fixed logics, lacking the adaptive ability to complex environments and the global optimization ability, resulting in poor control stability of transportation speed, thus leading to low efficiency of material transportation, and the unstable transportation speed will also cause large power losses. Summary of the Invention

[0004] The present invention provides a construction method and system based on a multi-agent game decision model, and its main purpose is to improve the transportation efficiency of material transportation and reduce the energy consumption during transportation.

[0005] To achieve the above object, a construction method based on a multi-agent game decision model provided by the present invention includes:

[0006] Receiving an agent game instruction, and determining a device to be controlled and a reference decision target based on the agent game instruction. Among them, the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network. The reference decision target includes a reference transportation speed;

[0007] Sequentially extracting original agents from the plurality of original agents, and constructing a historical transportation data matrix of the original agents. The historical transportation data matrix includes: historical acquisition time, historical output power, historical load mass, and historical motor speed;

[0008] Based on the historical transportation data matrix, training the deep learning network in the original agent to obtain a decision-making agent, and summarizing the decision-making agents to obtain a plurality of decision-making agents;

[0009] According to a preset initial output power set and a plurality of decision-making agents, performing material transportation, and detecting transportation parameters during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data set;

[0010] According to a preset decision target formula, a reference decision target, and a real-time transportation data set, decision penalty analysis is performed on multiple decision agents to obtain a decision transportation data set. Among them, the decision target formula aims to make the transportation speed closest to the reference transportation speed and minimize the loss power, and the decision transportation data set contains penalty factors;

[0011] Sequentially extract decision transportation data groups from the decision transportation data set, identify the target agents corresponding to the decision transportation data groups, and input the decision transportation data groups into the target agents to obtain target output powers;

[0012] Summarize the target output powers to obtain a target output power set, use the target output power set as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and multiple decision agents until the device to be controlled receives a preset transportation stop instruction, completing the construction of the multi-agent game decision model.

[0013] Optionally, constructing the historical transportation data matrix of the original agent includes:

[0014] Determine the material transportation device corresponding to the original agent, and extract the work log of the material transportation device. The work log includes the output power, motor speed, and load mass carried by the material transportation device at different times;

[0015] According to a preset historical query period, determine an original transportation data set in the work log, and filter the original transportation data set according to preset target constraint conditions to obtain a historical transportation data set;

[0016] Construct a historical transportation data matrix based on the historical transportation data set.

[0017] Optionally, the step of determining the original transportation data set in the work log according to the preset historical query period includes:

[0018] Record the instruction receiving moment of the step of receiving the agent game instruction, and determine a query time node group based on the instruction receiving moment and the historical query period. The time span of the query time node group is the historical query period;

[0019] Sequentially extract query time nodes in the query time node group, and determine an original transportation data group in the work log based on the query time node. The original transportation data group includes: original acquisition moment, original output power, original load mass, and original motor speed, and the original acquisition moment is the query time node;

[0020] Summarize the original transportation data groups to obtain an original transportation data set.

[0021] Optionally, screening the original transportation data set according to preset target constraint conditions to obtain a historical transportation data set, including:

[0022] Construct target constraint conditions, where the target constraint conditions include: load constraint conditions, rotational speed constraint conditions, and combined constraint conditions;

[0023] Successively extract original transportation data groups from the original transportation data set;

[0024] Based on the original output power, original load mass, and original motor rotational speed in the original transportation data group, determine whether the load constraint conditions, rotational speed constraint conditions, and combined constraint conditions in the target constraint conditions are all satisfied;

[0025] If it is confirmed that the load constraint conditions, rotational speed constraint conditions, and combined constraint conditions in the target constraint conditions are not all satisfied, mark the original transportation data group as an abnormal transportation data group;

[0026] Summarize the abnormal transportation data groups to obtain an abnormal transportation data set, and remove the abnormal transportation data set from the original transportation data set to obtain a historical transportation data set.

[0027] Optionally, the constructing of the target constraint conditions includes:

[0028] Obtain the maximum rated power, motor rotational speed range, and load mass range of the material transportation equipment corresponding to the original intelligent agent;

[0029] Determine the no-load loss power of the material transportation equipment in the no-load state;

[0030] According to the maximum rated power, motor rotational speed range, load mass range, and no-load loss power, respectively construct load constraint conditions, rotational speed constraint conditions, and combined constraint conditions using the following formulas:

[0031] G min ≤G x ≤G max

[0032] V min ≤V x ≤V max

[0033]

[0034] Among them, G min represents the minimum load mass in the load mass range, G x represents a preset load mass variable, G max represents the maximum load mass in the load mass range, V minRepresents the minimum motor speed in the motor speed range, V x Represents the preset motor speed variable, V max Represents the maximum motor speed in the motor speed range, P loss Represents the no-load loss power, π represents pi, r represents the preset transportation radius, γ represents the preset motor efficiency, P x Represents the preset output power variable, P max Represents the maximum rated power;

[0035] Based on the load constraint conditions, speed constraint conditions and joint constraint conditions, the construction of the target constraint conditions is completed.

[0036] Optionally, the historical transportation data matrix is expressed as:

[0037]

[0038] Wherein, R represents the historical transportation data matrix, Represents the first historical transportation data group in the historical transportation data group set, And Respectively represent the historical acquisition time, historical output power, historical load mass and historical motor speed in the first historical transportation data group, Represents the nth historical transportation data group in the historical transportation data group set, n represents the number of historical transportation data groups in the historical transportation data group set, And Respectively represent the historical acquisition time, historical output power, historical load mass and historical motor speed in the nth historical transportation data group.

[0039] Optionally, the decision penalty analysis is performed on multiple decision agents according to the preset decision target formula, reference decision target and real-time transportation data group set to obtain the decision transportation data group set, including:

[0040] Extract the real-time transportation data groups in the real-time transportation data group set in sequence, wherein the real-time transportation data group includes multiple real-time transportation data, and the real-time transportation data includes: real-time acquisition time, real-time output power, real-time load mass and real-time motor speed;

[0041] Identify the reference transportation speed in the reference decision target;

[0042] Calculate the penalty factor according to the reference transportation speed, real-time transportation data group and decision target formula;

[0043] Supplement the penalty factor to each real-time transportation data of the real-time transportation data group to obtain the decision transportation data group, and summarize the decision transportation data groups to obtain the decision transportation data group set.

[0044] Optionally, calculating a penalty factor according to a reference transportation speed, a real-time transportation data set, and a decision target formula includes:

[0045] Extracting a real-time acquisition time set from the real-time transportation data set, sequentially extracting real-time acquisition times in the real-time acquisition time set, obtaining an overall transportation speed of the device to be controlled based on the real-time acquisition times, and summarizing the overall transportation speeds to obtain an overall transportation speed set;

[0046] Calculating a speed dimension coefficient and a power dimension coefficient according to the overall transportation speed set;

[0047] Calculating a penalty factor based on the speed dimension coefficient, the power dimension coefficient, and the decision target formula, where the decision target formula is expressed as:

[0048]

[0049] where K represents the penalty factor, α represents the speed dimension coefficient, N represents the number of real-time transportation data in the real-time transportation data set, v s,i represents the i-th overall transportation speed in the overall transportation speed set, v ref represents the reference transportation speed, β represents the power dimension coefficient, and P s,i represents the real-time output power corresponding to the i-th real-time transportation data in the real-time transportation data set.

[0050] Optionally, calculating the speed dimension coefficient and the power dimension coefficient includes:

[0051] Extracting a real-time output power set from the real-time transportation data set, calculating an average output power of the real-time output power set, and an average transportation speed of the overall transportation speed set;

[0052] Calculating the speed dimension coefficient and the power dimension coefficient according to the average output power and the average transportation speed, where the speed dimension coefficient and the power dimension coefficient are respectively expressed as:

[0053]

[0054] where e v represents a preset speed decision weight, represents the average output power, represents the average transportation speed, and e P represents a preset power decision weight.

[0055] To achieve the above object, the present invention further provides a construction system based on a multi-agent game decision model, including:

[0056] A historical data query module, which is used to receive intelligent agent game instructions, determine the equipment to be controlled and the reference decision target based on the intelligent agent game instructions. Among them, the equipment to be controlled includes multiple original intelligent agents, and the original intelligent agents include material transportation equipment and a deep learning network. The reference decision target includes a reference transportation speed. The original intelligent agents are sequentially extracted from the multiple original intelligent agents to construct a historical transportation data matrix of the original intelligent agents. The historical transportation data matrix includes: historical acquisition time, historical output power, historical load quality, and historical motor speed;

[0057] A real-time parameter monitoring module, which is used to train the deep learning network in the original intelligent agent based on the historical transportation data matrix to obtain a decision-making intelligent agent, summarize the decision-making intelligent agents to obtain multiple decision-making intelligent agents, perform material transportation according to a preset initial output power set and the multiple decision-making intelligent agents, and perform transportation parameter detection during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data set;

[0058] A penalty mechanism calculation module, which is used to perform decision penalty analysis on multiple decision-making intelligent agents according to a preset decision target formula, reference decision target, and real-time transportation data set to obtain a decision transportation data set. The decision target formula aims at the transportation speed being closest to the reference transportation speed and the minimum loss power, and the decision transportation data set includes penalty factors;

[0059] A model iteration and optimization module, which is used to sequentially extract decision transportation data groups from the decision transportation data set, identify the target intelligent agent corresponding to the decision transportation data group, input the decision transportation data group into the target intelligent agent to obtain a target output power, summarize the target output power to obtain a target output power set, use the target output power set as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and the multiple decision-making intelligent agents until the equipment to be controlled receives a preset transportation stop instruction.

[0060] To solve the above problems, the present invention also provides an electronic device, which includes:

[0061] A memory, which stores at least one instruction;

[0062] A processor, which executes the instructions stored in the memory to implement the above-mentioned method for constructing a multi-agent game decision model.

[0063] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned method for constructing a multi-agent game decision model.

[0064] To solve the problems described in the background art, the present invention first determines the device to be controlled and the reference decision target. This step clarifies the target and the control object for decision-making, providing a direction for subsequent decisions. Then, a historical transportation data matrix of the original agent is constructed, and based on this matrix, the deep learning network in the original agent is trained to obtain a decision-making agent. This step provides rich historical data for model training, and at the same time, the training of the decision-making agent improves the adaptability and intelligence level of the entire decision-making process. After obtaining the real-time transportation data set, according to the decision target formula, the reference decision target, and the real-time transportation data set, decision penalty analysis is performed on multiple decision-making agents to obtain a decision transportation data set. This step introduces a penalty mechanism, which quantifies the deviation of the decision-making agent during decision-making through a penalty factor, thereby guiding the decision-making agent to optimize its decision-making strategy to make it closer to the reference decision target, and further realizing the optimization of transportation speed and loss power. Further, the target agent corresponding to the decision transportation data set is identified, and the decision transportation data set is input into the target agent to obtain the target output power. Through this step of feedback adjustment, the decision-making agent dynamically optimizes the output power according to real-time data and the penalty factor, further improving the accuracy and adaptability of the decision-making, enhancing the stability of the transportation speed and reducing the loss power. Finally, the target output power set is used as the initial output power set, and the above steps are repeated until the device to be controlled receives a transportation stop instruction. This step realizes the dynamic iterative optimization of multiple decision-making agents, ensuring that multiple decision-making agents continuously learn and improve during multiple iterations, and finally constructs an efficient and stable multi-agent game decision model, thereby achieving global optimal transportation control. Therefore, the present invention can improve the transportation efficiency of material transportation and reduce the energy consumption during transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 FIG. is a flowchart of a method for constructing a multi-agent game decision model provided by an embodiment of the present invention;

[0066] Figure 2 FIG. is a functional module diagram of a system for constructing a multi-agent game decision model provided by an embodiment of the present invention;

[0067] Figure 3 FIG. is a structural diagram of an electronic device for implementing the method for constructing a multi-agent game decision model provided by an embodiment of the present invention.

[0068] DESCRIPTION OF REFERENCE NUMERALS:

[0069] 1, electronic device; 10, processor; 11, memory; 12, bus.

[0070] The implementation, functional characteristics, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed implementation manners

[0071] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0072] An embodiment of the present application provides a method for constructing a multi-agent game decision-making model. The execution subject of the method for constructing a multi-agent game decision-making model includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the method for constructing a multi-agent game decision-making model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0073] Referring to Figure 1 shown is a schematic flowchart of a method for constructing a multi-agent game decision-making model provided by an embodiment of the present invention. In this embodiment, the method for constructing a multi-agent game decision-making model includes:

[0074] S1. Receive an agent game instruction, and determine a device to be controlled and a reference decision-making target based on the agent game instruction. Among them, the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network. The reference decision-making target includes a reference transportation speed.

[0075] It can be understood that the agent game instruction refers to an instruction for making an intelligent decision on a specific mechanism initiated by a person. The device to be controlled refers to the specific mechanism included in the agent game instruction, which represents the object to be controlled. The reference decision-making target refers to the decision-making purpose that needs to be achieved set by a person. For example, taking a material transportation system as the device to be controlled, the material transportation system refers to a device that conveys an object through a conveyor belt. In order to ensure the stability of the transportation speed of the material transportation system during the transportation of materials, the transportation speed of this material transportation system can be used as the reference decision-making target. The original agent refers to a component unit of the device to be controlled, and each original agent can independently make a decision on the current behavior.

[0076] Exemplarily, in a material transportation system, due to the long transportation distance (the transportation distance exceeds the maximum transportation distance of a single material transportation device), multiple material transportation devices need to work together. Thus, each material transportation device can be regarded as an independent and originally operating intelligent agent, and this originally intelligent agent also includes a decision control algorithm, which is established based on a deep learning network and can control the power output of the material transportation device according to the current operating parameters of the material transportation device, so as to achieve the purpose of the originally intelligent agent making independent behavioral decisions. Here, the purpose of the decision is to control the overall transportation speed of the material transportation system to be the same as the reference transportation speed. Among them, the reference transportation speed refers to the overall speed of the material transportation system for transporting materials set by humans. Since the material transportation system consists of multiple material transportation devices and each material transportation device can make independent decisions, that is, the output power of each material transportation device is different, resulting in different transportation speeds for each material transportation device. In order to ensure that the overall transportation speed of the material transportation system can be stabilized at the reference transportation speed under the cooperation of all material transportation devices, it is necessary to train the originally intelligent agent according to the decision results of the originally intelligent agent, so that all originally intelligent agents can cooperate better with each other.

[0077] S2. Sequentially extract the originally intelligent agents from multiple originally intelligent agents to construct the historical transportation data matrix of the originally intelligent agent, where the historical transportation data matrix includes: historical acquisition time, historical output power, historical load mass, and historical motor speed.

[0078] It is understandable that the historical transportation data matrix refers to a matrix containing the historical decision data of the originally intelligent agent. When the originally intelligent agent includes a material transportation device, the historical decision data includes: historical acquisition time, historical output power, historical load mass, and historical motor speed. Among them, the historical acquisition time refers to the time when the historical decision data is collected, the historical output power refers to the output power of the material transportation device corresponding to the originally intelligent agent at the historical acquisition time, the historical load mass refers to the mass of the materials carried by the material transportation device corresponding to the originally intelligent agent at the historical acquisition time, and the historical motor speed refers to the rotational speed of the motor of the material transportation device corresponding to the originally intelligent agent at the historical acquisition time.

[0079] Specifically, the construction of the historical transportation data matrix of the originally intelligent agent includes:

[0080] Determine the material transportation device corresponding to the originally intelligent agent, and extract the work log of the material transportation device, where the work log includes the output power, motor speed, and the carried load mass of the material transportation device at different times;

[0081] Determine the original transportation data set in the work log according to a preset historical query period, and screen the original transportation data set according to preset target constraint conditions to obtain a historical transportation data set;

[0082] Construct a historical transportation data matrix based on the historical transportation data set.

[0083] It should be explained that the work log refers to a database that records the operating parameters of the material transportation equipment at different times, which includes the output power and the load mass carried by the material transportation equipment at different times. The historical query period refers to the artificially set duration, which represents the time span of this historical transportation data matrix. For example, if 20 hours is used as the historical query period, the original transportation data set refers to the set of original transportation data groups of the material transportation equipment at different times. Among them, the original transportation data group includes: the original acquisition time, the original output power, the original load mass, and the original motor speed.

[0084] Furthermore, since there may be abnormal data in the original transportation data set, such as data exceeding the normal working range, etc., in order to eliminate these abnormal data, target constraint conditions are introduced. The target constraint conditions define the normal working range of the original transportation data in the original transportation data set. The historical transportation data set refers to the original transportation data set after screening. Among them, screening the original transportation data set means eliminating the original transportation data groups that do not meet the target constraint conditions, and recording the original transportation data groups that meet the target constraint conditions as historical transportation data groups.

[0085] Specifically, determining the original transportation data set in the work log according to the preset historical query period includes:

[0086] Record the instruction receiving time of the step of receiving the intelligent agent game instruction, and determine the query time node group based on the instruction receiving time and the historical query period. The time span of the query time node group is the historical query period;

[0087] Extract the query time nodes in the query time node group in sequence, and determine the original transportation data group in the work log based on the query time nodes. The original transportation data group includes: the original acquisition time, the original output power, the original load mass, and the original motor speed. The original acquisition time is the query time node;

[0088] Summarize the original transportation data groups to obtain the original transportation data set.

[0089] It is understandable that the instruction receiving time refers to the time when the agent game instruction is received. The query time node group refers to a combination of multiple time points, wherein an example of determining the query time node group based on the instruction receiving time and the historical query cycle is: if a certain instruction receiving time is 20:00 and the historical query cycle is 10 hours, then the instruction receiving time is taken as the starting point and the query time range is obtained by going back 10 hours: 10:00 to 20:00, then, according to the preset node extraction interval of 1min, the query time nodes are extracted in sequence in the query time range: 10:00, 10:01, 10:02, ..., 19:58, 19:59, 20:00, and these query time nodes constitute the query time node group.

[0090] It can be understood that the original transport data group refers to the operating parameters of the material transportation equipment recorded in the work log at the query time node. The original transport data group includes: the original collection time, the original output power, the original load mass and the original motor speed, which respectively refer to: the query time node, the output power of the material transportation equipment at the query time node, the mass of the material carried and the speed of the motor.

[0091] Specifically, the original transport data set is screened according to the preset target constraint conditions to obtain the historical transport data set, including:

[0092] Constructing target constraints, wherein the target constraints include: load constraints, speed constraints and combined constraints;

[0093] Extracting original transport data groups in sequence from the original transport data group set;

[0094] Based on the original output power, the original load mass and the original motor speed in the original transport data group, determining whether the load constraint condition, the speed constraint condition and the combined constraint condition in the target constraint condition are all established;

[0095] If it is confirmed that the load constraint condition, the speed constraint condition and the combined constraint condition in the target constraint condition are not all satisfied, the original transport data group is recorded as an abnormal transport data group;

[0096] The abnormal transportation data sets are aggregated to obtain an abnormal transportation data set, and the abnormal transportation data set is removed from the original transportation data set to obtain a historical transportation data set.

[0097] It should be explained that the load constraint condition refers to an inequality that restricts the mass of the material carried by the material transportation equipment. The rotational speed constraint condition refers to an inequality that restricts the rotational speed of the motor of the material transportation equipment. The combined constraint condition refers to an inequality that restricts the output power of the material transportation equipment by combining the material mass and the motor rotational speed. The specific forms of the above-mentioned constraint conditions will be given later. The detailed steps for determining whether the load constraint condition, the rotational speed constraint condition, and the combined constraint condition are all satisfied are as follows: Substitute the original load mass into the load constraint condition, substitute the original motor rotational speed into the rotational speed constraint condition, and substitute the original output power, the original load mass, and the original motor rotational speed into the combined constraint condition. After the data substitution is completed, if the corresponding inequalities of the above load constraint condition, rotational speed constraint condition, and combined constraint condition are all satisfied, it is determined that the load constraint condition, the rotational speed constraint condition, and the combined constraint condition are all satisfied; otherwise, the load constraint condition, the rotational speed constraint condition, and the combined constraint condition are not all satisfied.

[0098] Specifically, the construction of the target constraint conditions includes:

[0099] Obtain the maximum rated power, the motor rotational speed range, and the load mass range of the material transportation equipment corresponding to the original intelligent agent;

[0100] Determine the no-load loss power of the material transportation equipment in the no-load state;

[0101] According to the maximum rated power, the motor rotational speed range, the load mass range, and the no-load loss power, use the following formulas to construct the load constraint condition, the rotational speed constraint condition, and the combined constraint condition respectively:

[0102] G min ≤G x ≤G max

[0103] V min ≤V x ≤V max

[0104]

[0105] where G min represents the minimum load mass in the load mass range, G x represents the preset load mass variable, G max represents the maximum load mass in the load mass range, V min represents the minimum motor rotational speed in the motor rotational speed range, V x represents the preset motor rotational speed variable, V max represents the maximum motor rotational speed in the motor rotational speed range, P lossIt represents the no-load loss power, π represents the pi, r represents the preset transportation radius, γ represents the preset motor efficiency, and P x represents the preset output power variable, and P max represents the maximum rated power;

[0106] Based on the load constraint conditions, speed constraint conditions and combined constraint conditions, the construction of the target constraint conditions is completed.

[0107] It can be understood that the maximum rated power refers to the maximum power that the artificially set material transportation equipment can output, the motor speed range refers to the normal working range of the motor of the artificially set material transportation equipment, and this motor speed range can be determined through the product specifications of the material transportation equipment. The load mass range refers to the range of the material mass that the artificially set material transportation equipment can carry. The no-load loss power refers to the power output by the material transportation equipment when it is not carrying materials. The load mass variable, motor speed variable and output power variable respectively refer to the inputtable parameters of the load constraint conditions, speed constraint conditions and combined constraint conditions. The motor efficiency refers to the operation efficiency of the material transportation equipment, and it can be obtained by querying the relevant product specifications. The transportation radius refers to the radius of the drive pulley in the material transportation equipment.

[0108] Furthermore, by substituting the original output power, original load mass and original motor speed in the original transportation data group into the output power variable, load mass variable and motor speed variable respectively, it can be judged whether the corresponding target constraint conditions are established.

[0109] It can be understood that the above combined constraint conditions describe the value range of the output power of the material transmission equipment, which is restricted by the no-load loss power (the lowest limit) and the rated power of the motor (the highest limit) of the motor. In the inequality described by this constraint condition, the minimum value of the output power is composed of the no-load loss of the motor plus the additional power demand generated by carrying the object, and the maximum value is obtained by subtracting the no-load loss from the maximum rated power of the motor. This range ensures the safe operation of the material transportation equipment under different load and speed conditions, and at the same time reflects the direct influence of the material mass and motor speed on the output power.

[0110] Specifically, the historical transportation data matrix is expressed as:

[0111]

[0112] where R represents the historical transportation data matrix, represents the first historical transportation data group in the historical transportation data group set, and respectively represent the historical acquisition time, historical output power, historical load mass and historical motor speed in the first historical transportation data group, represents the nth historical transportation data group in the set of historical transportation data groups, where n represents the number of historical transportation data groups in the set of historical transportation data groups, and respectively represent the historical acquisition moment, historical output power, historical load mass, and historical motor speed in the nth historical transportation data group.

[0113] S3. Based on the historical transportation data matrix, train the deep learning network in the original agent to obtain a decision-making agent, and aggregate the decision-making agents to obtain multiple decision-making agents.

[0114] It should be explained that the training of the deep learning network in the original agent means: using the historical transportation data matrix of the original agent (which includes the historical acquisition moment, historical output power, historical load mass, and historical motor speed) as input features, and through the deep learning network, learning and modeling these data, so as to train an agent that can make optimal decisions according to the current input data, that is, the decision-making agent. Here, the decision-making refers to the magnitude of the power output by the decision-making agent.

[0115] Furthermore, a long short-term memory network (LSTM) can be selected as the deep learning network here. It can effectively capture the time-dependent relationships in the data, such as the variation laws of historical output power, load mass, and motor speed over time. By training the LSTM network, the decision-making agent can learn how to dynamically adjust the output power according to the current and historical data to optimize the transportation speed and minimize the loss power.

[0116] S4. According to the preset initial output power set and multiple decision-making agents, perform material transportation, and based on the preset monitoring duration, detect transportation parameters during the material transportation step to obtain a set of real-time transportation data groups.

[0117] It can be understood that the initial output power set refers to the power output by the decision-making agent set artificially in the initial state, and the monitoring duration refers to the duration set artificially, which represents the time span of the set of real-time transportation data groups. Among them, the set of real-time transportation data groups refers to the set of real-time transportation data groups corresponding to different decision-making agents. In the real-time transportation data group, it includes: real-time acquisition moment, real-time output power, real-time load mass, and real-time motor speed, and one real-time transportation data group corresponds to one decision-making agent.

[0118] It should be explained that the transportation parameter detection means: according to the preset acquisition interval, collect real-time transportation data of the decision-making agent until the cumulative acquisition duration reaches the monitoring duration.

[0119] S5. According to the preset decision target formula, reference decision target, and real-time transportation data set, perform decision penalty analysis on multiple decision agents to obtain a decision transportation data set. The decision target formula aims to make the transportation speed closest to the reference transportation speed and minimize the loss power, and the decision transportation data set includes penalty factors.

[0120] It can be understood that performing decision penalty analysis on multiple decision agents means evaluating the performance of the corresponding decision agent when making a decision on the output power through the real-time transportation data set, thereby obtaining the penalty factor and supplementing the penalty factor to the corresponding real-time transportation data set to obtain the decision transportation data set. The decision target formula aims to make the transportation speed closest to the reference transportation speed and minimize the loss power. By introducing the penalty factor, the degree of deviation of the agent from the target is quantified. The role of the penalty factor is to measure the deviation of the agent's decision. Through the penalty mechanism, the original agent is guided to adjust its decision-making behavior to make its decision closer to the optimization target.

[0121] Furthermore, the penalty factor quantifies the quality of the agent's decision by assigning additional "penalty costs" to the behaviors that deviate from the target. The introduction of the penalty factor enables the original agent to weigh different factors when pursuing the optimization target, avoid excessive deviation from the target, and at the same time guide the original agent to adjust its strategy in subsequent decisions to reduce the penalty and improve the overall decision accuracy. Therefore, the purpose of the penalty factor is to ensure that the behavior of the original agent is always adjusted towards the optimization target.

[0122] Specifically, performing decision penalty analysis on multiple decision agents according to the preset decision target formula, reference decision target, and real-time transportation data set to obtain a decision transportation data set includes:

[0123] Sequentially extract real-time transportation data sets from the real-time transportation data set. The real-time transportation data set includes multiple real-time transportation data, and the real-time transportation data includes: real-time acquisition time, real-time output power, real-time load quality, and real-time motor speed;

[0124] Identify the reference transportation speed in the reference decision target;

[0125] Calculate the penalty factor according to the reference transportation speed, real-time transportation data set, and decision target formula;

[0126] Supplement the penalty factor to each real-time transportation data in the real-time transportation data set to obtain the decision transportation data set, and summarize the decision transportation data sets to obtain the decision transportation data set.

[0127] It can be understood that the decision transportation data set refers to the real-time transportation data set after the penalty factor is supplemented. The detailed calculation process of the penalty factor will be given later.

[0128] Specifically, calculating the penalty factor according to the reference transportation speed, real-time transportation data set, and decision target formula includes:

[0129] Extract the real-time acquisition time group in the real-time transportation data set, sequentially extract the real-time acquisition times in the real-time acquisition time group, and based on the real-time acquisition times, obtain the overall transportation speed of the device to be controlled, and summarize the overall transportation speeds to obtain an overall transportation speed group;

[0130] Calculate the speed dimension coefficient and power dimension coefficient according to the overall transportation speed group;

[0131] Calculate the penalty factor based on the speed dimension coefficient, power dimension coefficient, and decision target formula, where the decision target formula is expressed as:

[0132]

[0133] where K represents the penalty factor, α represents the speed dimension coefficient, N represents the number of real-time transportation data in the real-time transportation data set, v s,i represents the i-th overall transportation speed in the overall transportation speed group, v ref represents the reference transportation speed, β represents the power dimension coefficient, and P s,i represents the real-time output power corresponding to the i-th real-time transportation data in the real-time transportation data set.

[0134] It can be understood that the real-time acquisition time group refers to the combination of all real-time acquisition times that appear in the real-time transportation data set, and the overall transportation speed refers to the overall speed of the device to be controlled for transporting materials. The way to obtain it is: confirm the starting point and ending point of the material transmission. The path distance between the starting point and the ending point is the overall transmission distance. The time taken for a certain material to be transported from the transmission starting point to the transmission ending point is the overall transmission time. Divide the overall transmission distance by the overall transmission time to obtain the overall transportation speed.

[0135] It should be explained that since there is a dimension difference between the speed-related term |v s,i -v ref | and the power-related term P s,i in the subsequent decision target decision formula, in order to ensure that the penalty factor calculated by the decision target formula is not interfered by the dimension difference, it is necessary to introduce the speed dimension coefficient and power dimension coefficient to compensate for this dimension difference.

[0136] Specifically, calculating the speed dimension coefficient and power dimension coefficient includes:

[0137] Extract the real-time output power group in the real-time transportation data set, calculate the average output power of the real-time output power group, and the average transportation speed of the overall transportation speed group;

[0138] Calculate the speed dimension coefficient and the power dimension coefficient according to the average output power and the average transportation speed, where the speed dimension coefficient and the power dimension coefficient are respectively expressed as:

[0139]

[0140] where, e v represents a preset speed decision weight, represents the average output power, represents the average transportation speed, e P represents a preset power decision weight.

[0141] It can be understood that the real-time output power group refers to the combination of all real-time output powers that appear in the real-time transportation data group, and the average output power refers to the average value of all real-time output powers in the real-time output power group. The average transportation speed refers to the average value of all overall transportation speeds in the overall transportation speed group. The speed decision weight refers to the weight constant of the speed set by humans when calculating the penalty factor, and the power decision weight refers to the weight constant of the power set by humans when calculating the penalty factor, where the sum of the speed decision weight and the power decision weight is 1.

[0142] It should be explained that since the decision target formula needs to achieve two decision-making purposes: one is to control the transportation speed to be the same as the reference transportation speed, and the other is to minimize the loss power, two constants need to be set to represent the importance relationship between these two decision-making purposes. And since these two decision-making purposes respectively correspond to the overall transportation speed and the agent output power, the speed decision weight and the power decision weight are introduced. When the speed decision weight is greater than the power decision weight, it means that the importance of purpose one is greater than that of purpose two, that is, the main purpose of this decision is to control the transportation speed to the reference transportation speed, and the secondary purpose is to minimize the loss power. Optionally, the speed decision weight is 0.6 and the power decision weight is 0.4.

[0143] S6. Sequentially extract decision transportation data groups from the decision transportation data group set, identify the target agent corresponding to the decision transportation data group, and input the decision transportation data group into the target agent to obtain the target output power.

[0144] It should be explained that the target agent refers to the decision-making agent corresponding to the decision-making transportation data set. Inputting the decision-making transportation data set into the target agent means that the target agent receives its corresponding decision-making transportation data set as input. Neural networks inside the target agent, such as LSTM, perform forward propagation based on these input data to generate a preliminary output power. Subsequently, the target agent evaluates the preliminary output power through a designed loss function. Among them, the loss function contains a penalty term corresponding to a penalty factor, which is used to quantify the deviation between the output power and the optimization goal. Then, according to the feedback of the loss function, the target agent adjusts the network parameters through backpropagation to optimize its decision-making strategy, thereby generating a new target output power.

[0145] Furthermore, the specific form of the loss function in this solution can be expressed as: LOSS = K, where LOSS represents the loss function.

[0146] S7. Aggregate the target output power to obtain a set of target output powers, use the set of target output powers as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and multiple decision-making agents until the device to be controlled receives a preset transportation stop instruction, completing the construction of the multi-agent game decision-making model.

[0147] It can be understood that the transportation stop instruction refers to an instruction to stop transportation initiated manually.

[0148] It can be understood that by aggregating the target output power to form a new initial output power set and feeding it back to the device to be controlled for the next round of transportation tasks, the dynamic optimization and iterative improvement of multiple decision-making agents are realized. This process reflects the coordination between individuals and the collective in multiple decision-making agents, as well as the game mechanism of gradually achieving global optimality by continuously learning and adjusting strategies. As the iteration progresses, the decision-making strategies of the decision-making agents gradually converge, and finally, the model construction is completed when the transportation stop instruction is received.

[0149] To solve the problems described in the background art, the present invention first determines the device to be controlled and the reference decision-making goal. This step clarifies the goal and the control object for which decision-making is required, providing a direction for subsequent decision-making. Then, a historical transportation data matrix of the original agent is constructed, and based on the historical transportation data matrix, the deep learning network in the original agent is trained to obtain a decision-making agent. This step provides rich historical data for model training, and at the same time, the training of the decision-making agent improves the adaptability and intelligence level of the entire decision-making process. After obtaining the real-time transportation data set, according to the decision-making target formula, the reference decision-making goal, and the real-time transportation data set, decision-making penalty analysis is performed on multiple decision-making agents to obtain a decision-making transportation data set. This step introduces a penalty mechanism, which quantifies the deviation of the decision-making agent during decision-making through a penalty factor, thereby guiding the decision-making agent to optimize its decision-making strategy to make it closer to the reference decision-making goal, and further realizing the optimization of transportation speed and loss power. Further, the target agent corresponding to the decision-making transportation data set is identified, and the decision-making transportation data set is input into the target agent to obtain the target output power. Through this step, through feedback adjustment, the decision-making agent dynamically optimizes the output power according to real-time data and the penalty factor, further improving the accuracy and adaptability of decision-making, enhancing the stability of transportation speed and reducing loss power. Finally, the target output power set is used as the initial output power set, and the above steps are repeated until the device to be controlled receives a transportation stop instruction. This step realizes the dynamic iterative optimization of multiple decision-making agents, ensuring that multiple decision-making agents continuously learn and improve during multiple iterations, and finally constructing an efficient and stable multi-agent game decision-making model, thereby achieving global optimal transportation control. Therefore, the present invention can improve the transportation efficiency of material transportation and reduce the energy consumption during transportation.

[0150] As Figure 2 shown, it is a functional module diagram of a construction system based on a multi-agent game decision-making model provided by an embodiment of the present invention.

[0151] The construction system 100 based on the multi-agent game decision-making model described in the present invention can be installed in an electronic device. According to the functions achieved, the construction system 100 based on the multi-agent game decision-making model can include a historical data query module 101, a real-time parameter monitoring module 102, a penalty mechanism calculation module 103, and a model iterative optimization module 104. The modules described in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0152] The historical data query module 101 is configured to receive an agent game instruction, determine a device to be controlled and a reference decision target based on the agent game instruction. Among them, the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network. The reference decision target includes a reference transportation speed. The original agents are sequentially extracted from the plurality of original agents to construct a historical transportation data matrix of the original agents. The historical transportation data matrix includes: a historical acquisition time, a historical output power, a historical load mass, and a historical motor speed;

[0153] The real-time parameter monitoring module 102 is configured to train the deep learning network in the original agent based on the historical transportation data matrix to obtain a decision-making agent, aggregate the decision-making agents to obtain a plurality of decision-making agents, perform material transportation according to a preset initial output power set and the plurality of decision-making agents, and perform transportation parameter detection during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data set;

[0154] The penalty mechanism calculation module 103 is configured to perform decision penalty analysis on a plurality of decision-making agents according to a preset decision target formula, a reference decision target, and the real-time transportation data set to obtain a decision transportation data set. The decision target formula aims at the transportation speed being closest to the reference transportation speed and the loss power being the least, and the decision transportation data set includes a penalty factor;

[0155] The model iteration and optimization module 104 is configured to sequentially extract decision transportation data groups from the decision transportation data set, identify the target agent corresponding to the decision transportation data group, input the decision transportation data group into the target agent to obtain a target output power, aggregate the target output power to obtain a target output power set, use the target output power set as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and the plurality of decision-making agents until the device to be controlled receives a preset transportation stop instruction.

[0156] Specifically, each module in the construction system 100 of the multi-agent game decision model described in the embodiments of the present invention adopts the same technical means as those in the above-mentioned Figure 1 The construction method of the multi-agent game decision model described in the above has the same technical effects and will not be elaborated here.

[0157] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the construction method of the multi-agent game decision model provided by an embodiment of the present invention.

[0158] The electronic device 1 may include a processor 10, a memory 11, and a bus 12, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for a method of constructing a multi-agent game decision model.

[0159] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 also includes the internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can not only be used to store application software installed on the electronic device 1 and various types of data, such as the code of the program for a method of constructing a multi-agent game decision model, etc., but can also be used to temporarily store data that has been output or will be output.

[0160] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the program for a method of constructing a multi-agent game decision model, etc.), and calling data stored in the memory 11, to execute various functions of the electronic device 1 and process data.

[0161] The bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to implement connection communication between the memory 11 and at least one processor 10, etc.

[0162] Figure 3 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component layout.

[0163] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charge management, discharge management, and power consumption management through the power management system. The power source may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0164] Further, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0165] Optionally, the electronic device 1 may further include a user interface. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0166] The program of the method for constructing a multi-agent game decision model stored in the memory 11 in the electronic device 1 is a combination of multiple instructions, and when running in the processor 10, it can implement:

[0167] Receive an agent game instruction, and determine the device to be controlled and a reference decision target based on the agent game instruction. Among them, the device to be controlled includes multiple original agents, and the original agents include a material transportation device and a deep learning network. The reference decision target includes a reference transportation speed;

[0168] Extract the original agents in sequence among the multiple original agents, and construct a historical transportation data matrix of the original agents. The historical transportation data matrix includes: historical acquisition time, historical output power, historical load mass, and historical motor speed;

[0169] Train the deep learning network in the original agent based on the historical transportation data matrix to obtain a decision agent, and summarize the decision agents to obtain multiple decision agents;

[0170] Perform material transportation according to a preset initial output power set and multiple decision agents, and detect transportation parameters during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data set;

[0171] Perform decision penalty analysis on multiple decision agents according to a preset decision target formula, reference decision target, and real-time transportation data set to obtain a decision transportation data set. The decision target formula aims to make the transportation speed closest to the reference transportation speed and the loss power the least, and the decision transportation data set contains penalty factors;

[0172] Extract the decision transportation data groups in sequence from the decision transportation data set, identify the target agent corresponding to the decision transportation data group, and input the decision transportation data group into the target agent to obtain a target output power;

[0173] Summarize the target output power to obtain a target output power set, use the target output power set as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and multiple decision agents until the device to be controlled receives a preset transportation stop instruction, and complete the construction of the multi-agent game decision model.

[0174] Specifically, the specific implementation method of the above instructions by the processor 10 can refer to Figures 1 to 3 The description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0175] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or system capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0176] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor of an electronic device, can implement:

[0177] Receiving an agent game instruction, determining a device to be controlled and a reference decision target based on the agent game instruction, where the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network, and the reference decision target includes a reference transportation speed;

[0178] Successively extracting original agents from the plurality of original agents to construct a historical transportation data matrix of the original agents, where the historical transportation data matrix includes: a historical acquisition time, a historical output power, a historical load mass, and a historical motor speed;

[0179] Training the deep learning network in the original agent based on the historical transportation data matrix to obtain a decision-making agent, and summarizing the decision-making agents to obtain a plurality of decision-making agents;

[0180] According to a preset initial output power set and a plurality of decision-making agents, perform material transportation, and based on a preset monitoring duration, detect transportation parameters during the material transportation step to obtain a real-time transportation data set;

[0181] According to a preset decision target formula, a reference decision target, and the real-time transportation data set, perform decision penalty analysis on each of the plurality of decision-making agents to obtain a decision transportation data set, where the decision target formula aims to make the transportation speed closest to the reference transportation speed and the loss power least, and the decision transportation data set includes penalty factors;

[0182] Successively extracting decision transportation data groups from the decision transportation data set, identifying the target agent corresponding to the decision transportation data group, and inputting the decision transportation data group into the target agent to obtain a target output power;

[0183] Summarize the target output power to obtain a set of target output powers, use the set of target output powers as the initial output power set, and return the step of performing material transportation according to the preset initial output power set and multiple decision-making intelligent agents until the device to be controlled receives a preset transportation stop instruction, thus completing the construction of the multi-agent game decision-making model.

[0184] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and there may be other partitioning methods in actual implementation.

[0185] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0186] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0187] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A construction method of a multi-agent game decision-making model, characterized in that, The method includes: Receiving an agent game instruction, and determining a device to be controlled and a reference decision target based on the agent game instruction. Among them, the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network. The reference decision target includes a reference transportation speed; Sequentially extracting original agents among the plurality of original agents, and constructing a historical transportation data matrix of the original agents. The historical transportation data matrix includes: historical acquisition time, historical output power, historical load mass, and historical motor speed; Training the deep learning network in the original agent based on the historical transportation data matrix to obtain a decision agent, and summarizing the decision agents to obtain a plurality of decision agents; Performing material transportation according to a preset initial output power set and a plurality of decision agents, and detecting transportation parameters during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data set; Performing decision penalty analysis on a plurality of decision agents according to a preset decision target formula, a reference decision target, and the real-time transportation data set to obtain a decision transportation data set. The decision target formula aims at the transportation speed being closest to the reference transportation speed and the minimum loss power, and the decision transportation data set includes a penalty factor; Sequentially extracting decision transportation data groups in the decision transportation data set, identifying the target agent corresponding to the decision transportation data group, and inputting the decision transportation data group into the target agent to obtain a target output power; Summarizing the target output power to obtain a target output power set, using the target output power set as the initial output power set, and returning to the step of performing material transportation according to the preset initial output power set and a plurality of decision agents until the device to be controlled receives a preset transportation stop instruction, and completing the construction of the multi-agent game decision model.

2. The construction method based on the multi-agent game decision-making model according to claim 1, characterized in that, The constructing of the historical transportation data matrix of the original agent includes: Determining the material transportation device corresponding to the original agent, and extracting the work log of the material transportation device. The work log includes the output power, motor speed, and load mass carried by the material transportation device at different times; Determining an original transportation data set in the work log according to a preset historical query period, and screening the original transportation data set according to a preset target constraint condition to obtain a historical transportation data set; Constructing a historical transportation data matrix based on the historical transportation data set.

3. The construction method of the multi-agent game decision-making model according to claim 2, characterized in that The determining of the original transportation data set in the work log according to a preset historical query period includes: Recording the instruction receiving time of the receiving agent game instruction step, and determining a query time node group based on the instruction receiving time and the historical query period. The time span of the query time node group is the historical query period; Sequentially extracting query time nodes in the query time node group, and determining an original transportation data group in the work log based on the query time node. The original transportation data group includes: original acquisition time, original output power, original load mass, and original motor speed. The original acquisition time is the query time node; Summarize the original transport data groups to obtain a set of original transport data groups.

4. The construction method based on the multi-agent game decision model according to claim 3, characterized in that According to the preset target constraint conditions, screen the set of original transport data groups to obtain a set of historical transport data groups, including: Construct target constraint conditions, where the target constraint conditions include: load constraint conditions, rotational speed constraint conditions, and combined constraint conditions; Successively extract original transport data groups from the set of original transport data groups; Based on the original output power, original load mass, and original motor rotational speed in the original transport data group, determine whether the load constraint conditions, rotational speed constraint conditions, and combined constraint conditions in the target constraint conditions are all satisfied; If it is confirmed that the load constraint conditions, rotational speed constraint conditions, and combined constraint conditions in the target constraint conditions are not all satisfied, mark the original transport data group as an abnormal transport data group; Summarize the abnormal transport data groups to obtain a set of abnormal transport data groups, and remove the set of abnormal transport data groups from the set of original transport data groups to obtain a set of historical transport data groups.

5. The construction method of the multi-agent game decision-making model according to claim 4, characterized in that, The construction of the target constraint conditions includes: Obtain the maximum rated power, motor rotational speed range, and load mass range of the material transport equipment corresponding to the original intelligent agent; Determine the no-load loss power of the material transport equipment in the no-load state; According to the maximum rated power, motor rotational speed range, load mass range, and no-load loss power, use the following formulas to construct load constraint conditions, rotational speed constraint conditions, and combined constraint conditions respectively: G min ≤G x ≤G max V min ≤V x ≤V max Among them, G min represents the minimum load mass in the load mass range, G x represents the preset load mass variable, G max represents the maximum load mass in the load mass range, V min represents the minimum motor speed in the motor speed range, V x represents the preset motor speed variable, V max represents the maximum motor speed in the motor speed range, P loss represents the no-load loss power, π represents pi, r represents the preset transportation radius, γ represents the preset motor efficiency, P x represents the preset output power variable, P max represents the maximum rated power; Complete the construction of the target constraint conditions based on the load constraint conditions, rotational speed constraint conditions, and combined constraint conditions.

6. The construction method of the multi-agent game decision-making model according to claim 5, characterized in that, The historical transport data matrix is expressed as: Among them, R represents the historical transportation data matrix, represents the first historical transportation data group in the historical transportation data group set, and respectively represent the historical acquisition moment, historical output power, historical load quality, and historical motor speed in the first historical transportation data group, represents the nth historical transportation data group in the historical transportation data group set, where n represents the number of historical transportation data groups in the historical transportation data group set, and respectively represent the historical acquisition moment, historical output power, historical load quality, and historical motor speed in the nth historical transportation data group.

7. The construction method of the multi-agent game decision-making model according to claim 6, characterized in that According to the preset decision target formula, reference decision target, and real-time transport data group set, perform decision penalty analysis on multiple decision intelligent agents to obtain a set of decision transport data groups, including: Successively extract real-time transport data groups from the real-time transport data group set, where the real-time transport data group includes multiple real-time transport data, and the real-time transport data includes: real-time acquisition time, real-time output power, real-time load mass, and real-time motor rotational speed; Identify the reference transport speed in the reference decision target; Calculate the penalty factor according to the reference transport speed, real-time transport data group, and decision target formula; Supplement the penalty factor to each real-time transport data in the real-time transport data group to obtain a decision transport data group, and summarize the decision transport data groups to obtain a set of decision transport data groups.

8. The construction method of the multi-agent game decision-making model according to claim 7, characterized in that The calculation of the penalty factor according to the reference transport speed, real-time transport data group, and decision target formula includes: Extract the real-time acquisition time group in the real-time transport data group, successively extract the real-time acquisition time in the real-time acquisition time group, and based on the real-time acquisition time, obtain the overall transport speed of the device to be controlled, and summarize the overall transport speed to obtain an overall transport speed group; Calculate the speed dimension coefficient and power dimension coefficient according to the overall transport speed group; Calculate the penalty factor based on the speed dimension coefficient, power dimension coefficient, and decision target formula, where the decision target formula is expressed as: Among them, K represents the penalty factor, α represents the velocity dimension coefficient, N represents the number of real-time transportation data in the real-time transportation data group, and v s,i represents the i-th overall transportation velocity in the overall transportation velocity group, and v ref represents the reference transportation velocity, β represents the power dimension coefficient, and P s,i represents the real-time output power corresponding to the i-th real-time transportation data in the real-time transportation data group.

9. The construction method of the multi-agent game decision-making model according to claim 8, characterized in that, The calculation of the speed dimension coefficient and power dimension coefficient includes: Extract the real-time output power group from the real-time transportation data group, calculate the average output power of the real-time output power group, and the average transportation speed of the overall transportation speed group; Calculate the speed dimension coefficient and the power dimension coefficient according to the average output power and the average transportation speed, where the speed dimension coefficient and the power dimension coefficient are respectively expressed as: Among them, e v represents a preset speed decision weight, represents the average output power, represents the average transportation speed, e P represents a preset power decision weight.

10. A construction system based on a multi-agent game decision-making model, characterized in that, The system includes: A historical data query module, configured to receive an agent game instruction, determine a device to be controlled and a reference decision target based on the agent game instruction, where the device to be controlled includes a plurality of original agents, and the original agents include a material transportation device and a deep learning network, and the reference decision target includes a reference transportation speed. Extract the original agents in sequence among the plurality of original agents to construct a historical transportation data matrix of the original agents, where the historical transportation data matrix includes: a historical acquisition time, a historical output power, a historical load mass, and a historical motor speed; A real-time parameter monitoring module, configured to train the deep learning network in the original agent based on the historical transportation data matrix to obtain a decision-making agent, summarize the decision-making agents to obtain a plurality of decision-making agents, perform material transportation according to a preset initial output power set and the plurality of decision-making agents, and perform transportation parameter detection during the material transportation step based on a preset monitoring duration to obtain a real-time transportation data group set; A penalty mechanism calculation module, configured to perform decision penalty analysis on the plurality of decision-making agents according to a preset decision target formula, a reference decision target, and the real-time transportation data group set to obtain a decision transportation data group set, where the decision target formula aims at the transportation speed being closest to the reference transportation speed and the loss power being the least, and the decision transportation data group includes a penalty factor; A model iteration optimization module, configured to sequentially extract decision transportation data groups from the decision transportation data group set, identify the target agent corresponding to the decision transportation data group, input the decision transportation data group into the target agent to obtain a target output power, summarize the target output power to obtain a target output power set, use the target output power set as the initial output power set, and return to the step of performing material transportation according to the preset initial output power set and the plurality of decision-making agents until the device to be controlled receives a preset transportation stop instruction.