Intelligent agent control method and device and storage medium

Through the three-level agent collaborative control and dynamic path planning, the problem of coordinated control of the agent system in complex environments is solved, efficient communication and path planning is realized, and the stability and efficiency of the system are improved.

CN120335451AInactive Publication Date: 2025-07-18BEIJING JUNDE INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510515279.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing agent systems face a large number of agents and complex environments, it is difficult to achieve efficient coordinated control, resulting in communication delays, conflicts and control instructions lag, affecting the overall performance of the system.

Method used

The three-level agent collaborative control method is adopted, including pilot agents, follow agents and auxiliary agents. By building target calculation models, optimizing control data, and introducing timestamps and virtual random number generators for path planning, achieving efficient communication and collaborative control.

Benefits of technology

It improves the overall control accuracy and stability of the multi-agent system, reduces communication delays and conflicts, enhances the adaptability and reliability of the system, and improves the flexibility of path planning and the ability to disperse traffic flows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335451A_ABST
    Figure CN120335451A_ABST
Patent Text Reader

Abstract

The invention discloses an agent control method and device and a storage medium, and the method comprises the steps: constructing a target calculation model based on an agent target moving track, and obtaining first control data based on the constructed target calculation model; optimizing the obtained first control data by using an algorithm to obtain second control data; the method comprises the following steps: collecting trial flight data of intelligent agents, and dividing the intelligent agents into three levels, namely a navigation intelligent agent, a following intelligent agent and an auxiliary intelligent agent, based on the collected trial flight data; and inputting the optimized second control data into the divided pilot agent, following agent and auxiliary agent, wherein the pilot agent, following agent and auxiliary agent perform cooperative control on the agent queue. Through cooperative control of the three-level intelligent agents, the overall control accuracy and stability of the multi-agent system are improved, and the following intelligent agents can more accurately maintain the formation and the track due to the guidance of the navigation intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to an intelligent agent control method, device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent agent systems have been widely used in various fields, such as autonomous driving, unmanned aerial vehicle formation, and robot collaborative operations. An intelligent agent system usually consists of multiple intelligent agents with autonomous decision-making and communication capabilities, and they can work together to complete complex tasks. However, the control problem of intelligent agent systems has always been a research hotspot and difficulty, especially in aspects such as multi-agent collaborative control, path planning, and information synchronization.

[0003] For example, a Chinese patent with the publication number CN114326759B discloses a multi-agent formation control method, device and multi-agent system. The method includes: for each intelligent agent in the target intelligent agent queue, based on the target running trajectory of the target intelligent agent queue and the state data of each intelligent agent at the previous moment, obtaining the first control data of each intelligent agent at the current moment; based on the target constraint condition corresponding to each intelligent agent and the ADMM algorithm, optimizing the first control data to obtain the second control data of each intelligent agent at the current moment; and controlling each intelligent agent based on the second control data. The multi-agent formation control method, device and multi-agent system provided by the present invention can decouple the target constraint condition corresponding to each intelligent agent based on the ADMM algorithm, can decompose complex problems into multiple sub-problems, can achieve more accurate multi-agent formation control, can reduce the calculation difficulty, improve the calculation speed and efficiency, and have lower requirements for computing devices.

[0004] The control methods in existing patent technical documents often have difficulty dealing with the situation of a large number of intelligent agents and complex and changeable environments. This is because each intelligent agent has its own state information, such as position, speed, attitude, etc., and this information needs to be updated and shared in real time to ensure the coordination and consistency of the entire system. In addition, communication delays and conflicts between intelligent agents may also cause lags and errors in control instructions, thus affecting the overall performance of the system. Summary of the Invention

[0005] The present application provides an intelligent agent control method, device and storage medium, which improve the overall control accuracy and stability of a multi-agent system through three-level intelligent agent collaborative control. The leading of the leading intelligent agent enables the following intelligent agents to more accurately maintain the formation and trajectory, and the environmental information and data support provided by the auxiliary intelligent agents enhance the adaptability of the system.

[0006] The present application provides an intelligent agent control method, including: S101. Based on the target running trajectory of the agent, obtain the data status information, construct a target calculation model according to the obtained data status information, and obtain the first control data based on the constructed target calculation model; S102. Optimize the obtained first control data using an algorithm to obtain the second control data; S104. Collect the trial voyage data of the agent, and classify the agent into three levels based on the collected trial voyage data. The first level is the leading agent, the second level is the following agent, and the third level is the auxiliary agent; S106. Input the optimized second control data into the divided leading agent, following agent, and auxiliary agent to perform cooperative control on the agent queue.

[0007] Preferably, the method of optimizing using an algorithm is as follows: decompose each agent optimization problem into multiple sub-problems. For each sub-problem, assume that the goal is to minimize an objective function containing constraint conditions , where is the control variable of the i-th agent and satisfies the global constraint condition Ax = b. At the same time, calculate the corresponding Lagrangian function, and the formula is: ( ) = + (A - ) + , where ρ > 0 is the penalty parameter, which is used to control the coupling degree between the original variable and the dual variable; is the original variable; is the auxiliary variable; is the dual variable. The formula for optimizing the original variable is: = , where k represents the number of iterations, A is the global linear transformation matrix. The formula for optimizing the auxiliary variable is: = . The formula for updating the dual variable is: = + .

[0008] Preferably, step S103 includes: collecting the trial voyage data of the agent, comparing the collected trial voyage data with the target data of the agent, calculating the position error, speed error, and direction error of the agent, weighting the position error, speed error, and direction error, calculating the comprehensive error value of the agent, presetting a threshold range, and classifying the agent with a comprehensive error value less than the preset threshold range as a leading agent; classifying the agent with a comprehensive error value equal to the preset threshold range as a following agent; and classifying the agent with a comprehensive error value greater than the preset threshold range as an auxiliary agent.

[0009] Preferably, the method for cooperative control of the agent queue further includes: S201, adding a time dimension to the agent communication system; S202, the leading agent simultaneously sends information to the following agent and the auxiliary agent, the following agent and the auxiliary agent verify the received information, and the leading agent adjusts the path planning information according to the current environmental information and the target position; S203, based on the verification of the received information by the following agent and the auxiliary agent in step S202, sending the verified error information to the leading agent so that the leading agent can make adjustments according to the collected error information.

[0010] Preferably, in step S202, the information received by the following agent and the auxiliary agent includes a timestamp, a sequence number, and path information.

[0011] Preferably, adjusting the path planning information includes: S301, the auxiliary agent collects real-time environmental information, the following agent collects its own state information, and the leading agent integrates the environmental information of the auxiliary agent and the own state information of the following agent; S302, based on the integrated environmental information and the state information of the following agent in step S301, identifying path decision points; S303, based on the identified path decision points, the leading agent generates six paths according to the environmental information provided by the auxiliary agent; S304, based on the six generated paths, setting up a virtual random number generator that can generate a random integer between 1 and 6, and when the leading agent moves to the decision point, dynamically selecting a path according to the leading agent's use of the virtual random number generator.

[0012] Preferably, according to the dynamic selection of the path by the leading agent, a path dynamically selected using the virtual random number generator is transmitted to the leading agent so that the leading agent can perform cooperative control on the agent queue.

[0013] Preferably, the path decision point is that the agent encounters several paths during driving and makes a path selection based on the several paths.

[0014] The present application also provides an agent control device, including an acquisition module, a control module, a preprocessing module and a training module. The acquisition module is connected to the preprocessing module, the preprocessing module is connected to the control module, and the control module is connected to the training module.

[0015] The present application also provides an agent storage medium, on which a computer program is stored.

[0016] One or more technical solutions provided in the present application have at least the following technical effects or advantages: Through the collaborative control of three-level agents, the overall control accuracy and stability of the multi-agent system are improved. The leading of the leading agent enables the following agents to maintain the formation and trajectory more precisely. The environmental information and data support provided by the auxiliary agents enhance the adaptability of the system. The rationality of agent role allocation improves the overall performance and efficiency of the system. The control accuracy is improved, and the position error and direction error are reduced. Efficient communication and collaborative control between agents are realized, improving the real-time performance, response speed and stability of the system, reducing communication delay and conflict. Through the introduction of timestamps and information synchronization verification, the accuracy and consistency of communication messages are significantly improved. By sending error information, accurate adjustment basis is provided for the leading agent, improving the overall collaborative efficiency and stability of the system, enabling the agents to adapt to environmental changes and execute tasks faster, and at the same time enhancing the reliability and availability of the system. By identifying the path decision point, the leading agent can generate multiple candidate paths according to real-time environmental information, which greatly improves the flexibility of path planning. The introduction of a virtual random number generator makes the path selection process random, thus increasing the diversity of path planning, avoiding the agents always choosing the same path, helping to disperse traffic flow and relieve road congestion. By dynamically adjusting the path, the leading agent can guide the following agents to drive along the optimal path in real time, realizing the collaborative control of the agent queue, and helping to improve the driving efficiency and safety of the entire agent queue. Brief Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a method for controlling an agent according to the present invention; Figure 2 It is a schematic flowchart of verifying the information synchronization between agents by adding a time dimension to the agents in an embodiment of the present invention; Figure 3 It is a schematic flowchart of dynamically planning a path for the leading agent in an embodiment of the present invention. Detailed Embodiments

[0018] To facilitate the understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0019] It should be noted that the terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only embodiments.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs; the terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0021] Embodiment 1: Figure 1 It is a schematic flowchart of an agent control method according to an embodiment of the present invention, including: S101, based on the target running trajectory of the agent, obtain data state information, construct a target calculation model according to the obtained data state information, and obtain first control data based on the constructed target calculation model; Furthermore, the target running trajectory is the path that the agent queue expects to follow during the execution of the task. For each agent, state information such as its position, speed, and acceleration at different times is obtained through a path planning algorithm. Based on the above information, a target calculation model is constructed using a machine learning algorithm. The state data (such as position, speed, attitude, etc.) of each agent at the previous moment is measured using a sensor, and the obtained state data of the agent at the previous moment is input into the constructed target calculation model. The target calculation model calculates the first control data of each agent at the current moment according to the input state data, and the target calculation model outputs the calculated first control data.

[0022] S102, optimize the obtained first control data using an algorithm to obtain second control data; Specifically, complex agent optimization problems such as the position and speed of each agent are decomposed into multiple sub-problems. Each sub-problem corresponds to the local optimization of one agent or a group of agents. The ADMM algorithm is applied to each sub-problem for iterative solution. In each iteration, the original variables and dual variables are alternately optimized. For each sub-problem, the goal is to minimize an objective function that includes constraint conditions , where is the control variable of the $i$-th agent (such as position, velocity, etc.), and these control variables need to satisfy the global constraint $Ax = b$ (where $A$ and $b$ are the global linear transformation matrix and the constraint vector, but they may only involve the control variables of some agents in each sub-problem). To apply the ADMM algorithm, we introduce the auxiliary variable and the dual variable . According to the current iteration results, update the control variables of the agents, such as position, velocity, etc., and substitute the updated iteration results into the initialized objective function , and at the same time calculate the corresponding Lagrangian function, the formula is: ( ) = + (A - ) + , where $\rho>0$ is the penalty parameter, which is used to control the coupling degree between the primal variable and the dual variable; is the primal variable; is the auxiliary variable; is the dual variable. The formula for optimizing the primal variable is: = , where $k$ represents the number of iterations, and $A$ is the global linear transformation matrix. The formula for optimizing the auxiliary variable is: = . The formula for updating the dual variable is: = + . It is used to handle the constraint conditions. The Lagrange multiplier reflects the relaxation or tightness of the constraint conditions. By alternately optimizing the primal variable and the dual variable, the optimal solution of the problem is gradually approximated. After each iteration, it is judged whether the current iteration result meets the predetermined convergence criterion. If it is satisfied, the iterative calculation is stopped; otherwise, the next iteration is continued. After reaching the predetermined number of iterations or convergence criterion, the current iteration result is extracted as the second control data of each agent at the current moment. The iteration result includes control instructions such as the optimal position, velocity, and acceleration of the agent.

[0023] S103. Collect the trial voyage data of the agents, and classify the agents into three levels based on the collected trial voyage data. The first level is the leading agent, the second level is the following agent, and the third level is the auxiliary agent; Specifically, collect the trial voyage data of the agents, compare the collected trial voyage data with the target data of the agents, calculate the position error, speed error, and direction error. The position error is the deviation between the actual position and the target position of the agent, the speed error is the deviation between the actual speed and the target speed of the agent, and the direction error is the deviation between the actual direction and the target direction of the agent. Using the method of weighted summation, weight the position error, speed error, and direction error according to their importance, and then calculate the comprehensive error value of each agent. Sort the agents in ascending order according to the calculated comprehensive error value. Preset a threshold range. For the agents with a comprehensive error value less than the preset threshold range, classify them as leading agents; for the agents with a comprehensive error value equal to the preset threshold range, classify them as following agents; for the agents with a comprehensive error value greater than the preset threshold range, classify them as auxiliary agents.

[0024] S104, input the optimized second control data into the classified leading agents, following agents, and auxiliary agents to perform cooperative control on the agent queue; Furthermore, configure the basic parameters and initial calculations of the leading agents, following agents, and auxiliary agents according to the objectives and task requirements, establish a communication link between the leading agents, following agents, and auxiliary agents to enable information transfer between the agents. Input the optimized second control data to the leading agents. The leading agents adjust their driving directions and speeds according to the second control data. The leading agents send the adjusted driving direction and speed reference signals to the following agents through the communication link. After receiving the instructions, the following agents adjust their own driving directions and speeds to maintain the relative position and formation with the leading agents. The auxiliary agents continuously collect and process environmental information, such as terrain, climate, obstacles, etc., and send the real-time environmental information to the leading agents and following agents.

[0025] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: Through the cooperative control of three-level agents, the overall control accuracy and stability of the multi-agent system are improved. The leading of the leading agents enables the following agents to maintain the formation and trajectory more precisely. The environmental information and data support provided by the auxiliary agents enhance the adaptability of the system. The rationality of the agent role allocation improves the overall performance and efficiency of the system, the control accuracy is improved, and the position error and direction error are reduced.

[0026] Embodiment 2: Based on the leading agents, following agents, and auxiliary agents in Embodiment 1, in this embodiment, by introducing the time dimension, verify the synchronization of information transmission, and can achieve efficient communication and cooperative control between agents.

[0027] Such as Figure 2As shown in the figure, the steps of adding a time dimension to the agent and verifying the information synchronization between agents are as follows: S201. Add a time dimension to the agent communication system; Furthermore, set the format of the timestamp. Use a 64-bit integer to represent the timestamp. The 64-bit integer enables the agent communication system to have higher precision. Divide the 64-bit timestamp into a high part and a low part. The high part is used to represent the number of seconds, represented by a 48-bit integer. The high part can accommodate the total number of seconds from a certain reference time (such as the start time of the Unix epoch) to the present. The low part is used to represent the number of nanoseconds or finer time units, represented by the remaining 16 bits. The low part represents the number of nanoseconds within each second. Determine the Unix epoch (00:00:00 UTC on January 1, 1970) as the reference time.

[0028] In the message structure definition of the agent communication system, add a new field to store the timestamp. The type of this field matches the timestamp format, so this field is 64 bits. When the sender is ready to send a message, obtain the current time of its local clock. The obtained time usually includes two parts: the number of seconds and the number of nanoseconds. According to the designed timestamp format, combine the number of seconds and the number of nanoseconds into a 64-bit unsigned integer, assign the converted timestamp to the timestamp field in this message structure, and add the timestamp field to the head of the message structure to ensure that the timestamp field has been correctly filled before the message is sent.

[0029] S202. The leader agent sends information to the follower agent and the auxiliary agent simultaneously. The follower agent and the auxiliary agent verify the received information. The leader agent adjusts the path planning information according to the current environmental information and the target location; Specifically, the leader agent determines a fixed communication cycle, which is set according to the real-time requirements of the system and the movement speed of the agent. Before the start of each communication cycle, the leader agent calculates the path planning information for the next stage according to the current environmental information and the target location. The path planning information includes key parameters such as the expected movement trajectory, speed, and acceleration of the leader agent. The leader agent synchronizes with its local clock before sending the information. At the start of the communication cycle, the leader agent encapsulates the prepared path planning information into a message and sends it to all follower agents and auxiliary agents simultaneously through the communication network. During the sending process, ensure the reliability and integrity of the message.

[0030] The follower agent and the auxiliary agent verify the timestamps. After receiving the path information sent by the leader agent, the follower agent and the auxiliary agent store it in the local memory. When the follower agent and the auxiliary agent receive the path information, they also receive and store the timestamp of this information at the same time. The follower agent and the auxiliary agent respectively compare the received timestamp with their own system time to obtain the timestamp difference value. Set a time threshold according to the requirements of the agent system, and compare the timestamp difference value with the preset time threshold. If the timestamp difference value is less than or equal to the preset time threshold, it is considered that the timestamps are consistent and the time is synchronized; if the timestamp difference value is greater than the preset time threshold, it is considered that the timestamps are inconsistent and the time is not synchronized. When the time is not synchronized, the follower agent and the auxiliary agent issue a warning and record the abnormal information.

[0031] The follower agent and the auxiliary agent verify the sequence numbers. The purpose of verifying the sequence numbers is to ensure that the path information can be transmitted in the correct order during the transmission process and there is no loss or duplication. Before sending each path information, the leader agent assigns a unique sequence number to it. This sequence number is usually an increasing integer and increases each time a new information is sent. When the follower agent and the auxiliary agent receive the path information, they receive and store the sequence number of this information at the same time. The follower agent and the auxiliary agent respectively check the received path information in the order of the sequence numbers. If the sequence numbers are consecutive and increasing, it means that the information is received in the correct order. If a sequence number is missing (that is, the information with a certain sequence number is not received) when checking the order, the follower agent and the auxiliary agent will issue a warning, and the warning information includes the missing sequence number so that the leader agent can identify and re-send the missing information.

[0032] After the follower agent and the auxiliary agent receive the path information, they compare the path information received by the follower agent with the path information received by the auxiliary agent. After comparison, it is found that the received path information is inconsistent. The follower agent and the auxiliary agent will issue a warning and record the abnormal information.

[0033] S203. Based on the verification of the received information by the follower agent and the auxiliary agent in step S202, send the verified error information to the leader agent so that the leader agent can adjust according to the collected error information; Further, collect the error information obtained after verification in step S202. The error information includes time difference, data inconsistency, and missing or duplicate serial numbers. Organize the collected error information into a message format and send the collected and organized error information to the pilot agent. Based on the received error information, the pilot agent makes adjustments. The pilot agent adjusts the communication cycle to enable more frequent transmission of information. By shortening the communication cycle, the information update speed can be increased, thereby reducing latency. The pilot agent optimizes the path planning algorithm to reduce the calculation time and improve the efficiency of path selection. At the same time, for the error information of missing or duplicate serial numbers, the pilot agent reduces the problem of missing or duplicate serial numbers by resending information to the follower agents and auxiliary agents.

[0034] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: realizing efficient communication and collaborative control between agents, improving the real-time performance, response speed and stability of the system, reducing communication latency and conflicts, significantly improving the accuracy and consistency of communication messages through the introduction of timestamps and information synchronization verification, providing an accurate adjustment basis for the pilot agent through the sending of error information, improving the overall collaborative efficiency and stability of the system, enabling the agents to adapt to environmental changes and execute tasks faster, and at the same time enhancing the reliability and usability of the system.

[0035] Embodiment 3: Based on the path planning information of the pilot agent in Embodiment 1 and Embodiment 2, this embodiment performs dynamic path planning, introducing a virtual random number generator to dynamically adjust the path, improving the efficiency and diversity of the pilot agent's path planning.

[0036] As Figure 3 shown, adjusting the path planning information includes: S301, the auxiliary agent collects real-time environmental information, the follower agent collects its own state information, and the pilot agent integrates the environmental information of the auxiliary agent and the own state information of the follower agent. Further, the auxiliary sensor uses a camera to collect the surrounding environmental information. The auxiliary agent uses an image processing algorithm to extract features in the image. The features include edges, corner points, textures, and colors. The auxiliary agent uses a machine learning algorithm to identify obstacles in the image, determine whether a certain area in the image contains an obstacle, and the type of the obstacle (such as a vehicle, a pedestrian, a building, etc.). The auxiliary agent analyzes consecutive frames of images, matches the obstacles in the current frame with the obstacles in the previous frame to determine their movement trajectories and speeds.

[0037] The auxiliary agent monitors the road conditions through camera and radar data, such as whether the road surface is slippery, whether there are potholes or construction areas. The traffic flow information is obtained by counting the number of vehicles and their speeds. The agent can identify the traffic flow in different lanes, providing an important basis for path planning.

[0038] S302. Based on the environmental information integrated in step S301 and the status information of the following agent, identify the path decision points; Specifically, the path decision points are the positions where the agent needs to make clear path choices during driving. These positions include intersections, bifurcated roads, in front of construction areas, speed limit areas, crosswalks, etc., all of which are key nodes encountered during the agent's driving. The leading agent constructs a real-time environmental model based on the environmental perception data of the auxiliary agent. The environmental model includes the detailed structure of the roads around the agent, traffic rules, traffic signals, and the real-time positions and dynamic information of obstacles (such as vehicles, pedestrians, etc.). Based on the environmental model, identify the structure of the roads in the environmental model. At intersections, since there are multiple driving direction options, they are identified as path decision points; on bifurcated roads, the leading agent needs to decide which branch road to take, so they are also path decision points; in front of construction areas, the leading agent needs to detour or wait, which are also identified as path decision points. The leading agent forms a path decision point database.

[0039] S303. Based on the identified path decision points, the leading agent generates six paths according to the environmental information provided by the auxiliary agent; Furthermore, based on the environmental information, the leading agent determines all possible directions starting from the current position, generates a basic path for each possible driving direction, scores each basic path based on whether it conforms to traffic rules, road physical restrictions (such as one-way driving, prohibited turning, etc.), and the vehicle's operating capabilities, sorts the scores, and selects the six paths with the highest scores.

[0040] S304. Based on the six generated paths, set up a virtual random number generator that can generate random integers between 1 and 6. When the leading agent moves to the decision point, the leading agent dynamically selects a path using the virtual random number generator; Specifically, the virtual random number generator is a virtual random number generator that can generate random integers between 1 and 6. Each integer corresponds to a candidate path. The random number generator is tested and verified multiple times to ensure that the generated random number sequence is uniformly distributed, that is, the number of times each path is selected in a large number of trials is close to the theoretical value. When the pilot agent is at a decision point, it will activate the virtual random number generator. When the mechanism is activated, the random number generator will start to operate and generate a random integer between 1 and 6. The pilot agent has a predefined list of candidate paths, and the paths in the list are numbered in order from 1 to 6. The generated random integer directly corresponds to a path number in the list. The pilot agent finds the corresponding path in the candidate path list according to the generated random number.

[0041] In some embodiments, the pilot agent is driving an autonomous vehicle on a road with multiple feasible directions, such as an upcoming intersection. There are options to go straight, turn left, turn right, and several other alternative routes (such as driving slightly to the left, driving slightly to the right, entering the access road, etc.) ahead. The pilot agent has generated six candidate paths based on environmental information and vehicle status and numbered them in order from 1 to 6, which are: 1. Go straight through the intersection, 2. Turn left into the left road, 3. Turn right into the right road, 4. Drive slightly to the left to avoid the construction area ahead, 5. Drive slightly to the right to overtake using the wider lane, 6. Enter the access road to bypass the congested section ahead. To select one of these six paths as the actual driving path, the pilot agent sets up a virtual random number generator. This mechanism is a virtual random number generator that can generate random integers between 1 and 6. Each integer corresponds to a candidate path. When the pilot agent moves near the decision point (i.e., the intersection), the virtual random number generator is activated, and the random number generator starts to operate. After a series of calculations, it generates a random integer between 1 and 6. Suppose this integer is 4. The pilot agent has a predefined list of candidate paths, and the paths in the list are numbered in order. The generated random integer 4 directly corresponds to the fourth path in the list, that is, "Drive slightly to the left to avoid the construction area ahead". Then, the pilot agent finds the corresponding path in the candidate path list according to the generated random number 4 and decides to take this path as the actual driving path. It adjusts the driving direction of the vehicle and starts to drive according to the selected path, successfully avoiding the construction area ahead and continuing to move forward towards the destination.

[0042] S305. Based on steps S301 - S304, the path of the pilot agent is dynamically adjusted, and the dynamically adjusted path is transmitted to the pilot agent in step S104. The pilot agent performs collaborative control on the intelligent agent queue.

[0043] The technical solutions in the embodiments of the present application have at least the following technical effects or advantages: By identifying path decision points, the leading agent can generate multiple candidate paths according to real-time environmental information, which greatly improves the flexibility of path planning. The introduction of a virtual random number generator makes the path selection process random, thereby increasing the diversity of path planning, avoiding the agent always choosing the same path, helping to disperse traffic flow and relieve road congestion. By dynamically adjusting the path, the leading agent can guide the following agent to drive along the optimal path in real time, realizing the collaborative control of the agent queue, and helping to improve the driving efficiency and safety of the entire agent queue.

[0044] Embodiment 4: This embodiment provides an agent control device for executing the above-mentioned agent control method, including an acquisition module, a control module, a preprocessing module, and a training module. The acquisition module is used to collect and process various information required by the agent, including environmental information, agent status information, etc. The control module is responsible for generating control instructions according to the acquired information and performing collaborative control on the agent queue. The preprocessing module is responsible for preliminarily processing and analyzing the acquired information for subsequent module use. The training module is responsible for training and optimizing machine learning algorithms to improve the control accuracy and adaptability of the agent.

[0045] Embodiment 5: This embodiment provides an agent storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0046] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An agent control method, characterized in that, Including: S101, based on the target running trajectory of the agent, obtain data status information, construct a target calculation model according to the obtained data status information, and obtain first control data based on the constructed target calculation model; S102, optimize the obtained first control data using an algorithm to obtain second control data; S103, collect the trial voyage data of the agent, and divide the agent into three levels based on the collected trial voyage data. The first level is the leading agent, the second level is the following agent, and the third level is the auxiliary agent; S104, input the optimized second control data into the divided leading agent, following agent, and auxiliary agent to perform cooperative control on the agent queue.

2. The intelligent agent control method according to claim 1, wherein The method of optimization using an algorithm is as follows: Decompose each agent optimization problem into multiple sub-problems. For each sub-problem, assume the goal is to minimize an objective function that includes constraint conditions , where is the control variable of the i-th agent and satisfies the global constraint condition Ax = b. At the same time, calculate the corresponding Lagrangian function, and the formula is: ( ) = + (A - ) + , where ρ > 0 is the penalty parameter, which is used to control the coupling degree between the primal variable and the dual variable; is the primal variable; is the auxiliary variable; is the dual variable. The formula for optimizing the primal variable is: = , where k represents the number of iterations, A is the global linear transformation matrix. The formula for optimizing the auxiliary variable is: = . The formula for updating the dual variable is: = + .

3. The intelligent agent control method according to claim 1, wherein Step S103 includes: collecting the trial voyage data of the agent, comparing the collected trial voyage data with the target data of the agent, calculating the position error, speed error, and direction error of the agent, weighting the position error, speed error, and direction error, and calculating the comprehensive error value of the agent. Preset a threshold range. For an agent with a comprehensive error value less than the preset threshold range, classify it as a leading agent; for an agent with a comprehensive error value equal to the preset threshold range, classify it as a following agent; for an agent with a comprehensive error value greater than the preset threshold range, classify it as an auxiliary agent.

4. The intelligent agent control method according to claim 1, characterized in that, The method for performing cooperative control on the agent queue further includes: S201, add a time dimension to the agent communication system; S202, the leading agent sends information to the following agent and the auxiliary agent simultaneously. The following agent and the auxiliary agent verify the received information. The leading agent adjusts the path planning information according to the current environmental information and target position; S203, based on the verification of the received information by the following agent and the auxiliary agent in step S202, send the verified error information to the leading agent so that the leading agent can make adjustments according to the collected error information.

5. The intelligent agent control method according to claim 4, wherein In step S202, the information received by the following agent and the auxiliary agent includes a timestamp, a serial number, and path information.

6. The intelligent agent control method according to claim 4, wherein Adjusting the path planning information includes: S301, the auxiliary agent collects real-time environmental information, the following agent collects its own status information, and the leading agent integrates the environmental information of the auxiliary agent and the own status information of the following agent; S302, based on the integrated environmental information and the status information of the following agent in step S301, identify path decision points; S303, based on the identified path decision points, the leading agent generates six paths according to the environmental information provided by the auxiliary agent; S304, based on the generated six paths, set up a virtual random number generator that can generate random integers between 1 and 6. When the leading agent moves to the decision point, dynamically select a path according to the leading agent's use of the virtual random number generator.

7. The intelligent agent control method according to claim 6, wherein, According to the dynamic selection of the path by the leading agent, transmit the dynamically selected path to the leading agent using the virtual random number generator for dynamic selection, so that the leading agent can perform cooperative control on the agent queue.

8. The intelligent agent control method according to claim 6, wherein The path decision point is that when the agent encounters several paths during driving, it selects a path based on these several paths.

9. An agent control device is applied to an agent control method according to any one of claims 1-8, characterized in that, It includes an acquisition module, a control module, a preprocessing module, and a training module. The acquisition module is connected to the preprocessing module, the preprocessing module is connected to the control module, and the control module is connected to the training module.

10. An intelligent agent storage medium is applied to an intelligent agent control method according to any one of claims 1-8, characterized in that A computer program is stored thereon.

Citation Information

Patent Citations

  • Multi-agent formation control method, device and multi-agent system

    CN114326759B