Preset-Time Formation Control Method for Multi-Agent Systems Based on Fuzzy Reinforcement Learning
Through the combination of fuzzy reinforcement learning and graph theory network, robust control strategies are designed, and the problem that multi-agent systems are difficult to achieve target formation within the expected time in complex environments is solved, and precise formation control and anti-interference capabilities are achieved within the preset time.
Patent Information
- Application Number
- CN202510279102.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Traditional multi-agent system control methods are difficult to ensure that the system achieves the target formation effect within the expected time in complex dynamic environments, and there is a problem of insufficient anti-interference ability.
Using a method based on fuzzy reinforcement learning, a robust control strategy is designed by constructing a fuzzy logic system and graph theory network, combining the optimal formation control theory and auxiliary functions, the precise formation control of the multi-agent system within the preset time is realized.
The robust control of multi-agent systems is realized in complex environments, ensuring that the target formation is reached within the preset time, and it has strong anti-interference ability, improving the system's adaptability and control accuracy.
Smart Images

Figure CN119781302B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of formation control of multi-agent systems, and specifically relates to a preset-time formation control method for multi-agent systems based on fuzzy reinforcement learning. Background Art
[0002] With the continuous progress of technology, multi-agent systems are increasingly widely used in fields such as unmanned aerial vehicle formations and robot teams. However, in these complex dynamic environments, achieving efficient coordination and formation control among agents still faces many challenges. Traditional multi-agent system control methods often perform poorly in dealing with dynamic uncertainties and external disturbances, and it is difficult to ensure that the system achieves the expected formation effect within a specified time. In addition, environmental uncertainties and non-linear interactions among agents further increase the control difficulty.
[0003] To solve these problems, there is a robust preset-time controller that combines a neural network and an output regulator in the prior art, realizing a preset-time tracking control method for heterogeneous multi-agent systems based on neural networks. Although it can optimize the cooperation strategy of agents through an adaptive learning mechanism and ensure that the system reaches the target formation within a preset time in a complex environment, such methods still have limitations in dealing with uncertainties such as system parameter or environmental changes; there are also control strategies for multi-agents by designing fuzzy logic systems in the prior art. According to the ability of fuzzy logic systems to handle uncertainties and fuzzy information, a dynamic adaptive update rate is provided to improve the efficiency of multi-agent formation, but such methods have limitations in the cooperative control among agents.
[0004] In view of the above problems, the present invention proposes a preset-time formation robust control method for multi-agent systems based on fuzzy reinforcement learning, effectively overcoming the limitations of traditional control methods, improving the adaptability and control accuracy of multi-agent systems in dynamic environments, and providing strong technical support for the formation tasks of multi-agent systems such as unmanned aerial vehicle swarms and robot teams. Summary of the Invention
[0005] In order to achieve the robust control of multi-agent systems in complex environments, so as to ensure that the system reaches the target formation within a preset time and has strong anti-interference ability, this application provides a preset-time formation control method for multi-agent systems based on fuzzy reinforcement learning.
[0006] In a first aspect, this application provides a preset-time formation control method for multi-agent systems based on fuzzy reinforcement learning, including:
[0007] S1. Based on graph theory, establish a communication network among multi-agents;
[0008] S2. Construct a fuzzy logic system;
[0009] S3. Construct a multi-agent system, including: constructing the state equations of multi-agents, constructing the formation error equations based on the multi-agent state equations, and designing a performance index function that introduces formation errors;
[0010] S4. With the goal of multi-agents achieving target formation control within a preset time, construct an auxiliary function and an error transformation function;
[0011] S5. Define the optimal formation control theory, including defining the optimal performance index function after updating the performance index function using the error change function, and introducing the HJB equation to solve for the optimal control input;
[0012] S6. Apply the ACI structure, and design identifiers, critics, actors, and corresponding update rules in combination with the fuzzy logic system to complete the update of the parameter matrix for optimal formation control.
[0013] By adopting the above scheme, the fuzzy logic system is used to accurately calculate the output of multi-agents in a complex dynamic environment; an auxiliary function is constructed to assist the multi-agent system to complete accurate formation control within a preset time; combined with the optimal formation control theory, the optimal control input is obtained; the update rules of identifiers, critics, and actors are designed to realize the update of the parameter matrix for optimal formation control and achieve robust control.
[0014] Preferably, the specific steps of S1 include:
[0015] S11. Define the vertex set , indicating the number of agents;
[0016] S12. Define the edge set , indicating the communication connection between agents;
[0017] S13. Define the adjacency matrix , indicating the communication connection between agent and agent . When there is a communication connection , otherwise, ;
[0018] S14. Define an undirected connected graph , characterizing the communication network topology of the multi-agent system; where , indicating the Laplacian matrix between the agent and the leader; Represents the communication matrix between the agents and the leader. Among them, it is assumed that at least one agent is connected to the leader, that is .
[0019] By adopting the above scheme, the vertex set, edge set, adjacency matrix and undirected connected graph are defined to effectively characterize the communication network structure of the multi-agent system, ensure that each agent can accurately receive and send information, avoid communication delays or interruptions, and thus improve the formation control performance of the entire multi-agent system.
[0020] Preferably, the S2 step specifically includes:
[0021] S21. Establish a fuzzy rule base; the rules in the fuzzy rule base are in the form: if is , is , is , then is ; where represents the input, represents the output, represents the number of fuzzy rules; represents the fuzzy set , represents 's membership function;
[0022] S22. Calculate the output quantity using singleton fuzzification, product inference and center-average defuzzification; the calculation formula includes:
[0023] where is the total number of fuzzy rules; satisfies , let
[0024] where and , is expressed as:
[0025] .
[0026] By adopting the above scheme, a detailed fuzzy rule base is established and the methods of singleton fuzzification, product inference and center-average defuzzification are used to improve the calculation efficiency and enhance the robustness and stability of the system.
[0027] Preferably, the S3 step includes:
[0028] S31. Construct the state equation of the multi-agent system, and the formula is: where Indicates the system state, Indicates the control output, Indicates an unknown continuous non - linear function, Indicates external interference;
[0029] S32. Construct a coordinate transformation equation:
[0030] Among them, Is the tracking error, Is the leader state, Indicates the relative position between the leader and the agent ;
[0031] S33. Construct a formation error equation, the formula is: Among them, , if there is a protocol such that For , Is a stable time, and Is the desired accuracy; Indicates the adjacency matrix of the i - th follower agent; Indicates the communication link weight between the i - th agent and its neighbor agents; Indicates the communication link weight between the leader and the i - th agent;
[0032] S34. Construct a performance index function, the formula is:
[0033] Among them, Indicates the discount factor, And Indicates a symmetric positive definite matrix; for a multi - agent system, if Is continuous, , Is stable on the set , Is finite, then on Allows a control protocol , expressed as .
[0034] By adopting the above - mentioned scheme, the states and their dynamic characteristics of each agent are constructed in detail, and the relative position relationship between the leader and the agent is reflected through the coordinate transformation equation, thereby quantifying the error in the formation process. Combining the constructed performance index provides a basis for subsequent robust control of the formation within a preset time.
[0035] Preferably, the S4 step includes:
[0036] S41. To achieve the preset time control performance of the multi-agent system, construct the auxiliary function M:
[0037] where is the design parameter, is the specified time; M is strictly decreasing on and when then and ; M is smooth and M is bounded at all times on ;
[0038] S42. Construct the error transformation function , is constructed as follows: where is a constant and .
[0039] By adopting the above scheme, it is ensured that the constructed auxiliary function is strictly decreasing and finally tends to zero within the preset time, and the auxiliary function is used to adjust the error transformation function, so as to effectively control the formation process of the multi-agent system to more accurately reach the target formation state within the predetermined time.
[0040] Preferably, the step S5 includes:
[0041] S51. Update the performance index function according to the error transformation function, and the updated one is further expressed as: where is expressed as through and substitute into the obtained result;
[0042] S52. Take the optimal group control to obtain the optimal performance index function: ;
[0043] S53. Calculate the time derivative of the optimal performance index function to obtain the HJB equation: where , , and is the unique solution of the HJB equation. When solving , the optimal control input is obtained: ;
[0044] S54. To achieve optimal swarm control, is divided into: where, is a design parameter, , ;
[0045] where, , , , substituting into , the optimal control input is obtained: ; where the unknown parameters and are continuous. For and , there exist and such that: where, and both represent the optimal parameter matrices; and are the fuzzy rule numbers; and represent the fuzzy basis function vectors; the approximation errors and correspondingly satisfy and ; and are constants.
[0046] By adopting the above scheme, the time derivative of the optimal performance index function is calculated, the HJB equation is derived, and the equation is solved to obtain the optimal control input, ensuring that the multi-agent system reaches the optimal formation control within the preset time; the optimal control input is divided into multiple parts, and the influences of the design parameters and the fuzzy basis functions are considered separately to further ensure the accuracy and robustness of the control output.
[0047] Preferably, the S6 step includes:
[0048] S61. Apply the ACI structure to design the identifier as: where, and respectively represent the outputs of the FLS and the identifier parameter matrix, and the identifier update rule is: where, represents the design parameter, represents the positive definite matrix;
[0049] S62. Design the critic to evaluate the control performance, and the formula is: where, For the critic parameter matrix; the critic update rule is designed as follows: Among them, and are design parameters;
[0050] S63. The design actor is used to implement the control behavior, and the formula is: Among them, represents the actor parameter matrix, and the design actor update rule is: Among them, is the design parameter;
[0051] S64. If falls on the values of 0 eigenvectors, the training is terminated.
[0052] By adopting the above scheme, the network update law of the design control strategy is designed to update the control parameter matrix and ensure the robustness of the control input.
[0053] Preferably, the step S41 further includes:
[0054] Using the fuzzy logic system to calculate the output of each agent at the current moment, comparing the output of each agent at the current moment with the preset threshold range to determine the state of the agent; the states of the agent include good, general, and poor;
[0055] Judging whether the proportion of the number of multi-agents with the determined state of poor is greater than the first preset ratio; if it is greater, on the basis of the originally constructed auxiliary function, select the first preset coefficient to adjust the preset specified time to extend the preset specified time, and obtain the adjusted auxiliary function, and the formula is:
[0056] Among them, is the first preset coefficient;
[0057] Judging whether the proportion of the number of multi-agents with the determined state of good is greater than the second preset ratio; if it is greater, on the basis of the originally constructed auxiliary function, select the second preset coefficient to adjust the preset specified time to shorten the preset specified time, and obtain the adjusted auxiliary function, and the formula is:
[0058] Among them, is the second preset coefficient.
[0059] By adopting the above scheme, the working state of each agent is judged by comparing the real-time output with the preset threshold, and then the preset time in the preset time formation control of the multi-agent system is dynamically adjusted to improve the adaptability and robustness of the system.
[0060] Preferably, the step S41 further includes:
[0061] Collect current environmental data; match a preset threshold range from a preset output threshold library according to the collected environmental data, and use the matched preset threshold range as the preset threshold range for comparing with the output of each agent; different preset threshold ranges under different environmental data conditions are stored in the preset output threshold library.
[0062] By adopting the above solution, the preset threshold range is dynamically adjusted according to the real-time output of the agent under different environmental conditions, so as to more accurately determine the working state of the agent.
[0063] Preferably, constructing the formation error equation in the step S33 further includes:
[0064] Determine the formation scenario of the target formation control, and the formation scenario includes: tight formation, loose formation;
[0065] Match a preset formation error coefficient weight combination according to the determined formation scenario of the target formation control; the preset formation error coefficient weight combination matched by the tight formation is the first weight combination, and the preset formation error coefficient weight combination matched by the loose formation is the second weight combination; wherein, compared with the second weight combination, the ratio of the weight corresponding to the tracking error between the leader and the agent and the weight corresponding to the tracking error between the agents in the first weight combination is larger;
[0066] Use the matched preset formation error coefficient weight combination to assign values to the weights of the constructed formation error equation.
[0067] By adopting the above solution, the weights in the formation error equation are flexibly adjusted according to different formation scenarios, improving the accuracy and stability of the formation control.
[0068] In summary, the present application has the following beneficial effects:
[0069] By constructing a fuzzy logic system to handle uncertainties for evaluating the accuracy of the agent state; formulating a prescribed time control law to ensure that the multi-agent system reaches the target formation within a predetermined time; combining a reinforcement learning algorithm to complete the optimal cooperation strategy between autonomous learning agents; designing a robust control strategy to enable the formation to suppress interference effects during movement, so as to achieve accurate formation control of the multi-agent system within a specified time in a complex environment and have interference suppression performance at the same time. Description of the Drawings
[0070] Figure 1 It is a flowchart of the method in a specific embodiment;
[0071] Figure 2 It is a flowchart of step 1 in the method in a specific embodiment;
[0072] Figure 3 It is the flowchart of step 2 in the method described in the specific embodiment;
[0073] Figure 4 It is the flowchart of step 3 in the method described in the specific embodiment;
[0074] Figure 5 It is the flowchart of step 4 in the method described in the specific embodiment;
[0075] Figure 6 It is the flowchart of step 5 in the method described in the specific embodiment;
[0076] Figure 7 It is the flowchart of step 6 in the method described in the specific embodiment. Detailed implementation manners
[0077] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0078] In the prior art, it is difficult for a multi-agent system to complete formation control within a preset time, especially in the face of complex environments and uncertain factors. Therefore, as Figure 1 shown, the embodiment of the present application discloses a method for preset-time formation control of a multi-agent system based on fuzzy reinforcement learning to achieve the effect of efficient and stable formation control of the multi-agent system in a complex environment; the specific steps include:
[0079] S1. Based on graph theory, establish a communication network among multi-agents.
[0080] Through the method of graph theory, construct a communication network of a multi-agent system to ensure smooth information transmission among agents. Specifically, as Figure 2 shown, the specific steps include:
[0081] S11. Define the vertex set , indicating the number of agents.
[0082] S12. Define the edge set , indicating the communication connection between agents.
[0083] S13. Define the adjacency matrix , indicating the communication connection between agent and agent . When there is a communication connection , otherwise, ; The set of neighbor nodes is represented by .
[0084] S14. Define an undirected connected graph, characterizing the communication network topology of the multi-agent system; where , represents the Laplacian matrix between the agent and the leader; represents the communication matrix between the agent and the leader, where it is assumed that at least one agent is connected to the leader, i.e., .
[0085] S2. Construct a fuzzy logic system.
[0086] The fuzzy logic system is used to calculate the output of the multi-agent in real time, evaluate the accuracy of the multi-agent state, so as to cope with the changes in the complex environment and reduce uncertainty. Specifically, as Figure 3 shown, the construction steps are as follows:
[0087] S21. Establish a fuzzy rule base.
[0088] Specifically, define the IF-THEN rules in the fuzzy rule base, and the rule form is: If is , is , is , then is ; where represents the input, represents the output, represents the number of fuzzy rules; represents the fuzzy set , represents 's membership function;
[0089] S22. Use singleton fuzzification, product inference, and center-average defuzzification to calculate the output ; The calculation formula includes:
[0090] where is the total number of fuzzy rules; satisfies , let
[0091] where and , is expressed as:
[0092] .
[0093] S3. Construct a multi-agent system.
[0094] To achieve robust formation control of multi-agents, the most important thing is to construct the mathematical model of the multi-agent system, including the state equation, formation error equation, and performance index function. As Figure 4 shown, the specific construction steps are as follows:
[0095] S31. Construct the state equation of the multi-agent system, and the formula is: where, represents the system state, represents the control output, represents the unknown continuous nonlinear function, represents the external disturbance.
[0096] S32. Construct the coordinate transformation equation:
[0097] where, is the tracking error, is the leader state, represents the relative position between the leader and the agent ; construct the mode description as , .
[0098] S33. Construct the formation error equation, and the formula is: where, , if there exists a protocol such that for , is a stable time, and is the desired accuracy; represents the adjacency matrix of the i-th follower agent; represents the communication link weight between the i-th agent and its neighbor agents; represents the communication link weight between the leader and the i-th agent.
[0099] S34. Construct the performance index function, and the formula is:
[0100] where, represents the discount factor, and represent symmetric positive definite matrices; for the multi-agent system, if is continuous, , is stable on the set , is finite, then on the control protocol is allowed , expressed as .
[0101] S4. With the goal of multi-agent achieving target formation control within a preset time, construct an auxiliary function and an error transformation function.
[0102] As Figure 5 shown, the specific steps include:
[0103] S41. To achieve the preset time control performance of the multi-agent system, construct an auxiliary function.
[0104] Construct the auxiliary function M:
[0105] Wherein, is a design parameter, is a specified time; M strictly decreases on , and when , , and ; M is smooth, and M and are bounded at all times on ;
[0106] S42. Construct the error transformation function , is constructed as follows: Wherein, is a constant, and .
[0107] S5. Define the optimal formation control theory, and solve to obtain the optimal control input based on the optimal formation control theory.
[0108] As Figure 6 shown, the specific steps include:
[0109] S51. Update the performance index function according to the error transformation function, and the updated one is further expressed as: Wherein, is expressed as through , and substitute into the obtained result.
[0110] S52. Define the optimal swarm control, that is, take the optimal swarm control , and obtain the optimal performance index function: ;
[0111] S53. Calculate the time derivative of the optimal performance index function to obtain the HJB equation: Among them, , , , and is the unique solution of the HJB equation. When solving , the optimal control input is obtained: ;
[0112] S54. Adjust and optimize the optimal control input. To achieve optimal swarm control, is divided into: Among them, is a design parameter, , ;
[0113] Among them, , , . Substitute into to obtain the optimal control input: ; Among them, the unknown parameters and are continuous. For and , there exist and such that: Among them, and both represent the optimal parameter matrix; and are the fuzzy rule numbers; and represent the fuzzy basis function vectors; the approximation errors and correspond to satisfying and ; and are constants.
[0114] S6. Apply the ACI structure, design the identifier, critic, actor and their corresponding update laws in combination with the fuzzy logic system, and complete the update of the parameter matrix for optimal formation control.
[0115] As Figure 7 shown, the specific steps include:
[0116] S61. Apply the ACI structure and design the identifier as: Among them, and respectively represent the outputs of the FLS and the identifier parameter matrix, and the identifier update rule is designed as follows: Among them, represents the design parameter, represents a positive definite matrix;
[0117] S62. The design critic is used to evaluate the control performance, and the formula is: Among them, is the critic parameter matrix; the design critic update rule is: Among them, and are design parameters;
[0118] S63. The design actor is used to achieve the control behavior, and the formula is: Among them, represents the actor parameter matrix, and the design actor update rule is: Among them, is the design parameter;
[0119] S64. If falls on the values of 0 eigenvectors, the training is terminated.
[0120] Using the above S1 - S6, the precise control of the preset - time formation can be completed. First, set , and initialize , set the parameters , and use S4 - S6 to calculate in sequence. Combining the determined fuzzy membership and fuzzy basis function vectors, calculate and obtain , and update .
[0121] A specific embodiment, different from the above - mentioned embodiment, is that a monitoring and adjustment mechanism for the agent state is added to further improve the adaptability of the multi - agent system in a complex environment and optimize the time and resource utilization rate during the control process. Specifically, the method further includes:
[0122] The step S41 further includes:
[0123] Using the fuzzy logic system to calculate the output quantity of each agent at the current moment, comparing the output quantity of each agent at the current moment with the preset threshold range to determine the state of the agent; the states of the agent include good, general, and poor; the preset threshold range can be specifically determined according to the statistical range of the real - time output quantities of historical agents in different states;
[0124] Determine whether the proportion of multi - agents with a determined poor state is greater than a first preset ratio; if it is greater, it indicates that the states of most multi - agents at the current moment are poor, and the difficulty coefficient of achieving precise control of the formation within the preset specified time increases. It is possible to choose to extend the preset time to ensure the safety and stability of the formation. Design a dynamic time adjustment factor to adjust the preset specified time, that is, on the basis of the originally constructed auxiliary function, select a first preset coefficient to adjust the preset specified time to extend the preset specified time, and obtain the adjusted auxiliary function, denoted as the first auxiliary function, and the formula is:
[0125] Wherein, is the first preset coefficient, which can be set manually;
[0126] Determine whether the proportion of multi - agents with a determined good state is greater than a second preset ratio; if it is greater, it indicates that the states of most multi - agents at the current moment are good, and the difficulty coefficient of achieving precise control of the formation within the preset specified time is less. It is possible to choose to shorten the preset time to improve the formation efficiency. Design a dynamic time adjustment factor to adjust the preset specified time, that is, on the basis of the originally constructed auxiliary function, select a second preset coefficient to adjust the preset specified time to shorten the preset specified time, and obtain the adjusted auxiliary function, denoted as the second auxiliary function, and the formula is:
[0127] Wherein, is the second preset coefficient, which can be set manually.
[0128] In addition, considering that during the statistical determination process of the real - time output range of historical agents in different states, environmental factors have an impact on the real - time output range in different states. To further improve the adaptive ability and formation accuracy of the multi - agent system, the step S41 further includes:
[0129] It is possible to use sensors to collect current environmental data; match the preset threshold range from the preset output threshold library according to the collected environmental data, and use the matched preset threshold range as the preset threshold range for comparing with the output of each agent; the preset output threshold library stores preset threshold ranges under different environmental data conditions.
[0130] In a specific embodiment, in order to meet the different requirements for formation control in different application scenarios and improve the overall formation effect, the method further includes: the construction of the formation error equation in the step S33 further includes:
[0131] Determine the formation scenario of the target formation control, and the formation scenario includes: tight formation, loose formation;
[0132] Match the formation scenario determined by the target formation control with a preset combination of formation error coefficient weights; the tight formation matches the preset combination of formation error coefficient weights as the first weight combination, such as: ; the loose formation matches the preset combination of formation error coefficient weights as the second weight combination, such as: ; Considering that for a tight formation, increasing the weights of the tracking errors between the leader and the agents and between the agents and agents can ensure that the distance between formation members is more compact and reduce error accumulation; for a loose formation, appropriately reducing these weights can make the formation looser and adapt to larger space requirements; therefore, compared with the second weight combination, the ratio of the weight corresponding to the tracking error between the leader and the agents and the weight corresponding to the tracking error between the agents and agents in the first weight combination is set to be larger;
[0133] Use the matched preset combination of formation error coefficient weights to assign values to the weights of the constructed formation error equation.
[0134] The above are all preferred embodiments of the present application. Without limiting the protection scope of the present application accordingly, any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.
Claims
1. A preset-time formation control method for multi-agent systems based on fuzzy reinforcement learning, characterized in that Including: S1. Based on graph theory, establish a communication network among multiple agents; S2. Construct a fuzzy logic system; S3. Construct a multi-agent system, including: constructing the state equation of multiple agents, constructing the formation error equation based on the multi-agent state equation, and designing a performance index function introducing formation error; S4. Aiming at the goal of completing the target formation control by multiple agents within a preset time, construct an auxiliary function and an error transformation function; S5. Define the optimal formation control theory, including: defining the optimal performance index function after updating the performance index function using the error change function, and introducing the HJB equation to solve the optimal control input; S6. Apply the ACI structure, combine the fuzzy logic system to design identifiers, critics, actors and corresponding update rules, and complete the update of the parameter matrix for optimal formation control; The step S3 includes: S31. Construct the state equation of the multi-agent system, and the formula is: where represents the system state, represents the control output, represents the unknown continuous nonlinear function, represents the external disturbance; S32. Construct a coordinate transformation equation: Among them, is the tracking error, is the leader state, represents the relative position between the leader and the agent ; S33. Construct the formation error equation, with the formula: where, , if there exists a protocol such that For , is a stable time, and is the desired accuracy; represents the adjacency matrix of the i-th follower agent; represents the communication link weight between the i-th agent and its neighbor agents; represents the communication link weight between the leader and the i-th agent; S34. Construct a performance index function, and the formula is: Among them, represents the discount factor, and represents a symmetric positive definite matrix; for a multi-agent system, if is continuous, , is stable on the set , is finite, then a control protocol is allowed on , denoted as ; The step S4 includes: S41. In order to achieve the preset time control performance of the multi-agent system, construct an auxiliary function M: Among them, is a design parameter, is a specified time; M is strictly decreasing on and when t = T, ; when 0 ≤ t < T, as t approaches T infinitely, there exists a value of M that approaches 1 + infinitely; M is smooth and M is bounded at all times. S42. Construct an error transformation function , is constructed as follows: wherein, is a constant and ; The step S41 further includes: Calculate the output quantity of each agent at the current moment using the fuzzy logic system, and compare the output quantity of each agent at the current moment with the preset threshold range to determine the state of the agent; the states of the agent include good, general and poor; Judge whether the proportion of the number of multi-agents whose states are determined to be poor is greater than the first preset proportion; if it is greater, on the basis of the originally constructed auxiliary function, select the first preset coefficient to adjust the preset specified time to extend the preset specified time, and obtain the adjusted auxiliary function, and the formula is: Among them, is the first preset coefficient; Judge whether the proportion of the number of multi-agents whose states are determined to be good is greater than the second preset proportion; if it is greater, on the basis of the originally constructed auxiliary function, select the second preset coefficient to adjust the preset specified time to shorten the preset specified time, and obtain the adjusted auxiliary function, and the formula is: Among them, is the second preset coefficient; The step S41 further includes: Collect the current environmental data; match the preset threshold range from the preset output threshold library according to the collected environmental data, and use the matched preset threshold range as the preset threshold range for comparison with the output quantity of each agent; different preset threshold ranges under different environmental data conditions are stored in the preset output threshold library; The construction of the formation error equation in the step S33 further includes: Determine the formation scenario of the target formation control, and the formation scenario includes: tight formation, loose formation; Match the preset formation error coefficient weight combination according to the determined formation scenario of the target formation control; the preset formation error coefficient weight combination matched by the tight formation is the first weight combination, and the preset formation error coefficient weight combination matched by the loose formation is the second weight combination; among them, compared with the second weight combination, the ratio of the weight corresponding to the tracking error between the leader and the agent and the weight corresponding to the tracking error between the agent and the agent in the first weight combination is larger; Use the matched preset formation error coefficient weight combination to assign values to the weights of the constructed formation error equation.
2. The preset-time formation control method for a multi-agent system based on fuzzy reinforcement learning according to claim 1, wherein The step S1 specifically includes: S11. Define the vertex set , which represents the number of agents; S12. Define the edge set , indicating the communication connections between agents; S13. Define the adjacency matrix , denotes an agent and agent the communication connection therebetween, when there is a communication connection otherwise, ; S14. Define an undirected connected graph , which represents the communication network topology of the multi-agent system; where , denotes the Laplacian matrix between the agents and the leader; denotes the communication matrix between the agents and the leader, where it is assumed that at least one agent is connected to the leader, i.e., .
3. The preset-time formation control method for a multi-agent system based on fuzzy reinforcement learning according to claim 1, characterized in that The step S2 specifically includes: S21. Establish a fuzzy rule base; the rules in the fuzzy rule base are in the form: If is , is , is , then is ; where represents the input, represents the output, represents the number of fuzzy rules; represents the fuzzy set , represents 's membership function; S22. Use singleton fuzzification, product inference, and center average defuzzification to calculate the output quantity ; The calculation formula includes: Among them, is the total number of fuzzy rules; Satisfy , let Among them, and , are expressed as: 。 4. The preset-time formation control method for a multi-agent system based on fuzzy reinforcement learning according to claim 1, characterized in that The step S5 includes: S51. Update the performance index function according to the error transformation function, and after the update, it is further expressed as: where through is expressed as , and substitute into what is obtained; S52. Obtain the optimal population control , and obtain the optimal performance index function: ; S53. Calculate the time derivative of the optimal performance index function to obtain the HJB equation: where, , , , and is the unique solution of the HJB equation. When solving , the optimal control input is obtained: ; S54. To achieve optimal swarm control, is divided into: where is a design parameter, , ; Among them, , , , substitute into to obtain the optimal control input: ; among them, the unknown parameters and are continuous. For and , there exist and such that: Among them, and both represent the optimal parameter matrices; and are the fuzzy rule numbers; and represent the fuzzy basis function vectors; the approximation errors and correspondingly satisfy and ; and are constants.
5. The preset-time formation control method for a multi-agent system based on fuzzy reinforcement learning according to claim 4, wherein The step S6 includes: S61. Apply the ACI structure, and the design identifier is: Wherein, and respectively represent the outputs of the FLS and the identifier parameter matrix, and the design identifier update rule is: Wherein, represents the design parameter, represents a positive definite matrix; S62. The design critic is used to evaluate the control performance, and the formula is: Among them, is the critic parameter matrix; the update rule of the design critic is: Among them, and are design parameters; S63. Design an actor to achieve control behavior, with the formula: Among them, represents the actor parameter matrix, and the designed actor update rule is: Among them, is the design parameter; S64. If falls on the values of 0 eigenvectors, the training terminates.