Multi-target test flat car cooperative control method and system
Through formal modeling and multi-objective reward function methods, the optimal control strategy is generated, which solves the problem of unstable control and difficult work in complex environments for autonomous driving test flatbeds, and realizes the stable and coordinated operation of flatbed vehicles in complex environments and efficient arrival of target positions.
Patent Information
- Application Number
- CN202510169327.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the control strategy optimization method of autonomous driving test flatbed vehicles is unstable, and the number of iterations is numerous, making it difficult to adapt to the complex and changing driving environment and the demand for the coordinated work of multiple flatbed vehicles, resulting in the flatbed vehicles entering areas with many obstacles, increasing the risk of collision, and being unable to quickly reach the target position.
Through formal modeling of the flatbed car collaborative control process, the state space and action space are constructed, the multi-objective reward function is determined, the optimal control strategy under different state information is generated, and the flatbed car collaborative control is carried out in combination with the current state information.
Effectively avoid flatbed vehicles entering areas with many obstacles, reduce the probability of collision, prompt the flatbed vehicles to reach the preset target position as soon as possible, improve overall transportation efficiency, and improve the stability of the update of strategy function values, and reduce the number of iterations.
Smart Images

Figure CN120044850A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of flatbed vehicle control, and more specifically, relates to a multi-objective test flatbed vehicle cooperative control method and system. Background Art
[0002] With the rapid development of autonomous driving technology, the construction of intelligent transportation systems has become an important development direction. As an important part of this field, test flatbed vehicles have broad application prospects. They can not only efficiently transport goods in logistics and manufacturing, but also play an important role in scenarios such as smart cities and unmanned delivery. During the autonomous driving test process, due to the influence of various factors such as traffic conditions and vehicle driving trajectories, how to efficiently and accurately coordinate the movement trajectories of flatbed vehicles has become an important problem that needs to be solved urgently. The existing control strategy optimization methods are unstable, have a large number of iterations, and are inefficient. It is difficult to adapt to the complex and changeable driving environment and the requirements of multiple flatbed vehicles working together. It is difficult to comprehensively make control decisions based on various factors during the driving of flatbed vehicles (such as the distance from the target position, collision situation, etc.), which may cause flatbed vehicles to enter areas with more obstacles, increasing the collision risk and making it impossible to reach the target position quickly.
[0003] The invention patent with the publication number CN116985802A in the prior art proposes a multi-objective dynamic weighting full-speed-range vehicle adaptive cruise control method, which collects the kinematic information of the vehicle itself and the vehicle in front; gives a method for establishing a longitudinal following discrete state space model considering the acceleration interference of the vehicle in front, and formulates control objectives of safety, following performance, economy and comfort; based on the objective benefit functions of each control objective, solves the probability distribution of the mixed strategy through non-cooperative game and uses it as the objective weight coefficient; considering the driver's subjective preference for vehicle control objectives, establishes a subjective weight model based on the blind number theory to solve the subjective weight coefficient; based on game theory, establishes a subjective and objective combined weighting model, dynamically assigns the optimal subjective and objective weights to the model predictive controller for real-time solution, and obtains the optimal expected acceleration under the subjective and objective conditions. This solution cannot handle complex environmental areas with more obstacles. Summary of the Invention
[0004] The present invention aims to overcome the problems in the prior art that the control strategy optimization method is unstable, has a large number of iterations, and is difficult to adapt to the complex and changeable driving environment and the cooperation of multiple flatbed vehicles, and provides a multi-objective test flatbed vehicle cooperative control method and system.
[0005] The primary object of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:
[0006] The first aspect of the present invention provides a multi-objective test flatbed vehicle cooperative control method, including the following steps:
[0007] Formalize the collaborative control process of the test flatbed vehicle by combining the driving direction and speed of the test flatbed vehicle with collision avoidance and dynamic obstacle avoidance requirements to obtain a collaborative control model;
[0008] Construct the state space and action space of the collaborative control process of the test flatbed vehicle by using the collaborative control model, and determine the multi-objective reward function by using the state space and the action space;
[0009] Generate the optimal control strategy of the test flatbed vehicle under different state information by using the state space, action space and objective reward function of the collaborative control process of the test flatbed vehicle;
[0010] Carry out collaborative control of the test flatbed vehicle according to the optimal control strategy of the test flatbed vehicle combined with the current state information.
[0011] Furthermore, the method for formalizing the collaborative control process of the test flatbed vehicle is as follows: Calculate the dynamic control variables in the collaborative control process of the test flatbed vehicle by using the driving direction and speed of the test flatbed vehicle, and obtain the representation forms of the driving direction and speed of each test flatbed vehicle at any moment. The expressions are as follows:
[0012]
[0013] where n ∈ [1, N], N represents the number of test flatbed vehicles in the test scenario. In this example, the value of N is 5, represents the dynamic control variable of the nth test flatbed vehicle at time t, represents the driving direction of the nth test flatbed vehicle at time t, represents the driving speed of the nth test flatbed vehicle at time t; The collaborative control target expression in the collaborative control process of the test flatbed vehicle is as follows:
[0014]
[0015] where, represents the position of the nth test flatbed vehicle at time t, respectively represent the horizontal position coordinate and the vertical position coordinate in the position ; L n represents the preset target position of the nth test flatbed vehicle, x n , y n respectively represent the horizontal position coordinate and the vertical position coordinate in the position L n ;
[0016] Combined with collision avoidance and dynamic obstacle avoidance requirements, the constraints in the collaborative control process of the test flatbed vehicle are:
[0017]
[0018] Among them, represents the distance between the nth test flatbed vehicle and the nearest obstacle at time t. If is less than or equal to 0, it means that the nth test flatbed vehicle collides with an obstacle at time t; represents the distance between the nth test flatbed vehicle and the nearest test flatbed vehicle at time t. If is less than or equal to 0, it means that the nth test flatbed vehicle collides with a test flatbed vehicle at time t.
[0019] Furthermore, the state space and action space of the collaborative control process of the test flatbed vehicle are constructed using the collaborative control model. Among them, the action space includes the set of control actions of the test flatbed vehicle in the collaborative control process, and the set of control actions includes the driving direction control action and the driving speed control action of the test flatbed vehicle; the state space includes all state information of the test flatbed vehicle in different driving environments, and the state information includes the position, driving direction, driving speed, and surrounding environment information of the test flatbed vehicle. The expression of the state information of the nth test flatbed vehicle at time t is as follows:
[0020]
[0021] Among them, represents the position of the nth test flatbed vehicle at time t; represents the state information of the nth test flatbed vehicle at time t; represents the surrounding environment information of the nth test flatbed vehicle at time t; represents the driving direction of the nth test flatbed vehicle at time t; represents the driving speed of the nth test flatbed vehicle at time t; the surrounding environment information is presented in the form of an environment matrix, represents the value of the environment matrix element at position i. Position 5 is the position where the nth test flatbed vehicle is located at any time t. The expression is as follows:
[0022]
[0023] Among them,
[0024]
[0025] , and the center of the matrix is the position where the nth test flatbed vehicle is located at time t.
[0026] Furthermore, the expression of the multi-objective reward function is as follows:
[0027]
[0028] Among them, R(·) represents the multi-objective reward function for the collaborative control process of the test flatbed vehicle; A t represents the set of control actions adopted by N test flatbed vehicles at time t, represents the control action of the nth test flatbed vehicle at time t; represents the state information of the nth test flatbed vehicle at time t.
[0029] Furthermore, the multi-objective reward function value includes: the distance reward value of the preset target position, the collision reward value, the target position reward value, the collaborative control reward value of multiple test flatbed vehicles, and the environmental complexity reward value. Then, the expression of the target reward function is as follows:
[0030]
[0031] Among them, the expression of the distance reward value of the preset target position is as follows:
[0032]
[0033] The expression of the collision reward value is as follows:
[0034]
[0035] The expression of the target position reward value is as follows:
[0036]
[0037] The expression of the collaborative control reward value of multiple test flatbed vehicles is as follows:
[0038]
[0039] The expression of the environmental complexity reward value is as follows:
[0040]
[0041] Among them, represents the control action of the nth test flatbed vehicle at time t; represents the state information of the nth test flatbed vehicle at time t; ||·|| 2 represents the L2 norm; represents the position of the nth test flatbed vehicle at time t + 1 after adopting the control action ; L n represents the preset target position of the nth test flatbed vehicle; represents the distance between the nth test flatbed vehicle and the nearest obstacle at time t + 1 after adopting the control action ; Denote the distance between the nth test flatbed vehicle and the nearest test flatbed vehicle at time t+1 after adopting the control action at time t; exp(·) represents the exponential function with the natural constant as the base; Denote the distance between the nth test flatbed vehicle and the nearest test flatbed vehicle at time t+1 after adopting the control action at time t; exp(·) represents the exponential function with the natural constant as the base; Denote the environment matrix after the nth test flatbed vehicle adopts the control action at time t. Denote the environment matrix after the nth test flatbed vehicle adopts the control action at time t.
[0042] Furthermore, the method for generating the control strategy of the test flatbed vehicle under different state information by using the state space, action space and target reward function of the collaborative control process of the test flatbed vehicle includes the following steps:
[0043] Calculate the initial policy function value by using the state space and action space of the collaborative control process of the test flatbed vehicle, and obtain the initial policy function value of the control action under each state information;
[0044] Combine the initial policy function value with the target reward function, calculate the expected reward value of the control action under each state information, and determine the action advantage value under different state information by using the expected reward value;
[0045] Input the expected reward value and the action advantage value into the policy function, and output the policy function value and the policy ratio at the current iteration number;
[0046] Iterate the policy function value by using the action advantage value and the policy ratio to obtain the final policy function value of selecting different control actions under different state information;
[0047] Store the finally obtained policy function value after iteration to form the optimal control strategy of the test flatbed vehicle under different state information.
[0048] Furthermore, iterating the policy function value by using the action advantage value and the policy ratio includes the following steps:
[0049] Standardize the policy ratio;
[0050] Use the standardized policy ratio, action advantage value and discount factor to iteratively update the policy function value and gradually optimize the policy function value;
[0051] Until the maximum iteration number is reached, or the difference between the policy advantage values obtained by adjacent iterations is less than the preset iteration threshold, and the final policy function value is obtained.
[0052] Furthermore, scheduling the test flatbed vehicle for collaborative control according to the control strategy of the test flatbed vehicle includes the following steps:
[0053] Obtain the state information of N test flatbed vehicles at time t;
[0054] Select state information from the optimal test flatbed vehicle control strategy, and calculate the similarity between the selected state information and the state information of the test flatbed vehicle at time t;
[0055] Take the control action with the largest final policy function value associated with the state information with the highest similarity in the optimal test flatbed vehicle control strategy as the control action of the test flatbed vehicle at time t, and perform cooperative control of N test flatbed vehicles according to the control action;
[0056] Repeat the above steps until all test flatbed vehicles reach the preset target position.
[0057] The second aspect of the present invention provides a multi-objective test flatbed vehicle cooperative control system for implementing the steps of a multi-objective test flatbed vehicle cooperative control method. The system includes:
[0058] A model construction module for formally modeling the cooperative control process of the test flatbed vehicle by using the driving direction and driving speed of the test flatbed vehicle in combination with collision avoidance and dynamic obstacle avoidance requirements to obtain a cooperative control model;
[0059] A reward function construction module for constructing the state space and action space of the test flatbed vehicle cooperative control process by using the cooperative control model, and determining a multi-objective reward function by using the state space and action space;
[0060] A control strategy generation module for generating an optimal test flatbed vehicle control strategy under different state information by using the state space, action space and target reward function of the test flatbed vehicle cooperative control process;
[0061] A cooperative control module for performing cooperative control of the test flatbed vehicle according to the optimal test flatbed vehicle control strategy in combination with the current state information.
[0062] Furthermore, the multi-objective reward function values include: a distance reward value for the preset target position, a collision reward value, a target position reward value, a cooperative control reward value for multiple test flatbed vehicles, and an environmental complexity reward value.
[0063] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0064] By comprehensively considering multiple objectives such as the distance to the preset target position, collision situations, whether the target position has been reached, cooperation status, and environmental complexity, the flatbed truck can select the optimal control action under different state information, effectively avoid entering areas with more obstacles, reduce the collision probability, and at the same time prompt the flatbed truck to reach the preset target position as soon as possible, improving the overall transportation efficiency. In addition, the expected reward value is calculated by combining the multi-objective reward function values at multiple consecutive moments, the policy function value is initialized, and the action advantage value is calculated. The policy ratio is constructed using adjacent iteration results to limit the update step size. This method not only improves the stability of the policy function value update, reduces the number of iterations, but also provides a clear direction for the update of the policy function value. Finally, the rapid acquisition of the optimal control action under different state information is realized, and an efficient optimal test flatbed truck control strategy is constructed to ensure the stable and cooperative operation of the flatbed truck in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] To make the objectives and technical solutions of the present invention clearer, the following drawings are provided and described:
[0066] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] In order to more clearly understand the above objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0068] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0069] Embodiment 1:
[0070] The present invention provides a multi-objective test flatbed truck cooperative control method, as Figure 1 shown in a flowchart of a multi-objective test flatbed truck cooperative control method, and the specific steps are as follows:
[0071] S1: Using the driving direction and driving speed of the test flatbed truck, combined with the requirements of collision avoidance and dynamic obstacle avoidance, formalize the cooperative control process of the test flatbed truck to obtain a cooperative control model.
[0072] The specific process is as follows:
[0073] Using the driving direction and speed of the test flatbed vehicle, calculate the dynamic control variables during the cooperative control process of the test flatbed vehicle, and obtain the representation forms of the driving direction and speed of each test flatbed vehicle at any moment. The expressions are as follows:
[0074]
[0075] where n ∈ [1, N], and N represents the number of test flatbed vehicles in the test scenario. represents the dynamic control variable of the nth test flatbed vehicle at time t. represents the driving direction of the nth test flatbed vehicle at time t. represents the driving speed of the nth test flatbed vehicle at time t. The cooperative control objective expression during the cooperative control process of the test flatbed vehicle is as follows:
[0076]
[0077] where represents the position of the nth test flatbed vehicle at time t. respectively represent the horizontal position coordinate and the vertical position coordinate in the position ; L n represents the preset target position of the nth test flatbed vehicle, and x n , y n respectively represent the horizontal position coordinate and the vertical position coordinate in the position L n .
[0078] Combined with the requirements of collision avoidance and dynamic obstacle avoidance, the constraints during the cooperative control process of the test flatbed vehicle are:
[0079]
[0080] where represents the distance between the nth test flatbed vehicle and the nearest obstacle at time t. If is less than or equal to 0, it means that the nth test flatbed vehicle collides with the obstacle at time t. represents the distance between the nth test flatbed vehicle and the nearest test flatbed vehicle at time t. If is less than or equal to 0, it means that the nth test flatbed vehicle collides with the test flatbed vehicle at time t.
[0081] S2: Construct the state space and action space of the cooperative control process of the test flatbed vehicle using the cooperative control model, and determine the multi-objective reward function using the state space and action space.
[0082] Among them, the action space includes the set of control actions of the test flatbed vehicle during the collaborative control process, where the set of control actions includes the driving direction control action and the driving speed control action of the test flatbed vehicle; the state space includes all state information of the test flatbed vehicle in different driving environments, and the state information includes the position, driving direction, driving speed of the test flatbed vehicle, and surrounding environment information. The expression of the state information of the nth test flatbed vehicle at time t is as follows:
[0083]
[0084] Among them, represents the position of the nth test flatbed vehicle at time t; represents the state information of the nth test flatbed vehicle at time t; represents the surrounding environment information of the nth test flatbed vehicle at time t; represents the driving direction of the nth test flatbed vehicle at time t; represents the driving speed of the nth test flatbed vehicle at time t; the surrounding environment information is presented in the form of an environment matrix, represents the value of the environment matrix element at position i, and the expression is as follows:
[0085]
[0086] Among them,
[0087]
[0088] , and the center of the matrix is the position where the nth test flatbed vehicle is located at time t.
[0089] More specifically, the expression of the multi-objective reward function is as follows:
[0090]
[0091] Among them, R(·) represents the multi-objective reward function of the collaborative control process of the test flatbed vehicle; A t represents the set of control actions adopted by N test flatbed vehicles at time t, represents the control action of the nth test flatbed vehicle at time t; represents the state information of the nth test flatbed vehicle at time t.
[0092] In this example, the multi-objective reward function includes: the distance reward value of the preset target position, the collision reward value, the target position reward value, the collaborative control reward value of multiple test flatbed vehicles, and the environmental complexity reward value. Then the expression of the target reward function is as follows:
[0093]
[0094] Among them, the expression of the distance reward value for the preset target position is as follows:
[0095]
[0096] The expression of the collision reward value is as follows:
[0097]
[0098] The expression of the target position reward value is as follows:
[0099]
[0100] The expression of the collaborative control reward value for multiple test flatbed trucks is as follows:
[0101]
[0102] The expression of the environmental complexity reward value is as follows:
[0103]
[0104] Among them, represents the control action of the nth test flatbed truck at time t; represents the state information of the nth test flatbed truck at time t; ||·|| 2 represents the L2 norm; represents the position of the nth test flatbed truck at time t + 1 after adopting the control action ; L n represents the preset target position of the nth test flatbed truck; represents the distance between the nth test flatbed truck and the nearest obstacle at time t + 1 after adopting the control action ; represents the distance between the nth test flatbed truck and the nearest test flatbed truck at time t + 1 after adopting the control action ; exp(·) represents the exponential function with the natural constant as the base; represents the environmental matrix of the nth test flatbed truck after adopting the control action ; The larger it is, the closer the distance to the preset target position after taking the control action ; The larger it is, the smaller the probability of collision after taking the control action ; The larger it is, the more likely it is to reach the preset target position after taking the control action ; The larger it is, the more it indicates that control actions are taken. After that, there are multiple test flatbed trucks in a clustered state, and combined collaborative control can be performed. The larger it is, the more it indicates that control actions are taken. There are more surrounding obstacles after that.
[0105] The reasons for selecting these 5 reward values in this embodiment are as follows:
[0106] 1. The distance reward value for the preset target position is selected because one of the core tasks of the flatbed truck is to reach the preset target position. The distance from the target position is directly related to the progress and efficiency of task completion. By setting the distance reward value, the flatbed truck can be guided to select control actions that can make it approach the target position faster. In a complex driving environment, the flatbed truck may face multiple path choices, and this reward value provides it with a clear optimization direction, that is, to preferentially select actions that can shorten the distance to the target position.
[0107] From the perspective of optimizing control actions, the distance reward value prompts the flatbed truck to continuously adjust its own state (such as driving direction and speed) during driving, so as to move towards the target position. For example, when the current driving direction of the flatbed truck causes the distance from the target position to increase, the distance reward value will decrease, thus prompting the control strategy to adjust the driving direction of the flatbed truck and select a better path.
[0108] Regarding the overall collaborative control effect, each flatbed truck approaching the target position quickly helps to improve the transportation efficiency of the entire flatbed truck cluster. If some flatbed trucks reach the target position quickly due to the guidance of the distance reward value, more driving space and resources can be released, reducing traffic congestion, and thus improving the collaborative working efficiency of the entire cluster.
[0109] 2. The collision reward value is selected because collision is a situation that must be strictly avoided in the autonomous driving scenario. It will not only cause damage to the flatbed truck itself, but may also affect the normal progress of the entire collaborative work, causing traffic jams and even more serious accidents. Therefore, by setting the collision reward value, control actions that may lead to collisions are punished.
[0110] In an environment where multiple flatbed trucks are running simultaneously, the risk of collision between flatbed trucks and between flatbed trucks and obstacles exists at any time. The collision reward value can give an early warning and avoid the occurrence of such dangerous situations to ensure driving safety.
[0111] From a safety perspective, the collision reward value provides safety guarantees for the operation of the flatbed truck. When the control actions taken by the flatbed truck may lead to collisions with obstacles or other flatbed trucks (i.e., when the distance is too close or a collision has already occurred, the reward value is -1), the control strategy will tend to avoid such actions, thereby reducing the collision risk and protecting the safety of the flatbed truck and the surrounding environment.
[0112] In terms of cooperative control, avoiding collisions helps maintain a stable distance and orderly operation between flatbed trucks, ensuring the cooperation of the entire cluster. If flatbed trucks frequently collide, it will disrupt the driving rhythm of the entire cluster, affect the normal driving paths and speed selections of other flatbed trucks, and thus reduce the overall cooperative control effect.
[0113] The reward value for the selected target position is because it is crucial to determine whether the flatbed truck has successfully reached the preset target position for judging whether the task is completed. The target position reward value can directly reflect whether the flatbed truck has achieved the final goal, providing a clear termination signal or a basis for continuous optimization for the control strategy.
[0114] In a complex driving environment and during the multi-objective cooperative control process, the flatbed truck may be interfered by other factors when approaching the target position. The target position reward value can ensure that the flatbed truck finally reaches the target position accurately while meeting other conditions (such as avoiding collisions and maintaining cooperation).
[0115] For a single flatbed truck, when it obtains the target position reward value (1 when it reaches the target position), it means that it has completed the current task. The control strategy can stop further optimization or adjustment of this flatbed truck based on this information and allocate resources to other flatbed trucks that are still in motion.
[0116] In the cluster cooperative control, the flatbed trucks reaching the target position successively helps the orderly progress of the entire task. Each flatbed truck reaching the target position can be regarded as a phased achievement of the entire cooperative task. The target position reward value can help monitor and manage the task progress of the entire cluster, ensuring that all flatbed trucks can reach their respective target positions accurately and efficiently, thereby achieving the successful completion of the overall cooperative task.
[0117] 4. The reward value for the cooperative control of multiple test flatbed trucks is selected because in the scenario where multiple flatbed trucks work cooperatively, the relative positions and state relationships between the flatbed trucks have an important impact on the overall cooperative effect. When the flatbed trucks are in a cluster state, they can perform combined cooperative control, such as synchronously adjusting speeds and maintaining a safe distance. This cooperative method helps improve transportation efficiency, avoid collisions, and optimize traffic flow. Therefore, the cooperative control reward value is set to encourage the flatbed trucks to form a cluster state.
[0118] Different driving environments and mission requirements may require the flatbed trucks to maintain different cluster states at different times. For example, a tight cluster is needed in narrow passages, while a relatively dispersed formation can be adopted in open areas. The collaborative control reward value can be dynamically adjusted according to the distance between the flatbed trucks to meet the collaborative requirements in various scenarios.
[0119] From the perspective of collaborative effect, the collaborative control reward value prompts the flatbed trucks to maintain an appropriate distance relationship and form an orderly cluster. This helps to reduce the mutual interference between the flatbed trucks and improve the driving stability and predictability of the entire cluster. For example, in a convoy, a tight cluster can reduce air resistance and improve energy utilization efficiency. At the same time, a reasonable spacing also facilitates braking and avoidance operations in case of emergencies, enhancing the safety of the entire cluster.
[0120] In terms of path planning and resource allocation in complex environments, the collaborative control reward value can guide the flatbed trucks to select a better driving path based on the positions of their surrounding peers and avoid local congestion. For example, when there are obstacles or traffic jams in some areas, the flatbed trucks can adjust the cluster state (such as increasing the spacing, changing the formation, etc.) to adapt to the environmental changes, ensuring that the entire cluster can pass through the complex area smoothly and improving the robustness of the overall task execution.
[0121] 5. The reward value for environmental complexity is selected because the complexity of the driving environment has an important impact on the control decisions of the flatbed trucks. There are more obstacles and uncertain factors in complex environments, and the flatbed trucks need to choose control actions more carefully. The reward value for environmental complexity can reflect the difficulty of the surrounding environment and provide environmental information reference for the control strategy of the flatbed trucks.
[0122] By quantifying the environmental complexity, the flatbed trucks can adjust their driving strategies according to the reward value. For example, in a complex environment, the speed can be appropriately reduced and the safety spacing can be increased to improve driving safety. While in a relatively simple environment, a more efficient driving mode can be selected to speed up the arrival at the target position.
[0123] In terms of safety guarantee, the reward value for environmental complexity reminds the flatbed trucks to pay attention to potential dangers in the surrounding environment. When the environmental complexity is high (i.e., there are more obstacles around and the absolute value of the reward value is large), the flatbed trucks will drive more carefully to avoid collisions or getting into trouble due to ignoring environmental factors. This helps to improve the survival ability of the flatbed trucks in various complex scenarios and the reliability of task execution.
[0124] For overall collaborative control, considering the environmental complexity can make the collaboration between flatbed trucks more intelligent and adaptable. The environmental complexity levels of different flatbed trucks may vary. Through the environmental complexity reward value, the control strategy can be adjusted individually according to the specific environmental conditions of each flatbed truck to ensure that the entire cluster operates efficiently and safely in a complex and changing environment. For example, when some flatbed trucks face a complex environment, other flatbed trucks can adjust their own states according to the overall collaboration requirements to provide support or avoid interference with each other, improving the ability of the entire cluster to cope with complex environments.
[0125] By comprehensively considering multiple objectives such as the distance to the preset target position, collision situations, whether the target position has been reached, collaboration status, and environmental complexity, the flatbed truck can select the optimal control action under different state information, effectively avoid entering areas with more obstacles, reduce the collision probability, and at the same time prompt the flatbed truck to reach the preset target position as soon as possible, improving the overall transportation efficiency. In addition, by combining the multi-objective reward function values at multiple consecutive moments to calculate the expected reward value, initializing the policy function value and calculating the action advantage value, a policy ratio is constructed using adjacent iteration results to limit the update step size. This method not only improves the stability of the policy function value update, reduces the number of iterations, but also provides a clear direction for the update of the policy function value. Finally, it realizes the rapid acquisition of the optimal control action under different state information, constructs an efficient optimal test flatbed truck control strategy, and ensures the stable collaborative operation of the flatbed truck in a complex environment.
[0126] S3: Generate the optimal test flatbed truck control strategy under different state information using the state space, action space, and target reward function of the test flatbed truck collaborative control process.
[0127] The specific process is as follows:
[0128] Using the state space and action space of the test flatbed truck collaborative control process, combined with the initial value of the target reward function, calculate the initial policy function value to obtain the initial policy function value of the control action under each state information In this embodiment, the value of is 0.1.
[0129] Combine the initial policy function value with the target reward function to calculate the expected reward value of the control action under each state information The expression is as follows:
[0130]
[0131] Among them, represents the control action in the state information The multi-objective reward function value under
[0132] Determine the action advantage value under different state information by using the expected reward value. The expression is as follows:
[0133]
[0134] Where is the state information, is the selected control action, and its action advantage value is
[0135] Input the expected reward value and the action advantage value into the policy function, and output the policy function value and the policy ratio at the current iteration number.
[0136] Set the current iteration number of the policy function value to m, then the m-th iteration result of the policy function value is The corresponding action advantage value is The initial value of m is 0, and the maximum value is M.
[0137] Calculate the policy ratio of the policy function value obtained in the m-th iteration. The policy ratio of the policy function value is:
[0138]
[0139] Where represents the state information, represents the selected control action, represents the policy function value of the policy ratio.
[0140] Perform normalization processing on the policy ratio to limit the amplitude of policy update and prevent drastic changes in the policy function value during the optimization process. The normalization processing formula of the policy ratio is:
[0141]
[0142] Where represents the state information, represents the selected control action, represents the normalized result of the policy ratio ; ∈ represents the ratio control parameter; in this embodiment, ∈ is set to 0.8.
[0143] Use the normalized policy ratio, action advantage value and discount factor to iteratively update the policy function value and gradually optimize the policy function value. The iteration formula of the policy function value is:
[0144]
[0145] Among them, β represents the discount factor. In this embodiment, β is set to 0.1; represents an intermediate parameter in the iterative process; represents the action advantage value; represents the target reward function.
[0146] Until the maximum number of iterations is reached, or the difference between the policy advantage values obtained in adjacent iterations is less than the preset iteration threshold, the final policy function value is obtained.
[0147] Store the final policy function value obtained by iteration to form the optimal test flatbed truck control strategy. The form of the optimal test flatbed truck control strategy formed is:
[0148] {V(S k |A c )|k ∈ [1, K], c ∈ [1, C]}
[0149] Among them: V(S k |A c ) represents the final policy function value of selecting the control action A k under the state information S c ; S k represents the k-th state information in the optimal test flatbed truck control strategy, and K represents the total number of state information in the optimal test flatbed truck control strategy; A c represents the c-th control action in the optimal test flatbed truck control strategy, and C represents the total number of control actions in the optimal test flatbed truck control strategy.
[0150] S4: Coordinately control the test flatbed truck according to the optimal test flatbed truck control strategy combined with the current state information.
[0151] The specific process is as follows:
[0152] Obtain the state information of N test flatbed trucks at time t;
[0153] Select state information from the optimal test flatbed truck control strategy and calculate the similarity between the selected state information and the state information of the test flatbed truck at time t;
[0154] Take the control action with the largest final policy function value associated with the state information with the highest similarity in the optimal test flatbed truck control strategy as the control action of the test flatbed truck at time t, and coordinately control N test flatbed trucks according to the control action;
[0155] Repeat the above steps until all test flatbed trucks reach the preset target position.
[0156] Embodiment 2:
[0157] This embodiment provides a multi - target test flatbed vehicle cooperative control system, which adopts a multi - target test flatbed vehicle cooperative control method as described in Embodiment 1. The system includes:
[0158] A model construction module, which is used to formally model the cooperative control process of the test flatbed vehicle by combining the driving direction and speed of the test flatbed vehicle with the requirements of collision avoidance and dynamic obstacle avoidance, so as to obtain a cooperative control model;
[0159] A reward function construction module, which is used to construct the state space and action space of the cooperative control process of the test flatbed vehicle by using the cooperative control model, and determine the multi - target reward function by using the state space and action space;
[0160] A control strategy generation module, which is used to generate the optimal test flatbed vehicle control strategy under different state information by using the state space, action space and target reward function of the cooperative control process of the test flatbed vehicle;
[0161] A cooperative control module, which is used to perform cooperative control on the test flatbed vehicle according to the optimal test flatbed vehicle control strategy combined with the current state information.
[0162] More specifically, the multi - target reward function values include: the distance reward value of the preset target position, the collision reward value, the target position reward value, the cooperative control reward value of multiple test flatbed vehicles, and the environmental complexity reward value.
[0163] Obviously, the above - mentioned embodiments of the present invention are only examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made on the basis of the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A multi-objective test flatbed vehicle collaborative control method, characterized in that: The steps include: The test flatbed vehicle's driving direction and speed are combined with collision avoidance and dynamic obstacle avoidance requirements to formally model the test flatbed vehicle's cooperative control process and obtain a cooperative control model. The collaborative control model is used to construct the state space and action space of the collaborative control process of the test flatbed vehicle, and the multi-objective reward function is determined using the state space and action space; The state space, action space and target reward function of the test flatbed vehicle cooperative control process are used to generate the optimal test flatbed vehicle control strategy under different state information; The test flatbed vehicle is cooperatively controlled according to the optimal test flatbed vehicle control strategy combined with current state information.
2. A multi-target test flatbed vehicle cooperative control method according to claim 1, characterized in that: The method for formalizing the modeling of the test flatbed vehicle cooperative control process is: using the driving direction and driving speed of the test flatbed vehicle, calculating the dynamic control variables in the test flatbed vehicle cooperative control process, and obtaining the driving direction and driving speed of each test flatbed vehicle at any time. The expression is as follows: Where n∈[1,N], N represents the number of test flatbed trucks in the test scenario, represents the dynamic control variable of the nth test flatbed truck at time t, represents the driving direction of the nth test flatbed truck at time t, represents the speed of the nth test flatbed truck at time t; the cooperative control target expression in the cooperative control process of the test flatbed truck is as follows: ( L n =(x n ,y n )) in, represents the position of the nth test flatbed truck at time t, Represents the location The horizontal position coordinates and vertical position coordinates in L n represents the preset target position of the nth test flatbed truck, x n ,y n Represents the position L n The horizontal position coordinates and vertical position coordinates in ; Combined with the requirements of collision avoidance and dynamic obstacle avoidance, the constraints in the test flatbed vehicle cooperative control process are: in, represents the distance between the nth test flatbed truck and the nearest obstacle at time t. If If it is less than or equal to 0, it means that the nth test flatbed truck collides with the obstacle at time t; represents the distance between the nth test flatbed truck and the nearest test flatbed truck at time t. If If it is less than or equal to 0, it means that the nth test flatbed truck collides with the test flatbed truck at time t.
3. The method for cooperative control of a multi-target test flatbed vehicle according to claim 1, characterized in that: The collaborative control model is used to construct the state space and action space of the collaborative control process of the test flatbed vehicle. The action space includes the set of control actions of the test flatbed vehicle in the collaborative control process, and the set of control actions includes the driving direction control action and the driving speed control action of the test flatbed vehicle. The state space includes all the state information of the test flatbed vehicle under different driving environments. The state information includes the position, driving direction, driving speed and surrounding environment information of the test flatbed vehicle. The expression of the state information of the nth test flatbed vehicle at time t is as follows: in, represents the position of the nth test flatbed truck at time t; Represents the status information of the nth test flatbed truck at time t; Represents the surrounding environment information of the nth test flatbed truck at time t; represents the driving direction of the nth test flatbed truck at time t; represents the speed of the nth test flatbed truck at time t; the surrounding environment information is expressed in the form of an environment matrix, Represents the value of the environment matrix element at position i. The expression is as follows: in, , The center of the matrix is the position of the nth test flatbed truck at time t.
4. The method for cooperative control of a multi-target test flatbed vehicle according to claim 1, characterized in that: The expression of the multi-objective reward function is as follows: Where R(·) represents the multi-objective reward function of the test flatbed truck cooperative control process; A t represents the set of control actions taken by N test flatbed trucks at time t, represents the control action of the nth test flatbed truck at time t; Represents the status information of the nth test flatbed truck at time t.
5. The method for cooperative control of a multi-target test flatbed vehicle according to claim 1, characterized in that: The multi-objective reward function value includes: a distance reward value of a preset target position, a collision reward value, a target position reward value, a collaborative control reward value of multiple test flatbed vehicles, and an environmental complexity reward value. The expression of the objective reward function is as follows: The expression of the distance reward value of the preset target position is as follows: The expression for the collision reward value is as follows: The expression of the target position reward value is as follows: The expression of the reward value of the cooperative control of multiple test flatbed vehicles is as follows: The expression of the environment complexity reward value is as follows: in, represents the control action of the nth test flatbed truck at time t; represents the state information of the nth test flatbed truck at time t; ||·||2 represents the L2 norm; Represents the control action taken by the nth test flatbed truck at time t After that, at the time t+1; L n represents the preset target position of the nth test flatbed truck; Represents the control action taken by the nth test flatbed truck at time t After that, the distance to the nearest obstacle at time t+1; Represents the control action taken by the nth test flatbed truck at time t After that, the distance to the nearest test flatbed at time t+1; exp(·) represents an exponential function with a natural constant as the base; Represents the control action taken by the nth test flatbed truck at time t The subsequent environmental matrix.
6. The method for cooperative control of a multi-target test flatbed vehicle according to claim 1, characterized in that: The method of using the state space, action space and target reward function of the test flatbed vehicle collaborative control process to generate a test flatbed vehicle control strategy under different state information includes the following steps: The initial strategy function value is calculated by using the state space and action space of the test flatbed vehicle cooperative control process, and the initial strategy function value of the control action under each state information is obtained; Combine the initial policy function value with the target reward function, calculate the expected reward value of the control action under each state information, and use the expected reward value to determine the action advantage value under different state information; Input the expected reward value and the action advantage value into the strategy function, and output the strategy function value and the strategy ratio at the current number of iterations; The strategy function value is iterated using the action advantage value and strategy ratio to obtain the final strategy function value for selecting different control actions under different state information. The final strategy function value obtained by iteration is stored to form the optimal test flatbed vehicle control strategy under different state information.
7. A multi-target test flatbed vehicle cooperative control method according to claim 6, characterized in that: Using the action advantage value and the policy ratio to iterate the policy function value includes the following steps: Standardize strategy ratios; Using the standardized strategy ratio, action advantage value and discount factor, the strategy function value is iteratively updated to gradually optimize the strategy function value; The final strategy function value is obtained until the maximum number of iterations is reached or the difference between the strategy advantage values obtained in adjacent iterations is less than the preset iteration threshold.
8. The method for cooperative control of a multi-target test flatbed vehicle according to claim 1, characterized in that: The test flatbed vehicle is dispatched for cooperative control according to the test flatbed vehicle control strategy, including the following steps: Obtain the status information of N test flatbed trucks at time t; Select state information from the optimal test flatbed vehicle control strategy, and calculate the similarity between the selected state information and the state information of the test flatbed vehicle at time t; The control action with the largest final strategy function value associated with the state information with the highest similarity in the optimal test flatbed vehicle control strategy is used as the control action of the test flatbed vehicle at time t, and the cooperative control of N test flatbed vehicles is performed according to the control action; Repeat the above steps until all test flatbed vehicles reach the preset target position.
9. A multi-target test flatbed vehicle collaborative control system, using a multi-target test flatbed vehicle collaborative control method according to any one of claims 1 to 8, characterized in that: include: A model building module is used to formally model the cooperative control process of the test flatbed vehicle by using the driving direction and driving speed of the test flatbed vehicle in combination with collision avoidance and dynamic obstacle avoidance requirements to obtain a cooperative control model; A reward function construction module is used to construct a state space and an action space of a test flatbed vehicle cooperative control process using a cooperative control model, and determine a multi-objective reward function using the state space and the action space; A control strategy generation module is used to generate the optimal test flatbed vehicle control strategy under different state information by using the state space, action space and target reward function of the test flatbed vehicle collaborative control process; The collaborative control module is used to collaboratively control the test flatbed vehicle according to the optimal test flatbed vehicle control strategy combined with current state information.
10. A multi-target test flatbed cooperative control system according to claim 9, characterized in that: The multi-objective reward function value includes: a distance reward value of a preset target position, a collision reward value, a target position reward value, a collaborative control reward value of multiple test flatbed vehicles, and an environmental complexity reward value.
Citation Information
Patent Citations
Full-speed-domain vehicle adaptive cruise control method based on multi-target dynamic weighting
CN116985802A