Reinforced learning PID controller parameter adjustment method and system based on closed-loop self-learning

Through the closed-loop self-learning framework and continuous learning method, the problem of insufficient adaptability of traditional PID controllers in complex scenarios is solved, and a wider range of scenario adaptability and control effects in autonomous driving are achieved.

CN120255314APending Publication Date: 2025-07-04TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510186487.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The fixed parameters of traditional PID controllers are difficult to cope with dynamic and complex scenarios. The enhanced learning PID controller has limited feasibility of scenes in real open road environments, making it difficult to achieve good control results under extreme conditions.

Method used

The reinforcement learning PID controller parameter adjustment method based on closed-loop self-learning is adopted, and multiple iterative training is carried out by building a reinforcement learning model. The continuous learning method is combined with the generation of adversarial edge road scenarios in high-fidelity simulation software, and the model is optimized to adapt to complex and diverse practical road scenarios.

Benefits of technology

The scenario feasible domain of reinforcement learning PID controller in the field of autonomous driving control is improved, the model can be avoided from being disastrously forgotten in iterative training, and the control effect in complex scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255314A_ABST
    Figure CN120255314A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving control, in particular to a reinforcement learning PID controller parameter adjustment method based on closed-loop self-learning, and the method enables a reinforcement learning model used for adjusting PID parameters to adapt to more complex and diversified actual road scenes after multiple closed-loop iterations. The method is not limited to a scene range of a scene data set used by single training, is beneficial for reinforcement learning of landing application of a PID controller in the field of automatic driving control, and improves a scene feasible region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of autonomous driving control, and particularly to a method and system for adjusting parameters of a reinforcement learning PID controller based on closed-loop self-learning. Background Art

[0002] Autonomous driving technology involves multiple links such as decision-making and planning control. Control is a key link in converting the upper-layer decision-making intention into the actual driving behavior of the vehicle. With the continuous update and iteration of control theory, control methods such as Proportional-Integral-Derivative (PID), Linear Quadratic Regulator (LQR), and Model Predictive Control (MPC) have been widely applied. Among them, the PID controller still occupies an important position in many applications due to its simplicity, effectiveness, and model-free characteristics.

[0003] The parameters of traditional PID controllers are fixed, making it difficult to handle dynamic and complex scenarios and highly dependent on manual parameter tuning. Therefore, researchers have applied methods such as fuzzy control and Bayesian optimization to PID parameter tuning, which can adaptively adjust the controller parameters in real time according to the current state to optimize control. However, these methods are often complex to implement. Reinforcement learning methods have shown powerful performance in many fields due to their characteristics of dynamic interaction and continuous trial and error. Many studies have combined reinforcement learning with PID controllers to achieve adaptive parameter tuning and obtained good control effects. However, for autonomous driving tasks, the real and open driving environment is unknown, complex, variable, and infinite. When the reinforcement learning PID controller trained only on common road datasets is deployed to the real and open road environment, it is still difficult to achieve good control effects under various extreme conditions, and the scene feasible region is still limited.

[0004] Therefore, it is very necessary to continuously generate adversarial scenarios for the defects of the reinforcement learning PID controller, and then train the controller to overcome new adversarial scenarios without catastrophic forgetting, and repeatedly iterate and optimize the reinforcement learning model to gradually expand the scene feasible region and cover diverse scenarios. Summary of the Invention

[0005] The object of the present disclosure is to propose a method for tuning parameters of a reinforcement learning PID controller based on closed-loop self-learning. The existing reinforcement learning PID controller is closed-loop optimized through a closed-loop self-learning framework, so that the reinforcement learning model for adjusting PID parameters can adapt to more complex and diverse actual road scenarios after multiple closed-loop iterations, rather than being limited to the scenario range of the scenario dataset used in a single training, which helps the reinforcement learning PID controller to be applied in the field of autonomous driving control and improve its scenario feasible region. The specific technical solutions are as follows.

[0006] In a first aspect, the present disclosure proposes a method for tuning parameters of a reinforcement learning PID controller based on closed-loop self-learning. The method includes the following steps: constructing a PID controller for parameter tuning based on a reinforcement learning model, where the reinforcement learning outputs the PID parameters of the PID controller, including the proportional coefficient, integral coefficient, and differential coefficient; constructing first training data based on scenario simulation during normal driving to obtain a trained reinforcement learning model, which can give appropriate PID parameters in different scenarios; Step 100: constructing second training data based on edge road scenario simulation, inputting the second training data into the trained reinforcement learning model, determining the training data that cannot be successfully executed by autonomous driving, and using it as third training data; based on the third training data, continuously training the reinforcement learning model through a continuous learning method, using the first training data and the second training data to test the trained reinforcement learning model for autonomous driving, and obtaining the range of the scenario feasible region; if the scenario feasible region corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target, then return to Step 100.

[0007] In an implementation of the above technical solution, the reinforcement learning model takes the current control target error err and the vehicle state as inputs. The vehicle state consists of the current vehicle speed v, the current tire friction coefficient μ, and the current road curvature k. The reinforcement learning model takes the proportional adjustment K p of the PID controller, the integral adjustment K p and the differential adjustment K d as action outputs, sets a reward function based on the current control target error err, and the PID controller outputs a control amount based on the current control target error err and the action output of the reinforcement model.

[0008] In an implementation of the above technical solution, while obtaining the scenario feasible region, it also includes calculating test evaluation indicators to quantify the control performance of the reinforcement model. The test evaluation indicators include control error and task completion rate. The average control error is the average value of the error between the control target and the actual control result in a single test, and the task completion rate is the test success rate of all test scenarios.

[0009] In an implementation of the above technical solution, for the adversarial edge road scenario, the simulation steps include: initializing vehicles with different tire friction coefficients and equipped with a reinforcement learning PID controller, generating roads with different elements, where the roads differ at least in curvature, length, and slope; designing driving scenarios with different driving speeds, so that for the generated road scenarios, the current PID controller tuned based on the reinforcement learning model cannot control the vehicle to safely and accurately complete the driving task.

[0010] In an implementation of the above technical solution, the simulation is implemented using a high-fidelity simulation software.

[0011] In an implementation of the above technical solution, the continuous learning method includes a replay-based method, a regularization-based method, an architecture-based method, and a combined method of the above methods.

[0012] In a second aspect, the present disclosure proposes a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform any of the above methods.

[0013] In a third aspect, the present disclosure proposes a system for tuning the parameters of a reinforcement learning PID controller based on closed-loop self-learning. The system includes a construction module, a first training module, a second training module, a third training module, and a judgment module. Among them: the construction module is configured to construct a PID controller tuned based on a reinforcement learning model, and the reinforcement learning outputs the PID parameters of the PID controller, including a proportionality coefficient, an integral coefficient, and a differential coefficient; the first training module is configured to construct first training data based on the scenario simulation during normal driving, and obtain a trained reinforcement learning model that can give appropriate PID parameters in different scenarios; the second training module is configured to construct second training data based on the edge road scenario simulation, input the second training data into the trained reinforcement learning model, determine the training data that cannot be successfully executed by autonomous driving, and use it as the third training data; the third training module is configured to continue training the reinforcement learning model based on the third training data through a continuous learning method, and perform autonomous driving tests on the trained reinforcement learning model using the first training data and the second training data to obtain the range of the scenario feasible region; the judgment module is configured to execute the second training module when the range of the scenario feasible region corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target.

[0014] Advantageous technical effects of the present disclosure: (1) The closed-loop self-learning framework of the reinforcement learning PID controller constructed by the method of the present disclosure initializes vehicles with different tire friction coefficients through a high-fidelity simulation software, generates different element roads with parameters such as different curvatures, lengths, and slopes, and designs driving scenarios with different driving speeds, which can efficiently generate the training scenario data required for further optimization of the reinforcement learning model and get rid of the dependence on the existing edge scenario dataset. (2) The closed-loop self-learning framework of the reinforcement learning PID controller constructed by the method of the present disclosure can train the reinforcement learning model by combining the continuous learning method in multiple iterative trainings, avoiding the interference of policies and the phenomenon of catastrophic forgetting caused by training the reinforcement learning model in different types and different difficulty driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 、 One Schematic diagram of the structure of the reinforcement learning PID controller in a certain embodiment.

[0017] Figure 2 、 One Schematic diagram of the closed-loop self-learning framework for iterative optimization of the reinforcement learning PID controller in a certain embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following will clearly and completely describe how to implement the technical solutions of this case in conjunction with the drawings. Obviously, the described embodiments are only a part of the embodiments of this case, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this case without creative efforts belong to the scope of protection of the present application.

[0019] A method for tuning parameters of a reinforcement learning PID controller based on closed-loop self-learning, which can be applied to the field of autonomous driving control, includes the following steps.

[0020] S1: Construct an initial PID controller for tuning parameters based on a reinforcement learning model.

[0021] The PID controller for tuning parameters based on a reinforcement learning model is divided into a reinforcement learning part for adjusting PID parameters and a PID part for outputting control quantities, and is applicable to autonomous driving longitudinal and lateral control.

[0022] SeeFigure 1 , the reinforcement learning model takes the current control target error err and the vehicle state as inputs. The vehicle state consists of the current vehicle speed v, the current tire friction coefficient μ, and the current road curvature k. The reinforcement learning model takes the proportional regulation K of the PID controller p , the integral regulation K p , and the derivative regulation K d as action outputs, sets the reward function based on the current control target error err, and the PID controller outputs a control amount based on the current control target error err and the action output of the reinforcement model.

[0023] The reinforcement learning method used by the reinforcement learning model takes the SAC (Soft Actor Critic) algorithm as an example. The SAC algorithm is a maximum entropy-based deep reinforcement learning algorithm designed for Markov decision processes in continuous action spaces. Based on the Actor-Critic framework, SAC performs well in many tasks, and its objective function is designed as:

[0024]

[0025] where ρ π represents the distribution that the action pair (s t , a t ) under the policy π follows. a is the weight of entropy, called the temperature coefficient. H(π(·|s t )) represents the entropy of the policy π in the state s t . Entropy represents the randomness or uncertainty of the policy. The larger the entropy, the stronger the exploratory ability of the policy. In this framework, the policy needs to balance the relationship between reward and entropy, maximize the expected reward and entropy, and thus maximize the expected return. Introducing entropy into the objective function helps to promote the random exploration of the policy, prevent the policy from converging to the local optimum prematurely, and make the model more adaptable and robust.

[0026] The SAC algorithm contains an Actor and two Critics. The Actor network updates the policy by minimizing the following expectation:

[0027]

[0028] where, are the policy network parameters, Q θ (s t , a t ) is the action value function estimated by the Critic network, log(π φ (a t |s t )) is the log probability of the policy. The Critic network updates using the square of the temporal difference error as the loss function:

[0029]

[0030] where θ is the value network parameter, is the target value network parameter, is the target value function, and (1 - d t+1 ) indicates whether the termination state is reached.

[0031] The input state of SAC is s t = [err, v, μ, κ], including the current control target error err, the current vehicle speed v, the current tire friction coefficient μ, and the current road curvature K. The output action a of the reinforcement learning model t = [K p , K i , K d , that is, the proportional, integral, and differential coefficients of PID. The reward function of reinforcement learning is related to the control target error err. The smaller err is, the better the parameter adjustment is, and the larger the reward value is.

[0032] The input of the PID controller is the current control target error err, and the output is the control quantity. For the longitudinal control task, the control target error err represents the difference between the current longitudinal speed and the target value. For the lateral control task, err represents the deviation angle between the current driving direction of the vehicle and the target anchor point. The control law of the PID controller is:

[0033]

[0034] S2: Construct the first training data based on the scenario simulation in the normal driving process, and obtain a trained reinforcement learning model, which can give appropriate PID parameters in different scenarios.

[0035] Common road scenarios represent common driving scenarios in the normal driving process, such as straight road and curved road scenarios at normal speed. This type of scenario can be directly obtained from the existing real road driving scenario dataset or simulated and generated in various high-fidelity simulation platforms. The scenarios are relatively simple, and the PID controller tuned by reinforcement learning can handle them well.

[0036] For the above simulation, the Carla or Metadrive simulation platform can be used to quickly simulate and generate a large number of common road scenarios.

[0037] S3: Construct the second training data based on the edge road scenario simulation, input the second training data into the trained reinforcement learning model, determine the training data that cannot be successfully executed by autonomous driving, and use it as the third training data.

[0038] Edge road scenarios represent driving scenarios that are rare in normal driving behaviors but require high control effectiveness from the controller, such as high-speed rain / snow road scenarios prone to skidding, continuous high-curvature bend scenarios, etc. Due to the long-tail problem of autonomous driving, the probability of occurrence of such road scenarios is not high but they are very important for driving safety. The reinforcement learning PID controller trained only in common road scenarios may not be able to set good control coefficients and achieve good control effects in such edge scenarios, so further training is required in this type of scenario. It is difficult to efficiently and accurately screen such scenarios directly from existing driving datasets, and they can be generated specifically in high-fidelity simulation software.

[0039] A method for generating adversarial edge road scenarios is as follows: Initialize vehicles with different tire friction coefficients in the high-fidelity simulation platform Metadrive, equipped with a reinforcement learning PID controller, to generate different-element roads with different parameters such as curvature, length, slope, etc., and design driving scenarios with different driving speeds, making the generated road scenarios as difficult as possible so that the current reinforcement learning PID controller cannot control the vehicle to complete the driving task safely and accurately, thereby uncovering the defects of the current reinforcement learning PID controller. After generating a sufficient number of adversarial edge road scenarios, collect them into the edge road scenario library for the reinforcement learning PID controller to perform the next stage of learning.

[0040] S4. Based on the third training data, continue to train the reinforcement learning model through a continuous learning method, and use the first training data and the second training data to conduct autonomous driving tests on the trained reinforcement learning model to obtain the range of the scenario feasible region.

[0041] In this embodiment, when the reinforcement learning PID controller is trained in the specifically generated adversarial edge road scenarios, to avoid model style drift, it learns the coping methods for edge road scenarios but forgets the parameter adjustment methods in common road scenarios, that is, catastrophic forgetting occurs. A continuous learning method is used to train in the edge road scenarios.

[0042] The described continuous learning method includes but is not limited to methods based on replay, methods based on regularization, methods based on architecture, and combined methods of the above methods.

[0043] In this embodiment, a method for slowing down catastrophic forgetting is introduced by taking a classic regularized continuous learning method, Elastic Weight Consolidation (EWC), as an example. The core idea of ​​EWC is to estimate the importance of network weights for previous tasks before switching tasks, and to impose regularization constraints on them when updating weights for subsequent tasks, so that important weights can be consolidated, thereby enabling the network model to take into account different tasks without increasing the network structure. In order to weigh the importance of different weights, the EWC algorithm estimates the importance of each parameter through the diagonal elements of the Fisher Information Matrix (FIM), thereby constructing a regularization term L ewc For task T i Sample x i The experience pool composed of ~p(x|θ), the FIM calculation formula is:

[0044]

[0045] Task T i For a certain parameter θ of the current network j The importance of F i The j-th diagonal element F i,j estimate:

[0046]

[0047] So when switching to task T n When θ j For any task T i (i<n) The regularization constraint term brought about in the subsequent training is:

[0048]

[0049] Among them, λ is used to measure the importance of the old task compared to the new task. Indicates that after task T i The optimal parameters obtained by training. Unlike L1 regularization or L2 regularization, which treats all parameters equally, L ewc It can be constrained selectively to avoid changes in important weights as much as possible. Then, task T n The loss function L(θ) during training is constructed as:

[0050]

[0051] Where L(θ) represents the loss when only considering the current task, It represents the sum of the regularization constraints of the previous n - 1 old tasks for all parameters. The jointly constituted loss function L(θ) enables the network parameters to change not only in the direction beneficial to the current task during update, but also consolidates the performance on the old tasks.

[0052] During the automatic driving test in this step, the calculation of test evaluation metrics can also be carried out simultaneously. The test evaluation metrics include control error and task completion rate. The average control error is the average value of the errors between the control target and the actual control result in a single test, and the task completion rate is the success rate of all test scenarios. The control performance of the reinforcement learning model is evaluated through the calculation of test evaluation metrics to improve the safety of automatic driving.

[0053] S5: When the feasible region range corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target, return to step S3.

[0054] Test the reinforcement learning PID controller in common road scenarios and edge road scenarios. If the feasible region of the scenario meets the requirements, end the closed-loop; otherwise, repeat the closed-loop process of scenario generation and model training above to continuously optimize the reinforcement learning model.

[0055] In this example, the test evaluation metrics include average control error, task completion rate, and scenario feasible region, etc. The average control error is the average value of the errors between the control target and the actual control result in a single test; the task completion rate is the success rate of all test scenarios; the scenario feasible region can be defined by analyzing and clustering the driving scenario features, so as to analyze the distribution of driving scenarios that the reinforcement learning PID controller can handle and observe the expansion of the scenario feasible region.

[0056] In this example, the closed-loop self-learning process refers to a closed-loop process for a reinforcement learning PID controller that can achieve good control effects in common road scenarios, repeatedly generating adversarial edge road scenarios, training in the more difficult generated adversarial edge road scenarios based on the method of continuous learning, and testing evaluation metrics such as the average control error, task completion rate, and scenario feasible region of the controller. In this process, adversarial scenarios are continuously generated for the defects of the reinforcement learning PID controller, and then the controller is continuously trained to overcome the new adversarial scenarios without catastrophic forgetting, repeatedly iteratively optimizing the reinforcement learning model until the control effect of the reinforcement learning PID controller in the test evaluation process meets the requirements, obtaining a reinforcement learning PID controller that can perform well in more complex and diverse road scenarios, and exiting the closed-loop self-learning process.

[0057] Figure 2It illustrates the iterative optimization process of implementing the EWV SAC closed-loop self-learning framework in the above steps. During the training of the reinforcement learning model, the losses include the EWC loss and the network's own loss, and gradient descent is used for the continuous evolution of the model. For the trained reinforcement learning model, model testing and evaluation are performed to determine whether the feasible region meets the target requirements. The data used for training is obtained through simulation software, or the normal driving scenarios come from actual autonomous driving. For the training data that cannot be used for the successful completion of tasks by autonomous driving, this data is continued to be used for model training until the control effect of the reinforcement learning PID controller in the testing and evaluation process meets the requirements, obtaining a reinforcement learning PID controller that can perform well in more complex and diverse road scenarios, and then exiting the closed-loop self-learning process.

[0058] Through the description of the above embodiments, those skilled in the art can clearly understand that a corresponding system can be implemented according to the method of the present disclosure. Exemplarily, a system for tuning the parameters of a reinforcement learning PID controller based on closed-loop self-learning, the system includes a construction module, a first training module, a second training module, a third training module, and a judgment module; wherein: the construction module is configured to construct a PID controller for parameter tuning based on a reinforcement learning model, and the reinforcement learning outputs the PID parameters of the PID controller, including the proportional coefficient, integral coefficient, and differential coefficient; the first training module is configured to construct first training data based on the scenario simulation during normal driving to obtain a trained reinforcement learning model, which can give appropriate PID parameters in different scenarios; the second training module is configured to construct second training data based on the edge road scenario simulation, input the second training data into the trained reinforcement learning model, determine the training data that cannot be successfully executed by autonomous driving, and use it as the third training data; the third training module is configured to continue training the reinforcement learning model based on the third training data through a continuous learning method, perform autonomous driving tests on the trained reinforcement learning model using the first training data and the second training data, and obtain the range of the scenario feasible region; the judgment module is configured to execute the second training module when the range of the scenario feasible region corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target.

[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods and systems of the present disclosure can be implemented by means of software plus necessary general hardware. Of course, they can also be implemented by dedicated hardware including application specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally speaking, functions accomplished by computer programs can easily be implemented by corresponding hardware, and there can be various specific hardware structures for implementing the same function, such as analog circuits, digital circuits or dedicated circuits. However, in more cases for the present disclosure, implementation by software programs is a better embodiment.

[0060] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and all of these fall within the scope of protection of the present disclosure.

Claims

1. A method for tuning parameters of a reinforcement learning PID controller based on closed-loop self-learning, characterized in that, The method includes the following steps: Construct a PID controller whose parameters are adjusted based on a reinforcement learning model. The reinforcement learning outputs the PID parameters of the PID controller, including the proportional coefficient, integral coefficient, and differential coefficient; Construct first training data based on scenario simulation during normal driving to obtain a trained reinforcement learning model that can give appropriate PID parameters under different scenarios; Step 100: Construct second training data based on edge road scenario simulation, input the second training data into the trained reinforcement learning model, determine the training data for which autonomous driving cannot be successfully executed, and use it as third training data; Based on the third training data, continue to train the reinforcement learning model through a continuous learning method. Use the first training data and the second training data to test the trained reinforcement learning model for autonomous driving and obtain the range of the scenario feasible region; When the range of the scenario feasible region corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target, return to step 100.

2. The method according to claim 1, characterized in that The reinforcement learning model takes the current control target error err and the vehicle state as inputs. The vehicle state consists of the current vehicle speed v, the current tire friction coefficient μ, and the current road curvature k. The reinforcement learning model takes the proportional adjustment K p of the PID controller, the integral adjustment K p and the derivative adjustment K d as action outputs, sets the reward function based on the current control target error err, and the PID controller outputs a control quantity based on the current control target error err and the action output of the reinforcement model.

3. The method according to claim 1, characterized in that, When obtaining the scenario feasible region, it also includes calculating test evaluation metrics to quantify the control performance of the reinforcement model. The test evaluation metrics include control error and task completion rate. The average control error is the average value of the error between the control target and the actual control result for a single test, and the task completion rate is the success rate of all test scenarios.

4. The method according to claim 1, wherein The simulation steps of the adversarial edge road scenario include: Initialize vehicles with different tire friction coefficients equipped with a reinforcement learning PID controller, and generate roads with different elements. The roads differ at least in curvature, length, and slope; Design driving scenarios with different driving speeds so that the generated road scenarios cannot be controlled by the current PID controller whose parameters are adjusted based on the reinforcement learning model to safely and accurately complete the driving task.

5. The method according to claim 1, characterized in that, The simulation is implemented using a high-fidelity simulation software.

6. The method according to claim 1, characterized in that The continuous learning method includes a replay-based method, a regularization-based method, an architecture-based method, and a combined method of the above methods.

7. A computer-readable storage medium, characterized in that: There is a computer program stored that can be loaded and executed by a processor to perform any one of the methods in claims 1 to 6.

8. A parameter tuning system for a reinforcement learning PID controller based on closed-loop self-learning, characterized in that, The system includes a construction module, a first training module, a second training module, a third training module, and a judgment module; wherein: The construction module is configured to construct a PID controller whose parameters are adjusted based on a reinforcement learning model. The reinforcement learning outputs the PID parameters of the PID controller, including the proportional coefficient, integral coefficient, and differential coefficient; The first training module is configured to construct first training data based on scenario simulation during normal driving to obtain a trained reinforcement learning model that can give appropriate PID parameters under different scenarios; The second training module is configured to construct second training data based on edge road scenario simulation, input the second training data into the trained reinforcement learning model, determine the training data for which autonomous driving cannot be successfully executed, and use it as third training data; The third training module is configured to continue training the reinforcement learning model based on the third training data through a continuous learning method, and perform an autonomous driving test on the trained reinforcement learning model using the first training data and the second training data to obtain the range of the scene feasible region; The judgment module is configured to execute the second training module when the range of the scene feasible region corresponding to the PID parameters output by the trained reinforcement learning model does not reach the target.