Large language model heuristic Monte Carlo tree search controller design method
Through the controller design method of large language model heuristic Monte Carlo tree search, an enhanced knowledge base is built and fine-tuned. Combined with the Monte Carlo tree search optimization control operator combination, the problems of low design efficiency and unstable effect of traditional controllers are solved, and efficient and stable controller generation is achieved.
Patent Information
- Application Number
- CN202510380864.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional controller design methods are inefficient and have high computing resource requirements when dealing with nonlinear or time-varying systems, and the control effect of the model-free method is unstable and difficult to optimize.
The controller design method of Monte Carlo tree search is adopted to fine-tune the large language model by building an enhanced knowledge base and establishing a private data set in the control field to generate a preliminary controller design scheme, and use Monte Carlo tree search optimization control operator combinations to combine it with the simulation environment for global optimization.
Significantly reduces controller design debugging time, improves controller response speed, stability and robustness, ensuring the generated controller has optimal performance in complex environments.
Smart Images

Figure CN120335295A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of controllers, and in particular to a method for designing a controller for heuristic Monte Carlo tree search of large language models. Background Art
[0002] Controllers play an indispensable role in automated systems. They are not only responsible for monitoring the real-time state of the system, but also constantly adjusting and optimizing the behavior of the system to ensure its operation under safe and efficient conditions. Controllers usually receive data from sensors, analyze and judge the current state of the system, and then generate appropriate control signals to convey instructions to actuators or other related devices. In this way, the controller can ensure that the entire system operates within the set performance index range and effectively cope with changes in the external environment and disturbances within the system.
[0003] Traditional controller design is usually divided into model-based methods and model-free methods. Model-based methods rely on an accurate mathematical model of the controlled object, and the controller adjusts by predicting the system response through the model. For example, patent (CN108107713A, A proportional derivative lead intelligent model set PID controller design method, 2021.06.29) and literature (Zhou Xuesong, Wang Xin, Ma Youjie, etc. Design of a genetic active disturbance rejection controller for DC converters based on D-separation method [J]. Acta Energiae Solaris Sinica, 2024, 45(09): 378-385.). However, the disadvantage of this method is that the model construction process is complex, requiring a large amount of data and accurate experiments to obtain parameters. Once the model does not match the actual system, the control effect may be greatly reduced, especially when dealing with nonlinear or time-varying systems. In addition, model-based methods have high requirements for computing resources and may not be able to adapt to resource-constrained application scenarios.
[0004] In contrast, model-free methods do not rely on an accurate mathematical model, but directly design control based on input-output data. For example, patent (CN 112034715A, A model-free feedback controller design method for a motor servo system based on an improved Q-learning algorithm, 2021.07.13) and literature (Huang Tao, Ban Xiaojun, Wu Fen, etc. Design of a self-learning control system for maglev based on Q-network [J]. Electric Machines and Control, 2021, 25(09): 132-139.). Such methods have the advantage of strong adaptability and can achieve good control effects in unknown or complex environments. However, the main disadvantage of model-free methods is that the control effect is often unstable and difficult to predict, especially when the system operating conditions change greatly. In addition, due to the lack of a detailed understanding of the internal behavior of the system, fault diagnosis and optimization are difficult, and the design and debugging process of the controller is often time-consuming. Summary of the Invention
[0005] The purpose of the present invention is to provide a controller design method based on large language model heuristic Monte Carlo tree search in order to reduce the design and debugging time of the controller and improve the control effect of the controller.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A controller design method based on large language model heuristic Monte Carlo tree search, the method comprising the following steps:
[0008] S1. Construct an enhanced knowledge base;
[0009] S2. Establish a private dataset in the control field to fine-tune the large language model;
[0010] S3. Based on the requirements input by the user, the fine-tuned large language model extracts corresponding control knowledge from the knowledge base and generates a preliminary controller design scheme;
[0011] S4. Establish an operator library and construct a simulation environment;
[0012] S5. Select control operators from the operator library, and use the Monte Carlo tree search method to optimize the preliminary controller design scheme, and search for the optimal path as the optimal controller;
[0013] S6. Test in the simulation environment. If the simulation result fails to meet the expected performance, optimize the optimal controller again. Otherwise, execute S7;
[0014] S7. Output the mathematical structure and simulation results of the controller.
[0015] Furthermore, each node of the Monte Carlo tree represents a control operator and its parameter settings, and the path represents a complete operator combination and its corresponding parameter configuration scheme.
[0016] Furthermore, the specific steps of executing the Monte Carlo tree search method are as follows:
[0017] First, select a control operator, combine the operator coefficient and the system variable to generate a root node, the operator coefficient and the system variable are the parameter settings of the operator, then select the next node of the current node as the new current node through scoring. When the current node searched is a leaf node, generate new child nodes according to the operator of the current node, generate new paths for the new child nodes, evaluate the performance of the new paths, assign the upper confidence bound score obtained by the evaluation to each node on the new path, and update the access times of each node at the same time, and finally obtain the optimal path.
[0018] Furthermore, the optimal path satisfies:
[0019] When the upper confidence bound score J of the new path is less than the threshold or the upper limit of the search times is reached, the path with the smallest upper confidence bound score J is the optimal path, and the controller corresponding to the optimal path is the optimal controller, where the upper confidence bound score is:
[0020] J = α·OS + β·T s + γ·|e ss |
[0021] OS represents the overshoot, which describes the maximum deviation amplitude of the system relative to the target value during the response process, and T s represents the settling time, that is, the time required for the system output to first enter and continuously remain within a certain error band of the target value, and |e ss | represents the steady-state error, the deviation of the system from the target value after reaching the steady state, and α, β, and γ represent the weighting coefficients.
[0022] Further, the specific steps for generating new child nodes according to the operator of the current node are as follows:
[0023] The large language model uses the control domain knowledge obtained during the fine-tuning process and the experience of controller design to select multiple candidate operators and their parameter settings from the pre-constructed operator library as new child nodes.
[0024] Further, the control operators include proportional operators, integral operators, and derivative operators.
[0025] Further, the requirements input by the user include control objectives, system types, and performance indicators.
[0026] Further, the specific steps for fine-tuning the large language model by establishing a private dataset in the control domain are as follows:
[0027] Establish a private dataset in the control domain, where the dataset includes labeled control system design data, simulation results, performance evaluations, and case summaries, and the private dataset in the control domain is used to train and iterate the large language model multiple times to complete the fine-tuning.
[0028] Further, the simulation environment includes various factors in the actual working environment.
[0029] Further, re-optimizing the optimal controller specifically means re-executing S2 to S5.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] After receiving the input information, the system designs a preliminary controller scheme through the agent, automatically establishes a control operator library and a simulation environment, and combines the Monte Carlo tree search technology to globally optimize the controller. It can effectively explore combinations of various control algorithms, find the controller with the optimal performance, and improve the control effect of the controller. The further optimization of the Monte Carlo tree search provides a high-quality search direction. When the search reaches a leaf node, the large language model generates new child nodes according to the operators of the current node. This step forms multiple alternative control strategies by combining system variables, control operators, and operator coefficients. Each newly generated node represents a different combination scheme, thus greatly expanding the search space. The system evaluates the performance of the newly generated paths. By calculating the control performance in the scoring (such as response speed, steady-state error, overshoot, etc.), the actual performance of each candidate scheme can be quantified, providing a basis for subsequent selection. This method can not only improve the response speed and stability of the controller, but also ensure its robustness and anti-interference ability, thereby improving the overall performance of the control system, ensuring that the generated controller has the optimal performance, and at the same time, reducing the design and debugging time of the controller. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is the flowchart of the present invention;
[0033] Figure 2 is the Monte Carlo search path diagram;
[0034] Figure 3 is the experimental scenario diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0036] The present invention can be applied to an intelligent controller design platform, aiming to automatically generate and optimize controller design by combining a large language model (LLM) with fine-tuning and enhanced knowledge bases. Through the collaborative work of multiple modules, the system integrates steps such as knowledge base construction, fine-tuning data set, agent generation, controller simulation optimization, and verification, comprehensively improving the efficiency, intelligence, and applicability of controller design.
[0037] The present invention proposes a method for designing a controller based on large language model heuristic Monte Carlo tree search, and the flowchart is as Figure 1 shown. The method includes the following steps:
[0038] S1. Construct an enhanced knowledge base;
[0039] S2. Establish a private dataset in the control domain to fine-tune the large language model;
[0040] S3. Based on the requirements input by the user, the fine-tuned large language model extracts corresponding control knowledge from the knowledge base and generates a preliminary controller design scheme;
[0041] S4. Establish an operator library and construct a simulation environment;
[0042] S5. Select control operators from the operator library and use the Monte Carlo tree search method to search for the optimal path as the optimal controller;
[0043] S6. Test in the simulation environment. If the simulation results do not meet the expected performance, optimize the optimal controller again. Otherwise, execute S7;
[0044] S7. Output the mathematical structure and simulation results of the controller.
[0045] The specific process is as follows:
[0046] Step 1: Construct an enhanced knowledge base
[0047] The system first constructs an enhanced knowledge base to provide rich control domain knowledge support for controller design. The knowledge base covers various aspects such as academic literature, expert experience, historical design data, and simulation data, forming a huge knowledge repository.
[0048] This knowledge base is dynamically updated through continuous search and reinforcement learning to ensure its comprehensiveness, accuracy, and synchronization with the latest technologies and trends in the current control domain. The construction and dynamic enhancement process of the knowledge base adopt big data-based search technologies and reinforcement learning methods, enabling the knowledge base to not only contain static basic information but also reflect the development forefront of current control technologies, providing the most powerful support for the intelligent design of controllers.
[0049] Step 2: Establish a private dataset in the control domain to fine-tune the large language model
[0050] To improve the agent's ability in the controller design task, the system fine-tunes the large language model (LLM) through specific datasets. These datasets include the private control system design data, simulation results, performance evaluations, and case summaries of the present invention. After data annotation, segmentation, and enhancement, a private fine-tuning dataset is established.
[0051] Through multiple training iterations of the large language model using the dataset, the system injects rich professional knowledge in the control field into the model, enabling it to understand and adapt to the complex requirements of different control scenarios and generate efficient controller design solutions that meet the requirements. These fine-tuning datasets combine actual control experience and failure cases, enabling the model to avoid common problems during the design process and improve the accuracy and performance of the design. After fine-tuning, the model not only has general reasoning capabilities but also provides more targeted and adaptable solutions for specific control scenarios.
[0052] Step 3: User input of requirements and the intelligent agent generates a preliminary design
[0053] During the actual design process, the user first needs to input the control requirements for a specific scenario, including control objectives (such as maintaining stability, rapid response, etc.), system types (such as robotic arms, motors, etc.), and performance metrics (such as response time, energy consumption, robustness, etc.). This information provides a clear design direction for the intelligent agent.
[0054] Based on the requirements input by the user, the fine-tuned large language model extracts the corresponding control knowledge from the knowledge base and automatically generates a preliminary controller design solution that meets the user's requirements. This solution includes the mathematical model of the controller and its performance in preliminary simulations. In this way, the user can obtain a more targeted preliminary design without having to manually consult a large number of documents or manually select control algorithms, thus greatly saving design time.
[0055] Step 4: Establish an operator library and construct a simulation environment
[0056] An operator library is established in the system to support the design work of the controller. The operator library contains various control operators, system variables, and parameters required for designing the controller. These operators cover a variety of control methods from classical PID control to more complex fuzzy control, optimal control, etc.
[0057] The establishment of the operator library provides different algorithm modules for the intelligent agent as shown in Table 1, enabling it to flexibly combine and configure the most suitable algorithms according to the specific scenario requirements when generating the controller. In addition, the system constructs a simulation environment based on the user's prompts and the information in the operator library to evaluate the performance of the controller. The simulation environment simulates the characteristics of the actual physical system as realistically as possible, for example, considering nonlinear factors, noise interference, etc., to ensure the stability and effectiveness of the controller in actual applications.
[0058] Table 1 Operator library
[0059]
[0060] Step 5: Monte Carlo Tree Search (MCTS) to optimize the controller design
[0061] After designing the generation controller, to ensure its optimal performance, the system uses the Monte Carlo Tree Search (MCTS) algorithm to globally optimize the combination of control operators. Monte Carlo Tree Search is an effective global optimization method that can explore and compare multiple combinations in the operator library to find the optimal solution most suitable for a specific scenario, such as Figure 2 shown.
[0062] Each node of the tree represents a choice of a control operator or the parameter setting of the operator, while the path represents a complete operator combination and its corresponding parameter configuration scheme. Next, the search process of MCTS is divided into several steps:
[0063] Selection phase: Starting from the root node of the tree, MCTS selects the next node along the path with the highest current priority. The priority is calculated by the Upper Confidence Bound (UCB) formula, which is designed to balance the strategies of exploring new paths and exploiting the existing best paths. The design of the UCB formula can prevent the search from falling into local optima and ensure sufficient exploration space during the search process.
[0064] Expansion phase: When reaching a leaf node, the system generates several new child nodes according to the operator of the current node. When generating several new child nodes, the large language model utilizes the control domain knowledge obtained during the fine-tuning process and the experience of controller design to select multiple candidate operator combinations from the pre-constructed operator library. Specifically, the operator library contains system variables, control operators, and their corresponding operator coefficients, which can be flexibly combined to form various possible control strategies, and each combination represents a candidate node. After generating the candidate nodes, the system uses the Monte Carlo Tree Search (MCTS) algorithm to score and rank each candidate node through the upper confidence bound formula. Finally, the candidate node with the highest priority is selected and used as the basis for further expanding the search in the next step. It represents a new operator combination form or parameter adjustment. By expanding new nodes, the search tree is further developed.
[0065] Simulation phase: On the newly expanded node, the system generates a complete operator combination and conducts a simulation test to evaluate the performance of this combination. The specific performance evaluation is carried out by calculating the objective function value, such as the minimum steady-state error or the shortest response time, etc.
[0066] Backpropagation phase: After the simulation is completed, the simulation results are fed back to all nodes on the path, updating the score and visit count of each node, thereby optimizing the subsequent search direction. The visit count is used to record the operators used more frequently in multiple designs, and thus optimize the subsequent design process.
[0067] The method for determining the optimal path is as follows:
[0068] According to the upper confidence bound formula J, the smaller the formula J is, the better the performance of the controller. When the performance of the controller meets the specified conditions or reaches the upper limit of the search times, the controller with the smallest J is the optimal controller.
[0069] J = α·OS + β·T s + γ·|e ss |
[0070] OS: Overshoot, which describes the maximum deviation amplitude of the system relative to the target value during the response process. T s : Settling time, the time required for the system output to first enter and continuously remain within a certain error band (e.g., ±5%) of the target value. |e ss |: Steady-state error, the deviation of the system from the target value after reaching the steady state. α, β, and γ: Weighting coefficients (weight parameters), which can be flexibly adjusted according to specific application requirements to balance the importance of the three performance indicators.
[0071] Take Figure 2 as an example, e(t) is the functional relationship between the error and time, and K n is a constant coefficient.
[0072] At the top layer of the tree, an error integral operator is first selected. The root node represents the first operation in the controller design, the integral K1∫e(t)dt. This choice is because during the design, the integral operator has a significant impact on eliminating the steady-state error of the system and is a fundamental part of the controller design.
[0073] Next, the controller selects the error proportional operator K2e(t). This step is to select an appropriate operator according to the current priority in the controller design. Although there are multiple optional operators, the proportional operator with a higher score is selected through calculation. The proportional operator is usually used to respond to the current error, thereby adjusting the system in real time. By optimizing the proportional term, the response ability of the system to immediate changes can be improved.
[0074] In the process of selecting the next operator, the controller selects the error differential with a higher score The differential operator is usually used to predict the future behavior of the system, reduce the overshoot of the system, and ensure a faster response time and better stability. By optimizing the differentiation of the error, the system can respond to changes more quickly and reduce oscillations.
[0075] In further design, the system selects a proportional operator K4e(t) again. This repeated selection of the proportional term is to strengthen the system's ability to adjust the error and further optimize the performance of the controller.
[0076] Finally, the obtained controller is:
[0077]
[0078] Through such a selection process, the MCTS search tree gradually constructs a complete controller design scheme from the root node to the leaf node. Each step of the selection is based on the priority calculated by the Upper Confidence Bound (UCB) formula, ensuring that the search process can both explore new possible paths and utilize the existing best paths, and finally achieve a globally optimal control scheme under different control requirements.
[0079] This optimization process evaluates multiple performance metrics of the controller, including response speed, system stability, robustness, and anti-interference ability, etc., ensuring that the finally output controller can meet the requirements set by the user in all aspects. This optimization process can find the globally optimal solution in the complex controller design space, thus providing users with high-quality controller design schemes.
[0080] Step 6: Simulation Verification and Adjustment Optimization
[0081] Once the optimal controller of the preliminary design is found, the system will conduct detailed simulation verification on it to ensure that it meets the various performance metrics set by the user. During the simulation process, the system will simulate various factors in the actual working environment (such as noise interference, external load changes, etc.) to verify the performance of the controller in a complex environment.
[0082] If the simulation results fail to meet the expected performance, the system will automatically adjust the controller design for further optimization. This includes further fine-tuning of the large language model or adjusting the combination of operators to optimize the overall performance of the controller. Through this repeated cycle of verification and adjustment, the system can continuously improve the controller design and finally generate a controller that meets all performance requirements.
[0083] Step 7: Output Controller Design and Simulation Results
[0084] When the controller passes the simulation verification and meets all performance standards, the system will output the mathematical structure of the controller and the simulation results. The output content not only includes the mathematical model of the controller but also the performance evaluation data in various simulation environments. These data provide users with comprehensive reference information to help users better understand the performance of the controller in different scenarios.
[0085] These output results can be used for actual control system deployment, for example Figure 3 , to provide a control scheme for the system or as a basic reference for subsequent optimization. This link ensures that users not only obtain a controller that meets the design requirements but also have a clear understanding of the performance of the controller in actual applications.
[0086] Through the above series of steps, the system realizes the automation, intelligence, and systematization of controller design from knowledge base construction, agent fine-tuning, controller generation, simulation optimization to final verification. The whole process not only greatly improves the design efficiency but also significantly enhances the adaptability and practical application effect of the controller, and is widely applicable to various fields such as industrial automation, robot control, and autonomous driving.
[0087] The structure of the present invention is as follows:
[0088] The core feature of the system is to generate a controller based on an enhanced knowledge base and a fine-tuned large language model. First, the knowledge base contains rich data in the field of control, including literature, expert experience, and historical simulation data, etc., providing a solid knowledge foundation. Then, the system fine-tunes the large language model using a customized dataset, enabling it to design a controller according to the specific scenario requirements of the user, improve control performance, and generate an efficient controller mathematical model and simulation results.
[0089] During the controller generation process, the system supports the user to input control requirements, including specific control scenarios, control objectives, and performance indicators, etc. After receiving the input information, the system designs a preliminary controller solution through an agent, automatically establishes a control operator library and a simulation environment, and combines the Monte Carlo tree search technology to globally optimize the controller to ensure that the generated controller has the best performance. Performance evaluation includes multiple dimensions such as response time, stability, and robustness, and the optimization process ensures that the controller can meet all the needs of the user.
[0090] Finally, the system conducts simulation verification on the generated controller to ensure that it meets the set performance standards. If the controller does not meet the requirements, the system will further adjust and optimize it, and through continuous iteration and improvement until the optimal solution is found. The verified controller and its simulation data will be used as the final output for the user to carry out practical applications or further optimization.
[0091] Through the close combination of the knowledge base, agent, and optimization algorithm, this system realizes the intelligence, automation, and high efficiency of controller design, can provide customized control solutions for users, and is widely applied in fields such as industrial automation and robot control, helping to improve the overall performance and reliability of the control system.
[0092] Compared with the prior art, the controller automatic design system of the present invention has the following remarkable advantages and beneficial effects:
[0093] Intelligent Design and Automatic Optimization: The present invention utilizes a fine-tuned large language model (LLM) combined with an enhanced knowledge base to achieve the intelligent design and automatic optimization of controllers. Through the rich knowledge of the knowledge base and the reasoning ability of the intelligent agent, the system can automatically generate controller design schemes according to the control requirements input by the user, significantly reducing the design time and the need for manual intervention.
[0094] Strong Adaptability and Flexibility: By fine-tuning the large language model, the system has the ability to handle various control scenarios. Users only need to input specific control requirements, and the system can flexibly adapt to different types of control objectives (such as position control, speed control, etc.) and performance indicators, providing the best control scheme, with strong adaptability and meeting a wide range of application requirements.
[0095] Efficient Performance Optimization: The present invention uses the Monte Carlo tree search technique to globally optimize the combination of control operators, which can effectively explore various combinations of control algorithms and find the controller with the optimal performance. This method can not only improve the response speed and stability of the controller, but also ensure its robustness and anti-interference ability, thus enhancing the overall performance of the control system.
[0096] Data-Driven Fine-Tuning and Improvement: By constructing and utilizing a specific dataset to fine-tune the intelligent agent, the present invention realizes the continuous improvement of the controller design ability, enabling it to have high precision and effectiveness in different control scenarios. The fine-tuning dataset includes historical design data, simulation results, and performance evaluations, etc., ensuring that the generated controller can more accurately handle complex environments.
[0097] Simulation Verification and Closed-Loop Feedback: After the controller design is completed, the present invention comprehensively verifies its performance through a simulation environment. If the performance does not meet the set indicators, the system will automatically adjust the design and re-optimize and verify it, forming a closed-loop feedback mechanism. This process ensures that the output controller meets the expected performance in various scenarios and improves the reliability of the controller in practical applications.
[0098] Reducing the Burden of Manual Design: Traditional controller design requires a large amount of expert knowledge and manual adjustment, while the intelligent design platform of the present invention greatly reduces the dependence on manual intervention through automated design, simulation, and optimization processes. The system not only improves the efficiency of controller design, but also reduces the technical requirements for professionals, and has good popularization and application value.
[0099] Broad Application Prospects: Due to the intelligent, automated, and highly adaptable characteristics of the present invention, it can be applied to the control system design in multiple fields, such as industrial automation, robot control, unmanned driving, aerospace, etc. The system can quickly generate efficient controller schemes for these complex control tasks, greatly expanding the application scope of controller design.
[0100] In summary, the controller automation design system of the present invention provides significant improvements compared to the prior art through its efficient, accurate, flexible, and user-friendly features, offering a brand-new solution for the design and development of controllers.
[0101] The present invention realizes intelligent, efficient, and automated controller design through the enhancement and dynamic update of the knowledge base, the customized fine-tuning of the LLM, the automated process of controller design, the optimization method based on Monte Carlo tree search, the closed-loop simulation verification mechanism, and the deep collaboration between the agent and the knowledge base. The following key points and protected points together constitute the unique advantages and innovation of the present invention in controller design, providing a high-quality solution for intelligent control in multiple fields.
[0102] Controller design agent based on an enhanced knowledge base: The present invention constructs an enhanced knowledge base containing rich control domain knowledge to support the large language model (LLM) for controller design. The knowledge base covers various sources such as literature, historical simulation data, and expert experience, and is continuously updated and enhanced through search and reinforcement learning to provide comprehensive and up-to-date knowledge support for the agent. The construction and dynamic enhancement method of this knowledge base are one of the important innovation points and protected points of the present invention.
[0103] Fine-tuning the large language model for controller design: By using specific datasets (such as historical control system design data, simulation results, and performance evaluations) to fine-tune the large language model, it enables the model to understand and generate controller design schemes. After fine-tuning, the model can provide customized controller design schemes according to the specific needs of users. This data-driven fine-tuning method and the construction of specific datasets are important protected points of the present invention, ensuring that the agent has high adaptability and design capabilities in diverse control scenarios.
[0104] Automatic generation and simulation of controller design: Based on the prompt information such as control objectives, system types, and performance indicators input by the user, the agent automatically generates a controller design scheme, including the mathematical model of the controller and preliminary simulation results, by combining the control domain knowledge in the knowledge base. This automated design process significantly improves the efficiency and intelligence level of controller design and is one of the key innovation points and protected contents of the present invention.
[0105] Control Operator Library and Monte Carlo Tree Search (MCTS) Optimization: The system establishes an operator library containing various control operators (such as PID control, fuzzy control, etc.), which can be flexibly combined by the intelligent agent when designing the controller. Through global optimization of the operator combination by Monte Carlo tree search, the system can effectively explore various combinations of control algorithms, find the optimal controller design solution with the best performance, and ensure that its response speed, stability, and robustness reach the optimum. The application of Monte Carlo tree search and the combination of the operator library are important protection points of the present invention, significantly improving the overall performance of the controller. Standard MCTS mainly searches within a limited game state space, where each node represents a game state, and the algorithm evaluates the nodes and guides the search through multiple simulations (random game processes). The MCTS algorithm in this system not only searches for states in controller design but also needs to explore and optimize the combination and parameters of control algorithms. Controller design involves multiple complex control strategies, system parameters (such as PID control, fuzzy control, etc.), and performance evaluation indicators (such as system stability, response time, anti-interference ability, etc.). Therefore, the search space here is multi-dimensional, including not only the states of the system but also the selection of control strategies and the tuning of operators.
[0106] Simulation Verification and Closed-loop Feedback Mechanism: After the controller design is completed, the system comprehensively verifies the performance of the controller through a simulation environment. If the controller fails to meet the set performance standards, the system will automatically readjust the design for further optimization, forming a closed-loop feedback mechanism to ensure that the finally generated controller meets the expected performance requirements. This process of closed-loop verification and optimization is a key part of the present invention, ensuring the reliability and adaptability of the output controller and is an important protection point.
[0107] Collaboration between Intelligent Agent and Knowledge Base: The close combination of the intelligent agent and the enhanced knowledge base enables the intelligent agent to extract relevant control domain knowledge from the knowledge base and combine it with control operators for design. This process makes the generation of the controller significantly intelligent and targeted, ensuring that the system performs excellently in dealing with different control scenarios. The collaborative design mode of the intelligent agent combined with the knowledge base is also an innovation point and protected content of the present invention.
[0108] After being fine-tuned with a knowledge base in the control field and controller design experience, the large language model of the present invention can accurately identify the functions and applicable scenarios of different operators in the operator library. Based on the learned experience, it can not only select appropriate operators from numerous optional control operators, but also flexibly combine various candidate solutions according to specific control requirements and system variables. These candidate combinations are well-founded. The large language model uses historical data and design experience to judge the advantages of each combination in different scenarios, thus forming a set of solutions that are both representative and cover a variety of possibilities. However, it should be noted that this selection process does not enumerate all possible combinations, but is an intelligent screening by the large language model based on existing knowledge and experience. This strategy not only ensures that the number of combinations is large enough to effectively cover the complex control strategy space, but also avoids the waste of computing resources caused by enumeration.
[0109] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A method for designing a controller of heuristic Monte Carlo tree search for large language models, characterized in that, The method includes the following steps: S1. Construct an enhanced knowledge base; S2. Establish a private dataset in the control field to fine-tune the large language model; S3. Based on the requirements input by the user, the fine-tuned large language model extracts corresponding control knowledge from the knowledge base and generates a preliminary controller design scheme; S4. Establish an operator library and construct a simulation environment; S5. Select a control operator from the operator library and use the Monte Carlo tree search method to optimize the preliminary controller design scheme, and search for the optimal path as the optimal controller; S6. Test in the simulation environment. If the simulation result fails to meet the expected performance, optimize the optimal controller again. Otherwise, execute S7; S7. Output the mathematical structure of the controller and the simulation result.
2. The controller design method for heuristic Monte Carlo tree search inspired by large language models according to claim 1, characterized in that, Each node of the Monte Carlo tree represents a control operator and its parameter settings, and the path represents a complete operator combination and its corresponding parameter configuration scheme.
3. The controller design method of the large language model-inspired Monte Carlo tree search according to claim 2, wherein The specific steps for executing the Monte Carlo tree search method are as follows: First, select a control operator, combine the operator coefficients and system variables to generate a root node, where the operator coefficients and system variables are the parameter settings of the operator. Then, select the next node of the current node as the new current node through scoring. When the current node searched is a leaf node, generate new child nodes according to the operator of the current node, generate new paths for the new child nodes, evaluate the performance of the new paths, assign the upper confidence bound score obtained from the evaluation to each node on the new paths, and at the same time update the access times of each node, and finally obtain the optimal path.
4. The controller design method of large language model-inspired Monte Carlo tree search according to claim 3, characterized in that The optimal path satisfies: When the upper confidence bound score J of the new path is less than the threshold or the search times limit is reached, the path with the smallest upper confidence bound score J is the optimal path, and the controller corresponding to the optimal path is the optimal controller, where the upper confidence bound score is: J = α·OS + β·T s + γ·|e ss | OS represents the overshoot, which describes the maximum deviation amplitude of the system relative to the target value during the response process, T s represents the settling time, that is, the time required for the system output to first enter and continuously remain within a certain error band of the target value, |e ss | represents the steady-state error, which is the deviation between the system and the target value after reaching the steady state. α, β, and γ represent the weighting coefficients.
5. The controller design method of large language model-inspired Monte Carlo tree search according to claim 3, characterized in that The specific steps for generating new child nodes according to the operator of the current node are as follows: The large language model selects multiple candidate operators and their parameter settings from the pre-constructed operator library as new child nodes by using the control field knowledge and controller design experience obtained during the fine-tuning process.
6. The method for designing a controller of heuristic Monte Carlo tree search inspired by a large language model according to claim 3, characterized in that The control operators include proportional operators, integral operators, and derivative operators.
7. A method for designing a controller of heuristic Monte Carlo tree search inspired by a large language model, characterized in that, The requirements input by the user include control objectives, system types, and performance indicators.
8. A method for designing a controller of heuristic Monte Carlo tree search inspired by a large language model according to claim 1, characterized in that The specific steps for establishing a private dataset in the control field to fine-tune the large language model are as follows: Establish a private dataset in the control field. The dataset includes labeled control system design data, simulation results, performance evaluations, and case summaries. The private dataset in the control field conducts multiple training iterations on the large language model to complete the fine-tuning.
9. The controller design method of the large language model-inspired Monte Carlo tree search according to claim 1, characterized in that The simulation environment includes various factors in the actual working environment.
10. A method for designing a controller of heuristic Monte Carlo tree search inspired by a large language model according to claim 1, characterized in that, Re-optimizing the optimal controller specifically means executing S2 to S5 again.
Citation Information
Patent Citations
Design method of proportion differentiation lead intelligent model set PID controller
CN108107713A
Motor servo system model-free feedback controller design method based on improved Q learning algorithm
CN112034715A