An optimal pulse control method applied to an automated guided vehicle system

By analyzing the associated information data and periodic Riccati matrix of the automated guided vehicle system, and combining the Riccati equation and reinforcement learning, the optimal pulse control problem of the discrete-time AGV system was solved, achieving the system's optimality and stability, and improving operating efficiency and resource utilization.

CN122443503APending Publication Date: 2026-07-24SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610687352.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies face challenges in addressing the optimal pulse control problem of discrete-time AGV systems. These challenges include the lack of defined gradients for performance indicators due to the non-smoothness of the state trajectory at each pulse moment, insufficient convergence of adaptive dynamic programming algorithms, and a lack of research within the reinforcement learning framework. Consequently, it is difficult to construct a rigorous theoretical framework and analytical optimality conditions.

Method used

By acquiring the associated information data of the automated guided vehicle system, the pulse interval length is determined, and the periodic Riccati matrix is ​​iterated and its stability is analyzed until the optimal pulse control matrix is ​​output. Combining the Riccati equation and reinforcement learning, a hybrid triggering mechanism is designed to achieve optimal pulse control.

Benefits of technology

It improves the operating efficiency of the AGV system, reduces costs and mechanical wear, ensures the optimality and stability of pulse control, and enhances the robustness and resource efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122443503A_ABST
    Figure CN122443503A_ABST
Patent Text Reader

Abstract

The application discloses an optimal pulse control method applied to an automatic guided vehicle system, relates to the technical field of automatic control, and determines the pulse interval length in a current pulse period through associated information data of the automatic guided vehicle system; then, the period Lyapunov matrix is iterated based on the associated information data and the pulse interval length, and the stability of the iterated period Lyapunov matrix is analyzed; if the stability is unstable, iteration is repeated until the stability is stable, and the target period Lyapunov matrix is output; the optimal pulse control matrix corresponding to the automatic guided vehicle system is determined according to the target period Lyapunov matrix, and is provided to the automatic guided vehicle system to realize optimal pulse control. The application solves the problem that the optimal pulse cannot be solved and / or the convergence performance is poor in the prior art, ensures the optimality and stability of the pulse control of the automatic guided vehicle system through iteration and stability analysis of the period Lyapunov matrix, effectively improves the system operation efficiency, and reduces the cost and mechanical wear.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automatic control technology, specifically to an optimal pulse control method for use in automated guided vehicle systems. Background Technology

[0002] With the continuous deepening of the technological revolution, the guiding role of theoretical innovation in engineering practice is becoming increasingly prominent. Against this backdrop, Automated Guided Vehicle (AGV) systems, as core equipment for intelligent manufacturing and smart logistics, have seen the precision of motion control and the reliability of real-time path tracking become key bottlenecks determining the overall performance of the system. Faced with increasingly complex operating environments and highly dynamic task requirements, traditional control strategies are no longer sufficient, making the research and exploration of advanced control theories a hot topic of common concern in both academia and industry.

[0003] Constrained by limited resources and actuator wear in practical AGV systems, there is an urgent need to propose optimal control strategies with high resource efficiency. This would not only effectively avoid redundant data transmission and excessive energy consumption, significantly improving operational efficiency, but also reduce costs and mechanical wear. This places certain demands on the execution capability and stability of the controller. Pulse control, with its unique advantages, provides solutions to resource constraints and mechanical wear problems in control systems and has been widely used in various fields, including communication security, spacecraft rendezvous, and multi-agent systems. Unlike continuous control, pulse control applies control action only at discrete moments, allowing the AGV to evolve freely during pulse intervals. However, traditional analysis methods have limitations when applied to pulse systems; the non-differentiability of state transitions at pulse moments makes it difficult to directly apply classical optimal control theory. Therefore, it is necessary to develop a theoretical framework and computational tools specifically for pulse systems. Researchers have developed an optimal control method based on adaptive dynamic programming, approximating the optimal value function through offline training of neural networks, thereby enabling real-time online decision-making for optimal pulse triggering. Meanwhile, some scholars have used measure-driven differential inclusion to model impulse behavior and proposed optimality conditions based on the Hamilton-Jacobi-Bellman equations for impulse control systems. They have derived a comprehensive method for feedback optimal control using impulse Euler solutions and invariance results, constructing a feedback-based impulse control strategy and extending it from open-loop to closed-loop control. However, due to the non-smoothness of the state trajectory at the impulse moment, the coupling problem between continuous and discrete dynamics, and the difficulty in directly applying traditional optimization tools based on variational methods or differential rules, the optimal impulse control of discrete-time AGV systems has not been thoroughly studied. Furthermore, a rigorous theoretical framework has not yet been established for the necessary and sufficient conditions of the optimal impulse control problem for discrete-time AGV systems.

[0004] In summary, existing technologies still face the following key scientific problems in dealing with the optimal pulse control problem for discrete-time AGV systems: (1) How to solve the problem of missing performance index gradient definition caused by the non-smoothness of the state trajectory at the pulse moment, thereby avoiding the structural complexity or even the inability to obtain closed-form solutions in the context of the classical maximum principle or Hamilton-Jacobi-Bellman equation under the pulse system framework, and thus constructing analytical optimality conditions applicable to discrete-time pulse systems. (2) How to ensure that the convergence process of the adaptive dynamic programming algorithm to approach the optimal solution is rigorously solved from the overall perspective, and establish a complete analytical framework for the global convergence of the algorithm from the theoretical level, so as to ensure that the adaptive dynamic programming algorithm can monotonically approach the optimal performance index in the entire iterative optimization process and finally converge to the optimal solution. (3) How to systematically study the optimal pulse control problem of discrete-time systems under the reinforcement learning framework, so as to fill the research gap in the field of pulse control, which is mostly concentrated on continuous-time or standard discrete-time systems. Summary of the Invention

[0005] The purpose of this application is to provide an optimal pulse control method for automated guided vehicle systems, which solves the problems of the inability to solve for the optimal pulse and / or poor convergence performance in the prior art.

[0006] This application is achieved through the following technical solution:

[0007] An optimal pulse control method for use in automated guided vehicle (AGV) systems includes:

[0008] Acquire the associated information data of the automated guided vehicle system, and simultaneously determine the pulse interval length of the automated guided vehicle system within the current pulse cycle;

[0009] Based on the associated information data and the pulse interval length, the periodic Licatti matrix is ​​iterated to determine the periodic Licatti matrix after iteration;

[0010] A stability analysis is performed on the periodic Likati matrix after the iteration to determine the stability analysis result; the stability analysis result includes whether it is stable or unstable.

[0011] If the stability analysis result is unstable, the periodic Licati matrix is ​​iterated repeatedly until the stability analysis result is stable, and the target periodic Licati matrix is ​​output.

[0012] The optimal pulse control matrix corresponding to the automated guided vehicle system is determined based on the target period Licati matrix, and the optimal pulse control matrix is ​​provided to the automated guided vehicle system to achieve optimal pulse control of the automated guided vehicle system.

[0013] In one possible implementation, acquiring the associated information data of the automated guided vehicle system includes:

[0014] The input matrix, system matrix, state weight matrix, impulse state weight matrix, and control weight matrix of the automated guided vehicle system are obtained to acquire the associated information data of the automated guided vehicle system.

[0015] In one possible implementation, the pulse interval length of the automated guided vehicle system within the current pulse period is determined as follows:

[0016] ;

[0017] in, Indicates the pulse interval length. Indicates the first The moment before the next pulse Indicates the first The moment after the next pulse.

[0018] In one possible implementation, based on the associated information data and the pulse interval length, the periodic Riccati matrix is ​​iterated to determine the iterated periodic Riccati matrix as follows:

[0019] ;

[0020] in, Let the periodic Licatti matrix be the matrix in the j-th iteration. Let represent the periodic Licati matrix during the (j+1)th iteration, i.e., the periodic Licati matrix after the iteration. Represents the pulse state weight matrix. This represents the first intermediate matrix obtained through the state weight matrix. This represents the second intermediate matrix obtained through the system matrix and the sampling period. This represents the pulse interval length, and T represents the transpose operation on the matrix. R represents the third intermediate matrix obtained through the system matrix and the input matrix, and R represents the control weight matrix.

[0021] In one possible implementation, the first intermediate matrix obtained through the state weight matrix is:

[0022] ;

[0023] in, This represents the first intermediate matrix. Let m represent the state weight matrix, and m represent the index.

[0024] In one possible implementation, the second intermediate matrix obtained through the system matrix and the sampling period is:

[0025] ;

[0026] in, This represents the second intermediate matrix. Represents the matrix index. Indicates the sampling period. This represents the system matrix.

[0027] In one possible implementation, the third intermediate matrix obtained through the system matrix and the input matrix is:

[0028] ;

[0029] in, This represents the third intermediate matrix. Represents the matrix index. Indicates the dummy element of the integral. Indicates the sampling period. Represents the system matrix. Represents the input matrix,

[0030] In one possible implementation, a stability analysis is performed on the periodic Likati matrix after the iteration, and the stability analysis results are determined, including:

[0031] Determine whether the Karti matrix converges in the period after the iteration. If it does, the stability analysis result is determined to be stable; otherwise, the stability analysis result is determined to be unstable.

[0032] In one possible implementation, determining the optimal pulse control matrix corresponding to the automated guided vehicle system based on the target periodic Ricardi matrix includes:

[0033] Based on the target periodicity Likati matrix, the first intermediate matrix, and the second intermediate matrix, obtain the fourth intermediate matrix;

[0034] Based on the fourth intermediate matrix, the optimal pulse control matrix corresponding to the automated guided vehicle system is determined as follows:

[0035] ;

[0036] in, This represents the optimal pulse control matrix. This represents the fourth intermediate matrix in the j-th iteration process. Indicates the first The moment after the next pulse Indicates the first The state variables before the next pulse.

[0037] In one possible implementation, the fourth intermediate matrix is ​​obtained based on the target periodicity Licatti matrix, the first intermediate matrix, and the second intermediate matrix: .

[0038] Compared with the prior art, this application has the following advantages and beneficial effects:

[0039] This application provides an optimal pulse control method for automated guided vehicles (AGV) systems. The method determines the pulse interval length within the current pulse period using the associated information data of the AGV system. Then, iterates the periodic Licati matrix based on the associated information data and the pulse interval length, and performs stability analysis on the iterated periodic Licati matrix. If unstable, the iteration is repeated until stable, and a target periodic Licati matrix is ​​output. The optimal pulse control matrix corresponding to the AGV system is determined based on the target periodic Licati matrix and provided to the AGV system to achieve optimal pulse control. This application solves the problems of inability to solve for the optimal pulse and / or poor convergence performance in existing technologies. Through iteration and stability analysis of the periodic Licati matrix, it ensures the optimality and stability of the pulse control for the AGV system, effectively improving system operating efficiency and reducing costs and mechanical wear. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the exemplary embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0041] Figure 1 A flowchart illustrating an optimal pulse control method for an automated guided vehicle (AGV) system, provided in an embodiment of this application;

[0042] Figure 2 A schematic diagram of the control framework of an AGV system employing an optimal pulse control method applied to an automated guided vehicle system, provided for an embodiment of this application;

[0043] Figure 3 This is a schematic diagram comparing the status responses of AGVs provided in an embodiment of this application.

[0044] Figure 4 A comparative diagram illustrating the control strategies provided in the embodiments of this application;

[0045] Figure 5 This is a schematic diagram comparing state paths under different control gains provided in the embodiments of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the embodiments and accompanying drawings. The illustrative embodiments and descriptions of this application are only for explaining this application and are not intended to limit this application.

[0047] Example 1

[0048] like Figure 1 As shown, this application provides an optimal pulse control method for an automated guided vehicle (AGV) system, comprising:

[0049] S101. Obtain the associated information data of the automated guided vehicle system, and simultaneously determine the pulse interval length of the automated guided vehicle system in the current pulse cycle;

[0050] S102. Based on the associated information data and the pulse interval length, iterate the periodic Licatti matrix to determine the iterated periodic Licatti matrix.

[0051] S103. Perform stability analysis on the periodic Likati matrix after the iteration, and determine the stability analysis result; the stability analysis result includes stable or unstable.

[0052] S104. If the stability analysis result is unstable, repeatedly iterate the periodic Licati matrix until the stability analysis result is stable, and output the target periodic Licati matrix.

[0053] S105. Determine the optimal pulse control matrix corresponding to the automated guided vehicle system based on the target period Licati matrix, and provide the optimal pulse control matrix to the automated guided vehicle system to achieve optimal pulse control of the automated guided vehicle system.

[0054] To facilitate understanding of the technical solutions described in the embodiments of this application by those skilled in the art, the closed-loop system of the automated guided vehicle system involved in the embodiments of this application is first modeled as follows.

[0055] To reflect the instantaneous motion of the AGV, and without considering the effects of wind and longitudinal forces on lateral motion, the vehicle dynamics equations are as follows:

[0056] ;

[0057] in, This represents the total lateral force on the front axle. Indicates the front wheel steering angle. The total vehicle weight is expressed in... express, This represents the lateral velocity, the physical quantity above it. This represents the derivative of the physical quantity with respect to time, i.e., This indicates lateral acceleration, and the same applies later in the text. Indicates yaw rate. Indicates yaw acceleration. Indicates longitudinal velocity. The total lateral force on the rear axle is represented by . Indicates the moment of inertia of yaw rotation. and These represent the front wheelbase and the rear wheelbase, respectively.

[0058] When the deflection angle is very small, using a small angle approximation, we can obtain:

[0059] ;

[0060] in, This represents the sideslip angle at the center of mass, while and These represent the camber angles of the front and rear wheels, respectively. ;in, and These are the steering stiffnesses of the front and rear wheels, respectively.

[0061] From this, we can deduce that:

[0062] ;

[0063] in, Indicates the sideslip angular velocity. It represents the yaw acceleration.

[0064] To control the dynamics of a vehicle traveling along the lane centerline, it is also necessary to obtain information about the surrounding environment and additional conditions related to the vehicle's state. By simplifying the complexity of a specific application scenario (ignoring road curvature), the vehicle path tracking equation can be obtained as follows:

[0065] ;

[0066] in, Indicates lateral positional deviation. This represents the rate of change of the lateral deviation. Indicates the aiming distance. Indicates heading error. This indicates the rate of change of heading error.

[0067] Therefore, the following dynamic equation for a continuous-time AGV can be derived:

[0068] ;

[0069] in, Represents state variables The derivative with respect to time, and , Indicates the sideslip angle at the center of mass. Indicates yaw rate. Indicates heading error. Indicates lateral positional deviation. Representing a dimension as The matrix, It is a control input. For the system matrix, The input matrix is ​​denoted as .

[0070] The system matrix is:

[0071] ;

[0072] ;

[0073] , ;

[0074] , , ;

[0075] in, Representation matrix The element in the first row and first column Representation matrix The element in the first row and second column, Representation matrix The element in the second row and first column, Representation matrix The element in the second column and second row, Representation matrix The first row of elements Representation matrix Second row elements

[0076] Consider discretizing a typical continuous AGV system, assuming a sampling period of . The time step is The following solutions exist:

[0077] ;

[0078] in, Let e ​​be the integral variable, and let e represent the matrix exponent. This indicates that the system, without input, starts from... Evolved to The state transition. Indicates the initial time. Indicates the current time. Indicates from Evolved to The state transition matrix.

[0079] For any and have:

[0080] ;

[0081] in, Representing discrete time The state vector at time t, Representing discrete time The state vector at time t, Indicates the dummy element of the integral. This represents the discrete-time control input. , This represents the integral variable.

[0082] Let the AGV second intermediate matrix The third intermediate matrix .in, , Let the state variable express Control input express Here, It is pulse control. . Indicates at the sampling time The centroid sideslip angle, Indicates at the sampling time yaw rate, Indicates at the sampling time The heading error, Indicates at the sampling time Lateral positional deviation.

[0083] The discrete-time AGV system can then be represented as:

[0084] ;

[0085] in, express The state variable at any given time.

[0086] Furthermore, pulse control has the following definition:

[0087]

[0088] in, It is the Dirac discrete-time impulse function. Indicates the first Each pulse moment, Indicates at time The pulse size at that time.

[0089] For all All ( (where the pulse time is). Therefore, under pulse control, the closed-loop system can be described as:

[0090] ;

[0091] in, and These represent the states before and after the pulse. Indicates the first The instant after the next pulse. Indicates the first The instant before the next pulse. This is the initial state. According to the Kalman rank condition, It is pulse-controllable, which ensures the existence of an allowable control. Such that for any initial state, we have Next, the following performance metrics are constructed:

[0092] ;

[0093] in, and These are the state weight matrix, the impulse state weight matrix, and the control weight matrix, respectively. This represents the cost of the free evolution phase within the pulse interval. This represents the instantaneous state cost at each pulse moment, while The cost of corresponding pulse control. Matrix and All are symmetric positive definite matrices. k represents the k-th matrix. Sub-pulse control event index.

[0094] The closed-loop system obtained above can be used to derive the subsequent optimal pulse control solution formula, while the performance index can be used to evaluate the effect of the optimal pulse control.

[0095] In one possible implementation, acquiring the associated information data of the automated guided vehicle system includes:

[0096] The input matrix, system matrix, state weight matrix, impulse state weight matrix, and control weight matrix of the automated guided vehicle system are obtained to acquire the associated information data of the automated guided vehicle system.

[0097] In one possible implementation, the pulse interval length of the automated guided vehicle system within the current pulse period is determined as follows:

[0098] ;

[0099] in, Indicates the pulse interval length. Indicates the first The moment before the next pulse Indicates the first The moment after the next pulse.

[0100] In one possible implementation, based on the associated information data and the pulse interval length, the periodic Riccati matrix is ​​iterated to determine the iterated periodic Riccati matrix as follows:

[0101] ;

[0102] in, Let the periodic Licatti matrix be the matrix in the j-th iteration. Let represent the periodic Licati matrix during the (j+1)th iteration, i.e., the periodic Licati matrix after the iteration. Represents the pulse state weight matrix. This represents the first intermediate matrix obtained through the state weight matrix. This represents the second intermediate matrix obtained through the system matrix and the sampling period. This represents the pulse interval length, and T represents the transpose operation on the matrix. R represents the third intermediate matrix obtained through the system matrix and the input matrix, and R represents the control weight matrix.

[0103] In one possible implementation, the first intermediate matrix obtained through the state weight matrix is:

[0104] ;

[0105] in, This represents the first intermediate matrix. Let m represent the state weight matrix, and m represent the index.

[0106] In one possible implementation, the second intermediate matrix obtained through the system matrix and the sampling period is:

[0107] ;

[0108] in, This represents the second intermediate matrix. Represents the matrix index. Indicates the sampling period. This represents the system matrix.

[0109] In one possible implementation, the third intermediate matrix obtained through the system matrix and the input matrix is:

[0110] ;

[0111] in, This represents the third intermediate matrix. Represents the matrix index. Indicates the dummy element of the integral. Indicates the sampling period. Represents the system matrix. Represents the input matrix,

[0112] In one possible implementation, a stability analysis is performed on the periodic Likati matrix after the iteration, and the stability analysis results are determined, including:

[0113] Determine whether the Karti matrix converges in the period after the iteration. If it does, the stability analysis result is determined to be stable; otherwise, the stability analysis result is determined to be unstable.

[0114] In one possible implementation, determining the optimal pulse control matrix corresponding to the automated guided vehicle system based on the target periodic Ricardi matrix includes:

[0115] Based on the target period's Kalti matrix, the first intermediate matrix, and the second intermediate matrix, the fourth intermediate matrix is ​​obtained as follows: ;

[0116] Based on the fourth intermediate matrix, the optimal pulse control matrix corresponding to the automated guided vehicle system is determined as follows:

[0117] ;

[0118] in, This represents the optimal pulse control matrix. This represents the fourth intermediate matrix in the j-th iteration process. Indicates the first The moment after the next pulse Indicates the first The state variables before the next pulse.

[0119] Based on the above-mentioned content regarding obtaining the optimal pulse control matrix, the embodiments of this application derive and explain the principles involved, as follows.

[0120] To establish an optimal pulse controller for this AGV system, this application derives the Riccati equation based on the Bellman optimality principle and extends it to the periodic Riccati equation. A novel hybrid triggering mechanism is designed, which combines state feedback during pulseless periods with pulse compensation during pulsed periods. This effectively suppresses energy loss caused by frequent communication or actuator wear while ensuring the AGV path tracking accuracy.

[0121] The expression for the Riccati equation is as follows:

[0122] ;

[0123] in, It is a positive definite matrix. It is a standard formation. Indicates time The Ricardi matrix, Indicates time The Ricardi matrix, Indicates the first After one pulse, at the time... Indicates the first The previous sampling time before the next pulse. Indicates the first Ricardi matrix before the second pulse Indicates the first The Riccati matrix following the next pulse Represents the Riccati matrix at infinity.

[0124] Furthermore, the optimal value function has the following form:

[0125] ;

[0126] And the corresponding optimal pulse control is:

[0127] ;

[0128] in, Representing discrete time The optimal value function, Indicates the first The state variables before the next pulse.

[0129] The Riccati equation can be solved using the following procedure:

[0130] 1. Initialization: Let any in Let be any positive integer; where, Indicates the first The index of the time after each pulse.

[0131] 2. Reverse recursion: For all moments within the current pulse period Calculate sequentially: At the pulse moment First calculate Update again ;in, This represents the fifth intermediate matrix. Indicates the first The Riccati matrix at the next moment after the subpulse Indicates the first The Ricardi matrix following the second pulse.

[0132] 3. Update the loop variable: Let Then, proceed to the next earlier pulse interval and continue the reverse solution back to the initial time.

[0133] 4. Output results.

[0134] Therefore, the optimal pulse control can be determined.

[0135] Based on the above solution process of the Riccati equation, this application further provides a periodic Riccati iterative design based on reinforcement learning, as detailed below.

[0136] Since all matrix elements in the system are constants, and given the time interval... For a fixed value, define the following matrix sequence: in This corresponds to the instant before the pulse occurs, i.e., the start and end of the cycle. In each cycle of length... Within the pulse interval, Indicates the remaining pulse period The periodicity of the step is represented by the Kati matrix, and there exists the following simple correspondence: …, .in, This indicates the start point of the current pulse cycle, at which point the number of remaining steps in the cycle is... , This represents the first moment of the system's automatic smoke evolution after the pulse occurs, at which point the remaining number of steps in the cycle is... , This indicates the last step of the current pulse cycle, at which point the number of steps remaining in the cycle is... .

[0137] Therefore, the periodic Riccati equation is defined as follows:

[0138] ;

[0139] in, This indicates the number of steps remaining in the current cycle until the next pulse. Indicates the remaining pulse period The periodic Liccati matrix at step -1, This represents the periodic Licatti matrix at the start of the current pulse cycle.

[0140] According to the Bellman optimality principle, we have the following equation:

[0141] ;

[0142] in, Indicates the first The optimal value function before the next pulse. Indicates the first The moment before the next pulse Indicates the first The optimal value function before the next pulse.

[0143] The following is given for solving: Generalized equation:

[0144] ;

[0145] Among them, superscript Indicates the first iteration This indicates pulse control. The The next iteration, below Similarly, it is clear that determining the optimal control law... Therefore, the optimal value function must first be determined. .

[0146] Once the strategy is obtained The algorithm then performs the following iterations:

[0147] ;

[0148] in, Indicates the first The value function before the next pulse. iteration Indicates the first The value function before the next pulse. The next iteration.

[0149] make And iterate repeatedly as described above. Based on the Bellman optimality principle, when At that time, the following two key equations hold true:

[0150] ;

[0151] in, Indicates the first The infiniteth iteration of the value function before the next pulse. Indicates the first The optimal value function before the next pulse. Indicates the first The infinite number of iterations of pulse control before the next pulse. Indicates the first Optimal pulse control before the next pulse.

[0152] Because obtaining optimal pulse control requires consideration In general discrete-time systems, based on the closed-loop system equations, the following can be obtained using a recursive method: And thus and Establishing a connection. Among them, It is the state transition matrix. It is a constant.

[0153] Based on the aforementioned properties of the periodic Riccati equation, we can obtain:

[0154] ;

[0155] in, 'm' represents the step index within the current period. Next, the strategy is used... The iterative formula yields the following expression for solving the control update strategy:

[0156]

[0157] By simplifying the relationships and substituting the equations, we can obtain the equations for solving the pulse control problem and the equations for solving the performance indicators:

[0158]

[0159]

[0160] Therefore, we can conclude the following about Value iteration conclusion:

[0161]

[0162] For discrete-time systems, the above can be derived. The periodic Riccati equation with a fixed pulse interval is solved by iterative equations, and the equations for pulse control are solved iteratively to obtain the optimal pulse control. The iterative process is as follows:

[0163] 1. Initialization: Collect information, and set... in For any constant array, further define the pulse interval length. Define the remaining steps index. ;

[0164] 2. Value iteration: based on Perform iterations;

[0165] 3. Judgment: If If it does not converge, repeat step 2 until... Approaching ;

[0166] 4. Output control results.

[0167] Stability analysis is as follows:

[0168] Select As a candidate Lyapunov function, considering non-impulse times, we can obtain the following equation:

[0169] ;

[0170] At the instant of the pulse, the following relationship can be obtained:

[0171] ;

[0172] At the same time, given ,in, We can obtain:

[0173] ;

[0174] because , , can be obtained Therefore, it can be concluded that the Lyapunov function is non-increasing. Considering... The optimal control point is the only equilibrium point, and under optimal control, the controlled system is globally asymptotically stable. Based on this analysis, for a periodic pulse control system, since the closed-loop system is stable under optimal control, the system state remains bounded as it evolves along the optimal strategy within the pulse period, thus ensuring the iterative sequence... It converges asymptotically. Furthermore, using the principles of dynamic programming, theoretically, it would require an infinite number of iterations to find the optimal solution. In the specific simulation verification, a threshold was set. ,when This indicates the sequence. The system gradually converges along the optimal strategy, at which point the optimal solution is reached. .

[0175] Compared with the prior art, this application has the following outstanding advantages:

[0176] 1. This application focuses on the infinite-time domain optimal impulse control problem of discrete-time linear AGV systems. By utilizing the known linear structure of the AGV discrete-time model, the value function is expressed as a quadratic form, thus transforming the optimal impulse control problem into an iteratively solvable Riccati equation. This method not only retains the efficiency and flexibility of reinforcement learning (RL) methods in iterative computation but also fully utilizes system model information, thereby ensuring the convergence of control and the stability of the system at the algorithm level. It provides a new approach for high-precision AGV motion control that combines theoretical rigor with computational feasibility.

[0177] 2. Addressing the specific requirement of AGV systems to simultaneously handle continuous motion processes and discrete control events in pulse control, this application establishes the necessity and sufficiency conditions for the optimal control strategy based on the Riccati equations, seamlessly integrating inter-pulse dynamics and pulse-time dynamics. Furthermore, through a completeness analysis of the Riccati equations, a more comprehensive theoretical constraint system is presented.

[0178] 3. This application further extends the aforementioned Riccati equation framework to the scenario of periodic pulse action. By introducing periodic pulse time constraints and system dynamic evolution laws, a novel periodic hybrid Riccati equation system is constructed. This system not only characterizes the coupling relationship between continuous-time dynamics between pulses and discrete jumps in pulse timing, but also incorporates periodic boundary conditions into the solution process of the optimal control law. Based on this theory, this application designs an iterative solution algorithm with a recursive structure and proves the convergence of the algorithm in the infinite time domain.

[0179] Example 2

[0180] like Figure 2This application presents an optimal pulse control framework for Automated Guided Vehicles (AGVs) under discrete sampling and discrete execution conditions. The system consists of the AGV itself, a pulse controller, a communication network, sensors, and actuators. The AGV integrates multiple sensors to perceive the environmental state in real time and performs physical actions such as driving, steering, and loading / unloading through actuators. The pulse controller, as the core command issuing unit, connects to the AGV through a limited communication network to achieve vehicle scheduling and control. Given the limited communication resources and actuator wear in practical applications, traditional continuous-time control strategies are difficult to apply directly. Therefore, this framework introduces a pulse control mechanism at the front end of the communication network. This mechanism possesses adaptive and dynamic adjustment capabilities: during time intervals without pulse application, the system follows a "free evolution" law, naturally evolving based on the AGV's own dynamics and environmental interactions; after a periodically changing time step, the pulse controller generates discrete control inputs to regulate the system state. This pulsed intervention reduces communication frequency and actuator action frequency, while applying optimal control at critical moments, thereby achieving coordinated optimization of AGV path planning and behavior decision-making in complex environments and effectively improving the system's robustness and resource efficiency.

[0181] The front and rear wheelbases are respectively and The lateral stiffness of the front and rear wheels are respectively and longitudinal velocity Pre-aiming distance Select sampling period Discretize the continuous-time system, pulse interval The weight matrices are respectively set as and . Figure 2 It presents different control effects and The process of change.

[0182] Figure 3 This demonstrates that under pulse control, and It exhibits the smoothest characteristics, with advantages such as fast convergence speed and small error, outperforming other control methods. In contrast, optimal pulse control in control... While its performance is not as good as continuous control, it is far superior to the relaxed continuous-time actor-critic algorithm and the direct neural dynamic programming algorithm.

[0183] Figure 4This demonstrates a key characteristic of optimal pulse control: it triggers only at the pulse moment, minimizing the number of control executions to just 35. This not only effectively reduces resource waste but also lowers equipment wear, a crucial consideration in practical engineering applications.

[0184] (2) To verify the effectiveness of the proposed optimal control method, the system will be compared under the conditions of manually set gain matrix control and optimal gain matrix control. and The trajectory. Consider the following specific parameters for the discrete-time system: The weight matrices are set as follows: , , ; Initial value Set as Next, set up 5 different manual gain matrices: The optimal gain matrix obtained is: . Figure 5 Initial values ​​were displayed. Variations under 6 different control gains.

[0185] The attached figures compare the state variable trajectories under manual control and optimal control. It can be seen that the optimal control method proposed in this application enables the state variables to converge rapidly to near zero during a smooth transition process, and the optimal control gain matrix corresponds to the minimum cost, effectively proving its optimality.

[0186] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0187] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0188] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0189] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0190] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0191] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. An optimal pulse control method applied to an automated guided vehicle (AGV) system, characterized in that, include: Acquire the associated information data of the automated guided vehicle system, and simultaneously determine the pulse interval length of the automated guided vehicle system within the current pulse cycle; Based on the associated information data and the pulse interval length, the periodic Licatti matrix is ​​iterated to determine the periodic Licatti matrix after iteration; A stability analysis is performed on the periodic Likati matrix after the iteration to determine the stability analysis result; the stability analysis result includes whether it is stable or unstable. If the stability analysis result is unstable, the periodic Licati matrix is ​​iterated repeatedly until the stability analysis result is stable, and the target periodic Licati matrix is ​​output. The optimal pulse control matrix corresponding to the automated guided vehicle system is determined based on the target period Licati matrix, and the optimal pulse control matrix is ​​provided to the automated guided vehicle system to achieve optimal pulse control of the automated guided vehicle system.

2. The optimal pulse control method for an automated guided vehicle system according to claim 1, characterized in that, Obtain relevant information data about the automated guided vehicle (AGV) system, including: The input matrix, system matrix, state weight matrix, impulse state weight matrix, and control weight matrix of the automated guided vehicle system are obtained to acquire the associated information data of the automated guided vehicle system.

3. The optimal pulse control method for an automated guided vehicle system according to claim 1, characterized in that, The pulse interval length of the automated guided vehicle system within the current pulse period is determined as follows: ; in, Indicates the pulse interval length. Indicates the first The moment before the next pulse Indicates the first The moment after the next pulse.

4. The optimal pulse control method for an automated guided vehicle system according to claim 2, characterized in that, Based on the aforementioned associated information data and the pulse interval length, the periodic Licati matrix is ​​iterated to determine the periodic Licati matrix after iteration as follows: ; in, Let the periodic Licatti matrix be the matrix in the j-th iteration. Let represent the periodic Licati matrix during the (j+1)th iteration, i.e., the periodic Licati matrix after the iteration. Represents the pulse state weight matrix. This represents the first intermediate matrix obtained through the state weight matrix. This represents the second intermediate matrix obtained through the system matrix and the sampling period. This represents the pulse interval length, and T represents the transpose operation on the matrix. R represents the third intermediate matrix obtained through the system matrix and the input matrix, and R represents the control weight matrix.

5. The optimal pulse control method for an automated guided vehicle system according to claim 4, characterized in that, The first intermediate matrix obtained through the state weight matrix is: ; in, This represents the first intermediate matrix. Let m represent the state weight matrix, and m represent the index.

6. The optimal pulse control method for an automated guided vehicle system according to claim 4, characterized in that, The second intermediate matrix obtained through the system matrix and the sampling period is: ; in, This represents the second intermediate matrix. Represents the matrix index. Indicates the sampling period. This represents the system matrix.

7. The optimal pulse control method for an automated guided vehicle system according to claim 4, characterized in that, The third intermediate matrix obtained through the system matrix and the input matrix is: ; in, This represents the third intermediate matrix. Represents the matrix index. Indicates the dummy element of the integral. Indicates the sampling period. Represents the system matrix. This represents the input matrix.

8. The optimal pulse control method for an automated guided vehicle system according to claim 1, characterized in that, A stability analysis is performed on the periodic Likati matrix after the iteration, and the stability analysis results are determined, including: Determine whether the Karti matrix converges in the period after the iteration. If it does, the stability analysis result is determined to be stable; otherwise, the stability analysis result is determined to be unstable.

9. The optimal pulse control method for an automated guided vehicle system according to claim 4, characterized in that, The optimal pulse control matrix for the automated guided vehicle system is determined based on the target period Ricardi matrix, including: Based on the target periodicity Likati matrix, the first intermediate matrix, and the second intermediate matrix, obtain the fourth intermediate matrix; Based on the fourth intermediate matrix, the optimal pulse control matrix corresponding to the automated guided vehicle system is determined as follows: ; in, This represents the optimal pulse control matrix. This represents the fourth intermediate matrix in the j-th iteration process. Indicates the first The moment after the next pulse Indicates the first The state variables before the next pulse.

10. The optimal pulse control method for an automated guided vehicle system according to claim 9, characterized in that, Based on the target period's Kalti matrix, the first intermediate matrix, and the second intermediate matrix, the fourth intermediate matrix is ​​obtained as follows: .