Control apparatus, control system, control program, and control method
The control device and system address the challenge of adapting to changing probability distributions by constructing predictive distributions and extracting feature values to calculate control parameters, ensuring stable and efficient adaptive control.
Patent Information
- Application Number
- PCT/JP2025/026653
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2025-07-28
- Publication Date
- 2026-03-05
AI Technical Summary
Existing adaptive control systems fail to detect and adapt to continuous or discrete changes in the probability distribution of parameters affecting controlled objects, leading to instability and inefficiency.
A control device and system that includes a distribution prediction unit, an extraction unit, and a control unit to construct predictive distributions from sample values, extract feature values, and calculate control parameters based on these distributions, enabling stable adaptive control.
The system achieves stable adaptive control by dynamically adjusting to changes in the probability distribution of parameters, ensuring robustness and effectiveness in controlling complex systems.
Smart Images

Figure JP2025026653_05032026_PF_FP_ABST
Abstract
Description
Control device, control system, control program, and control method
[0001] The present disclosure relates to a control device, a control system, a control program, and a control method that perform adaptive control in response to changes over time in the probability distribution of a parameter that changes the characteristics of a controlled object.
[0002] Conventionally, a configuration for adaptive control based on temporal changes in the probability distribution of a parameter that changes the characteristics of a controlled object (adaptive control) has been known. For example, International Publication No. 2024 / 009418 (Patent Document 1) discloses a remote control device that estimates the probability distribution of a transmission delay using a transmission delay model based on the mode of the transmission delay between a mobile object and the remote control device. This remote control device can change the target behavior of the mobile object depending on the transmission delay mode, thereby improving the safety of remote control of the mobile object via a network.
[0003] In addition, in Non-Patent Document 1, a Lyapunov inequality is derived that characterizes the stability of a linear stochastic system without limiting the class of stochastic processes. In Non-Patent Document 2, a configuration related to a non-linear stochastic system is disclosed. In Non-Patent Document 3, H 2 H, an evaluation index of control performance based on the norm 2 Conditional expressions that can evaluate performance are being derived.
[0004] International Publication No. 2024 / 009418
[0005] Y. Hosoe and T. Hagiwara, "On Second-Moment Stability of Discrete-Time Linear Systems With General Stochastic Dynamics", IEEE Transactions on Automatic Control, vol. 67, no. 2, pp. 795-809, 2022. Y. Kawano and Y. Hosoe, "Contraction Analysis of Discrete-Time Stochastic Systems", IEEE Transactions on Automatic Control, vol. 69, no. 2, pp. 982-997, 2024. Y. Hosoe, T. Okamoto, and T. Hagiwara, "H2 Performance Analysis and Synthesis for Discrete-Time Linear Systems With Dynamics Determined by an iid Process", IEEE Control Systems Letters, vol. 7, pp. 751-756, 2023.
[0006] In Patent Document 1, a remote control device, which is a controller, outputs a control input calculated using a control gain based on information obtained from a mobile object, etc., to the mobile object, which is a controlled object. To achieve stable adaptive control of temporal changes in the probability distribution of a parameter, such as transmission delay, which changes the characteristics of the controlled object, it is necessary to detect continuous or discrete changes in the probability distribution and have the control gain follow those changes. However, Patent Document 1 does not disclose a configuration for detecting such changes in the probability distribution of transmission delay, and does not realize a relationship between the changes in the probability distribution and the control gain. Furthermore, Non-Patent Documents 1 to 3 assume that the distribution of parameters in the controlled object is known, and therefore do not realize control adaptive to temporal changes in the distribution.
[0007] This has been made to solve the above-mentioned problems, and its purpose is to realize stable adaptive control.
[0008] A control device according to one aspect of the present disclosure outputs a first parameter to a control object, the first parameter changing the state of the control object. The control device includes a distribution prediction unit, an extraction unit, and a control unit. The distribution prediction unit constructs a predictive distribution from at least one sample value of a second parameter changing the characteristic of the control object, the sample value corresponding to at least one time including the current time. The extraction unit extracts a feature value from the predictive distribution. The control unit calculates the first parameter based on the feature value and the state of the control object.
[0009] A control system according to another aspect of the present disclosure includes a control object and a control device. The control device outputs a first parameter to the control object, the first parameter changing the state of the control object. The control device includes a distribution prediction unit, an extraction unit, and a control unit. The distribution prediction unit constructs a predictive distribution from at least one sample value of a second parameter changing the characteristic of the control object, the sample value corresponding to at least one time including the current time. The extraction unit extracts a feature value from the predictive distribution. The control unit calculates the first parameter based on the feature value and the state of the control object.
[0010] A control program according to another aspect of the present disclosure, when executed by a processor, causes the processor to perform processing for outputting a first parameter that changes a state of the controlled object to the controlled object, the processing including: constructing a predictive distribution from at least one sample value of a second parameter that changes a characteristic of the controlled object, the sample value corresponding to at least one time including the current time; extracting a feature value from the predictive distribution; and calculating the first parameter based on the feature value and the state of the controlled object.
[0011] A control method according to another aspect of the present disclosure is a control method for performing processing to output a first parameter that changes a state of a controlled object to the controlled object, the control method including the steps of: constructing a predictive distribution from at least one sample value of a second parameter that changes a characteristic of the controlled object, the sample value corresponding to at least one time including the current time; extracting a feature value from the predictive distribution; and calculating the first parameter based on the feature value and the state of the controlled object.
[0012] According to the control device, control system, control program, and control method disclosed herein, stable adaptive control can be achieved by calculating the first parameter using a feature extracted from the predictive distribution of the second parameter.
[0013] 8 is a block diagram showing the configuration of a control system according to an embodiment. FIG. 9 is a block diagram showing the configuration of the controller of FIG. 1. FIG. 10 is a flowchart showing an example of the flow of processing performed by the controller of FIG. 1. FIG. 11 is a diagram showing the configuration of a remote control system, which is an example of the control system of FIG. 1. FIG. 12 is a diagram showing an example of a target route when the moving body of FIG. 4 is a vehicle. FIG. 13 is a block diagram showing an example of the configuration of the moving body of FIG. 4. FIG. 14 is a diagram showing an example of a specific configuration when the moving body of FIG. 6 is a vehicle. FIG. 15 is a block diagram showing an example of the configuration of the server of FIG. 4. FIG. 16 is a diagram showing an example of time change in sample values of transmission delay measured by the delay calculation unit of FIG. 8. FIG. 17 is a diagram showing an example of the configuration of the control unit of FIG. 18. FIG. 19 is a diagram showing an example of the hardware configuration of the server of FIG. 19. FIG. 20 is a diagram showing a schematic diagram of a time interval referenced when the number of samples is 100. FIG. 21 is a diagram showing a simulation result of adaptive stochastic control according to embodiment 1. FIG. 22 is a diagram showing a simulation result of robust stochastic control according to embodiment 1. FIG. 23 is a diagram showing a simulation result of nominal stochastic control according to a comparative example.
[0014] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals, and the description thereof will not be repeated in principle. + , N 0 represent the set of real numbers, positive numbers, and non-negative integers, respectively. n , Rm×n represent a set of real vectors of degree n and a set of real matrices of m×n, respectively. n×n , S n×n + represent the set of n × n real symmetric matrices and positive definite matrices, respectively. n represents the identity matrix. n When the size n of is known, the subscript n is omitted. ||·|| represents the Euclidean norm. i (·) represents the i-th row of a matrix. The row expansion row(·) of a matrix with m rows is 1 (・), …, row m (·)]. E[·] represents the expected value. For a real square matrix M, He(M) is M+M T (T denotes transpose). diag(·) denotes a (block) diagonal matrix.
[0015] I. Overview of the embodiment Fig. 1 is a block diagram showing the configuration of a control system CS according to the embodiment. As shown in Fig. 1, the control system CS includes a controller 100 (control device) and a controlled object 400. The control system CS calculates the characteristics of the controlled object 400 (particularly the state x) from data obtained from time to time. k The time-varying parameter ξ k The time-varying parameter ξ is used to predict the probability distribution of the time-varying parameter ξ. k is a parameter that varies due to external or internal factors that are not controllable by the controller 100.
[0016] The controller 100 receives the manipulated variable u as an input to the controlled object 400. k (first parameter) is output. k is the state x of the control object 400 k Change the state x k is the time-varying parameter ξ k The control target 400 is also affected by the manipulated variable u k and state x k Control amount y according to k The controller 100 outputs the state x k , control amount y k, and the time-varying parameter ξ k Measure or estimate the manipulated variable u from these parameters. k+1 The state v of the controller 100 is calculated. k is the control variable y k It changes depending on the state x k If it is possible to directly measure the controlled variable y k is state x k The state x k and the time-varying parameter ξ k When it is not possible to directly measure each of the variables, they are calculated by using an ensemble Kalman filter or a particle filter. k , control amount y k and other available information.
[0017] The control system CS sequentially predicts the probability distribution at time k+1, one step ahead, using data up to the current time k, and reflects the prediction results in the control of the controlled object 400. In such sequential control, the state x k+1 State x k and the manipulated variable u k and the time-varying parameter ξ k and the control variable y k State x k and the manipulated variable u k and the time-varying parameter ξ k Among the relationships between these parameters, a general relationship that includes both nonlinear and linear systems can be expressed as in the following equation (1) (see Non-Patent Document 2).
[0018]
[0019] Equation (1) can be further generalized as a nonlinear system with a stochastic mapping, as in equation (2) below.
[0020]
[0021] In equation (2), π op,k and ρ op,k represents a time-varying conditional probability. The relationship between the two sides, "~", indicates that the left side of equation (2) follows the corresponding distribution.
[0022] State V k+1 and state v k and the control amount y k and the time-varying parameter ξ k Prediction distribution PD k+1 The feature value η k+1 and the control variable u k and state v k and feature η k+1 Since the general relationship between can include both nonlinear and linear systems, it can be expressed as in the following equation (3): Note that the predictive distribution is a future probability distribution predicted using a plurality of available sample values.
[0023]
[0024] Equation (3) can also be further generalized as a nonlinear system based on a stochastic mapping, as in the following equation (4).
[0025]
[0026] In equation (4), π K,k and ρ K,k represents a time-varying conditional probability. The relationship "~" on both sides indicates that the left side of equation (4) follows the corresponding distribution. The expressions of conditional probability in equations (2) and (4) have the same format as the expressions used to implement probabilistic policies in reinforcement learning. Note that, regardless of whether they are deterministic or probabilistic, the above equations (1) to (4) include the expression of the manipulated variable u sequentially by solving an optimization problem, as in model predictive control. k The expression of the controller 100 in such a manner as to determine the manipulated variable u k By regarding the algorithm that determines σ as a deterministic or stochastic function, it is possible to express the controller 100 using equations (1) to (4).
[0027] II. Control Theory II-A. Non-stationary Stochastic System In the following, a non-stationary stochastic system (see Non-Patent Document 1) of a linear system included in equations (1) to (4) will be explained with reference to equations (5) to (9). ξ = (ξ k ) (k∈N 0) is a time-independent stochastic process of order Z (Z is a natural number). k is a time-varying parameter that determines the probability measure, then the time-varying parameter ξ k The cumulative distribution function F(ξ k ;η k ) is calculated by the following equation (5): k A plurality of predetermined base distributions F (r) (ξ k ) is expressed as a convex combination (linear combination) of
[0028]
[0029] basis distribution F (r) (ξ k ) satisfies the definition of a cumulative distribution function. (r) (ξ k ) the support of the random vector Ξ (r) (⊂R Z ) is defined as the following equation (6).
[0030]
[0031] From the definition of equation (6), the time-varying parameter ξ k The support of is given by the set Ξ or a subset thereof. k Since is time-varying, the feature η k The corresponding distribution is also time-varying. Therefore, the time-varying parameter ξ k is generally a non-stationary stochastic process, not an i.i.d. (independent and identically distributed) process. Such a time-varying parameter ξ k Assuming a stochastic system that depends on state x k+1 and state x k and the manipulated variable u k The relationship between is expressed as the following equation (7).
[0032]
[0033] Status x k is expressed as an n-dimensional (n is a real number) vector. k The initial state x0 is assumed to be deterministic. k is expressed as an m-dimensional (m is a real number) vector. OP (ξ k ), B OP (ξ k ) is the state x k , operation amount u k The function A that determines the coefficients has dimensions that match each other. OP , B OP is a matrix-valued Borel measurable function whose domain is the base Ξ. k is expressed by the control gain K(η k ) and state x k That is, equation (8) is expressed as the product of the state x k is expressed as the manipulated variable u via the control gain K. k This represents the state feedback that is fed back to
[0034]
[0035] The control gain K is an m×n matrix, and is expressed by a plurality of predetermined base gains K as shown in the following equation (9): (r) It is expressed as a convex combination of
[0036]
[0037] Fig. 2 is a block diagram showing the configuration of the controller 100 shown in Fig. 1. As shown in Fig. 2, the controller 100 includes a distribution prediction unit 110, an extraction unit 120, and a control unit 130. The distribution prediction unit 110 calculates a time-varying parameter ξ corresponding to at least one time including the current time k. k From at least one sample value of k+1 Predict the predicted distribution PD k+1 to the extraction unit 120. Each of the sample values is a time-varying parameter ξ k Alternatively, the time-varying parameter ξ may be calculated using at least one of filters such as an ensemble Kalman filter or a particle filter. kThe sample values may be estimated from measurements of other parameters that have a dependency relationship with the distribution prediction unit 110. The sample values may be measured or estimated by the distribution prediction unit 110, or may be provided by a configuration different from the distribution prediction unit 110.
[0038] The extraction unit 120 extracts the prediction distribution PD k+1 From the feature value η k+1 and extract the feature η k+1 is output to the control unit 130. k+1 is a predictive distribution PD k+1 or the predictive distribution PD k+1 may be extracted from
[0039] The control unit 130 calculates the feature amount η k+1 and state x k Based on this, the manipulated variable u k is calculated, and the manipulated variable u k to the control target 400. Specifically, the control unit 130 outputs the feature amount η k+1 The manipulated variable u k and state x k By applying the predetermined relational expression (8) between k In addition, the state x k If it is not possible to directly measure the manipulated variable u, at least one of a filter such as an ensemble Kalman filter or a particle filter is used to measure the manipulated variable u. k and the controlled variable y k From state x k may be estimated.
[0040] Fig. 3 is a flowchart showing an example of the flow of processing performed by the controller 100 in Fig. 1. The processing shown in Fig. 3 is called for each step (sampling time) by a main routine (not shown) that comprehensively performs adaptive control on the control target 400. Below, a step will simply be abbreviated as S.
[0041] As shown in FIG. 3, in step S101, the controller 100 calculates a time-varying parameter ξ corresponding to at least one time. kAt step S102, the controller 100 obtains at least one sample value of the time-varying parameter ξ from the plurality of sample values. k Future predicted distribution PD k In S103, the controller 100 predicts the predicted distribution PD k From the feature value η k In S104, the controller 100 extracts the feature quantity η k and state x k Based on this, the manipulated variable u k In S105, the controller 100 calculates the manipulated variable u k is output to the control object 400, and the process returns to the main routine.
[0042] FIG. 4 is a diagram showing the configuration of a remote control system RCS, which is an example of the control system CS of FIG. 1. As shown in FIG. 4, the remote control system RCS includes a server 100A, a mobile object 400A, and routers RT1, RT2, RT3, and RT4. In the remote control system RCS, adaptive control (e.g., autonomous driving) of the mobile object 400A is performed by the server 100A. The mobile object 400A includes, for example, an automobile, a robot, and a drone. The server 100A and the mobile object 400A are examples of the controller 100 and the controlled object 400, respectively, of FIG. 1.
[0043] The routers RT1 to RT4 are connected to one another via a network NW. The mobile unit 400A is connected to the routers RT1 and RT2. The server 100A is connected to the routers RT3 and RT4. The operation amount u output from the server 100A is k is transmitted to the mobile unit 400A via the router RT4, the network NW, and the router RT1. k is transmitted to the server 100A via the router RT2, the network NW, and the router RT3. In the communication between the server 100A and the mobile unit 400A, a difference between the transmission time and the reception time (transmission delay) occurs. The probability distribution of the transmission delay changes over time. That is, the transmission delay is expressed by the time-varying parameter ξ in FIG.k Therefore, in the following, we will use the time-varying parameter ξ in Figure 1 as an example. k Here, the case where includes a transmission delay will be described. Note that the transmission delay may be at least any of the amount of transmission delay, the average value of the transmission delay in a predetermined time interval, the variance of the transmission delay, the maximum value of the transmission delay, and the minimum value of the transmission delay.
[0044] 5 is a diagram showing an example of a target route TR when the moving body 400A in FIG. 4 is a vehicle. As shown in FIG. 5, the target route TR is defined in an absolute coordinate system having an X-axis and a Y-axis that are orthogonal to each other. The moving body 400A has an x-axis and a y-axis that are orthogonal to each other. b axis and y b A relative coordinate system with an x axis is set. b The axis represents the traveling direction of the moving object 400A. The target route TR is defined by the y b The lateral position deviation x1 is the distance between the intersection point INP and the moving object 400A. The angular deviation x3 is the difference between the tangent line that is in contact with the target route TR at the intersection point INP and the x axis. b is the angle between the axis.
[0045] 4, when a target route TR is given, the server 100A calculates a control amount u of the moving object 400A so that the moving object 400A follows the target route TR. k If the target route TR includes a series of target speeds, the control amount u that controls the accelerator and brake to achieve the target speed is calculated. k Calculate the following.
[0046] Fig. 6 is a block diagram showing an example of the configuration of the moving body 400A of Fig. 4. As shown in Fig. 6, the moving body 400A includes an internal sensor 401, a command value calculation unit 402, an actuator 403, a receiving unit 404, a transmitting unit 405, and a time synchronization unit 406.
[0047] The internal sensor 401 is, for example, an IMU (Inertial Measurement Unit) sensor, a speed sensor, an acceleration sensor, a steering angle sensor, a steering torque sensor, or the like, and detects a control amount y kThe internal sensor 401 detects internal information including the following: The internal sensor 401 outputs the internal information to the router RT2 via the transmitting unit 405 as mobile object information.
[0048] The command value calculation unit 402 calculates the control amount u k via the receiving unit 404. The command value calculation unit 402 receives the control amount u k is converted into an actuator command value for the actuator 403. For example, the control amount u k is the target steering angle, the command value calculation unit 402 calculates the control variable u k into a current command value or a voltage command value for an electric power steering (EPS). The actuator 403 is composed of a motor and the like that actually operates the moving body 400A.
[0049] The time synchronization unit 406 synchronizes the timing of data transmission and reception in cooperation with a time synchronization unit in the server 100A. The time synchronization unit in the server 100A will be described later with reference to FIG.
[0050] Figure 7 is a diagram showing an example of a specific configuration when the moving body 400A in Figure 6 is a vehicle. A steering wheel 1 is installed so that a driver can operate the vehicle, and is engaged with a steering shaft 2. The steering shaft 2 is engaged with a pinion shaft 13 of a rack and pinion mechanism 4. A rack shaft 14 of the rack and pinion mechanism 4 is movable back and forth in response to the rotation of the pinion shaft 13, and front knuckles 6 are connected to both left and right ends of the rack shaft 14 via tie rods 5. The front knuckles 6 rotatably support front wheels 15 as steered wheels, and are supported on the vehicle body frame so as to be able to steer.
[0051] Torque generated by the driver operating the steering wheel 1 rotates the steering shaft 2. The rack and pinion mechanism 4 moves the rack shaft 14 left and right in response to the rotation of the steering shaft 2. The movement of the rack shaft 14 causes the front knuckle 6 to rotate about a kingpin axis (not shown), thereby steering the front wheels 15 left and right. The driver can change the amount of lateral movement of the vehicle by operating the steering wheel 1 when the vehicle moves forward and backward. Note that in the case of non-rider-type mobile objects such as fully autonomous vehicles and drones, components for driver operation such as a steering wheel are not required.
[0052] The moving body 400A is provided with internal sensors 401 for recognizing the moving state of the moving body 400A, such as a vehicle speed sensor 20, an IMU sensor 21, a steering angle sensor 22, and a steering torque sensor 23. The detected values of these sensors are output to a command value calculation unit 402. The command value calculation unit 402 outputs actuator command values to the acceleration / deceleration control device 9 and the steering control device 12.
[0053] The moving body 400A is equipped with actuators such as an electric motor 3 for realizing lateral movement of the moving body 400A, a vehicle drive device 7 for controlling longitudinal movement of the moving body 400A, and a brake control device 10. An acceleration / deceleration control device 9 controls the vehicle drive device 7 and the brake control device 10. A steering control device 12 controls the electric motor 3.
[0054] The electric motor 3 generally comprises a motor and a gear. The electric motor 3 applies torque to the steering shaft 2, thereby freely rotating the steering shaft 2. The electric motor 3 can freely steer the front wheels 15 independently of the driver's operation of the steering wheel.
[0055] The vehicle drive device 7 is an actuator for driving the moving body 400A in the forward and backward directions. The vehicle drive device 7 rotates the front wheels 15 and the rear wheels 16 using driving force obtained from a driving source such as an engine or a motor via a transmission and a shaft (not shown). This allows the vehicle drive device 7 to freely control the driving force of the moving body 400A.
[0056] On the other hand, the brake control device 10 is an actuator for braking the moving body 400A, and controls the braking amount of the brakes 11 installed on the front wheels 15 and rear wheels 16 of the moving body 400A. A typical brake generates a braking force by using hydraulic pressure to press pads against disc rotors that rotate together with the front wheels 15 and rear wheels 16.
[0057] The above-described internal sensors and other devices form a network using a controller area network (CAN) or a local area network (LAN) within the mobile object 400A. Each device within the mobile object 400A shown in FIG. 7 can acquire its respective information via the network. The internal sensors can also transmit and receive data to and from each other via the network. Even if the mobile object is not a vehicle, the mobile object 400A has a configuration corresponding to the actuators, internal sensors, command value calculation unit, etc.
[0058] Fig. 8 is a block diagram showing an example of the configuration of server 100A of Fig. 4. With reference to Fig. 8 and Fig. 2 together, server 100A includes a distribution prediction unit 110A, an extraction unit 120A, a control unit 130A, a delay calculation unit 140, a receiving unit 151, and a transmitting unit 152. Distribution prediction unit 110A, extraction unit 120A, and control unit 130A are examples of distribution prediction unit 110, extraction unit 120, and control unit 130 of Fig. 2, respectively.
[0059] The information from router RT3 is output to control section 130A and delay calculation section 140 via receiving section 151. Transmitting section 152 outputs the information received from control section 130A to router RT4.
[0060] The delay calculation unit 140 uses the information received from the receiving unit 151 to calculate the total delay consisting of the transmission delay occurring between the mobile body 400A and the server 100A and the calculation delay required for processing within the server 100A, and outputs a sample value of the delay to the distribution prediction unit 110A and the control unit 130A.
[0061] The distribution prediction unit 110A calculates a predicted distribution PD of the transmission delay from a plurality of sample values of the delay calculated at each sampling time. k+1 Predict the predicted distribution PD k+1 The extraction unit 120A outputs the prediction distribution PD k+1 From the feature value η k+1 and extract the feature η k+1 The control unit 130A outputs the feature amount η k+1 The manipulated variable u k and state x k By applying the predetermined relational expression (8) between k is calculated, and the manipulated variable u k to the transmitter 152. k is calculated by using at least one filter such as an ensemble Kalman filter or a particle filter to calculate the control variable u k and the controlled variable y k It is estimated from
[0062] FIG. 9 is a diagram showing an example of the change over time in the sampled values of the transmission delay measured by the delay calculation unit 140 of FIG. 8. In FIG. 9, time interval Pr1 represents the time interval from time t1 to t2 (t1<t2). Time interval Pr2 represents the time interval from time t2 to t3 (t2<t3). Time interval Pr3 represents the time interval from time t3 to t4 (t3<t4). Time interval Pr4 represents the time interval from time t4 to t5 (t4<t5). Time interval Pr5 represents the time interval from time t5 to t6 (t5<t6). Time interval Pr6 represents the time interval after time t6.
[0063] As shown in Figure 9, when a dedicated line or the like is not used, transmission delay generally does not take a constant value, but often takes a variable value. The transmission delays in time intervals Pr1, Pr3, Pr4, and Pr6 vary to the same extent. On the other hand, the variation in transmission delay in time intervals Pr2 and Pr5 is greater than the variation in transmission delay in time intervals Pr1, Pr3, Pr4, and Pr6. This phenomenon of differences in transmission delay variation depending on the time interval is caused by the fact that the probability distribution of transmission delay changes over time.
[0064] Fig. 10 is a diagram showing an example of the configuration of the control unit 130A of Fig. 8. As shown in Fig. 10, the control unit 130A includes a state estimating unit 131 and a controlled variable calculating unit 132. The state estimating unit 131 receives the time-varying parameter ξ from the delay calculating unit 140. k and obtains the sample value of the control amount u from the control amount calculation unit 132. k and acquires mobile object information, surrounding information, and map data from the receiving unit 151. The surrounding information includes, for example, the time, the state of radio waves around the mobile object, the presence or absence of structures around the mobile object, the distance between the mobile object and the structures, conductors and obstacles around the mobile object, radio wave strength, weather, and traffic. The map data includes, for example, the shape of roads around the mobile object, and the positions and shapes of surrounding structures. The state estimation unit 131 estimates the time-varying parameter ξ using at least one filter such as an ensemble Kalman filter or a particle filter. k sample value, control variable u k , state x from mobile object information, surrounding information, and map data k and estimate the state x k is output to the control amount calculation unit 132.
[0065] The control amount calculation unit 132 receives the feature amount η from the extraction unit 120A. k+1 and also acquires moving object information, surrounding information, and map data from the receiving unit 151. The control amount calculation unit 132 calculates the feature amount η k+1 The manipulated variable u k and state x k By applying the predetermined relational expression (8) between k is calculated, and the manipulated variable uk to the transmitting unit 152.
[0066] Fig. 11 is a diagram showing an example of the hardware configuration of the server 100A of Fig. 8. As shown in Fig. 11, the server 100A includes a processing circuit 51, a memory 52, an input unit 53, an output unit 54, and a communication unit 55. These components are connected to each other by a bus 50.
[0067] The processing circuitry 51 includes a processor such as a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), or a digital signal processor (DSP). The processing circuitry 51 executes programs stored in the memory 52 to realize the functions of the server 100A. The processing circuitry 51 may be configured as dedicated hardware. When the processing circuitry 51 is dedicated hardware, the processing circuitry 51 may be configured as, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination thereof.
[0068] The memory 52 includes, for example, non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), HDD (Hard Disk Drive), magnetic disk, flexible disk, optical disk, compact disk, mini disk, DVD (Digital Versatile Disc) and its drive device, or any storage medium to be used in the future.
[0069] The memory 52 stores a control program Pg for implementing adaptive control for the mobile unit 400A, sample values Ds of transmission delays measured at each sampling time, and control data Dc required for the adaptive control. The processing circuit 51 that executes the control program Pg implements the distribution prediction units 110A and 120A and the control unit 130A shown in Fig. 8. Although not shown in Fig. 11, the memory 52 may also store an OS (Operating System), data required for executing each program, and the processing results of each program.
[0070] The input unit 53 receives input from the user, such as GUI (Graphical User Interface) operations, gestures, CUI (Character User Interface) input, command input, or voice input. The input unit 53 includes a mouse, keyboard, touch panel, camera, and microphone. The output unit 54 outputs the processing results of the processing circuit 51 to the user. The output unit 54 includes, for example, a touch panel, display, and speaker. The communication unit 55 includes a NIC (Network Interface Card) and realizes data transmission and reception between the routers RT3 and RT4 and the processing circuit 51.
[0071] II-B. Stabilization Problem Given a stochastic system expressed by equation (7) and a controller 100 performing feedback control expressed by equation (8), the closed-loop system is expressed by the following equations (10) to (12).
[0072]
[0073] Given the state feedback expressed by equation (8), the sequence η = (η k ) (k∈N 0 ), under which a = a(η)∈R + When there exists λ=λ(η)∈(0,1) and the following equation (13) holds, the closed-loop system expressed by equation (10) is said to be second-order moment exponentially stable under the series.
[0074]
[0075] In the control by the controller 100, the feature quantity η k+1 On the other hand, at the time of advance (initial time k = 0), all future features η k The series η = (η k ) (k∈N 0 ) is difficult to predict in practice, so it is not assumed that the sequence is known. k It is necessary to ensure that the stability of the closed loop system is not compromised, regardless of the value of in the future. k is set E R Therefore, the series η is unknown, but belongs to the set ε R Belongs to.
[0076]
[0077] Any sequence η∈ε R If the closed-loop system is second-moment exponentially stable under R In the following, we will consider the feature η k Inequality conditions are derived to design a controller 100 that performs state feedback and can adjust the control method sequentially depending on the value of .
[0078] II-C. Lyapunov inequality time-varying parameter ξ k Prediction distribution PD k The feature η that determines k The time-varying parameter ξ k The probability integral for is defined as follows:
[0079]
[0080] Theorem 1: Given the state feedback shown in equation (8), under the sequence η, there exists a scalar λ∈(0,1) and a mapping P:E R →S n×n +exists and the following equation (16) holds, the closed-loop system shown in equation (10) is second-moment exponentially stable under that sequence (this can be proven using the result of Non-Patent Document 1).
[0081]
[0082] In equation (16), P(η k ) is a Lyapunov matrix, and equation (16) is a Lyapunov inequality. When the sequence η is known, the stability of a closed-loop system can be evaluated by searching for a scalar λ and a mapping P that satisfy the Lyapunov inequality. However, as described above, the assumption that the sequence η is known is not realistic, and the Lyapunov inequality (16) is an infinite simultaneous equation with respect to time k, so it is difficult to analyze stability using the Lyapunov inequality (16) as it is.
[0083] To address the difficulty of such stability analysis, we derive practical inequality conditions that enable us to evaluate the stability of a closed-loop system given the state feedback shown in equation (8). We then further extend these inequality conditions to derive inequality conditions for designing the state feedback shown in equation (8) so that the closed-loop system is robustly stable.
[0084] II-D. Inequality Conditions for Analysis Lemma 1: Under the state feedback sequence η shown in Equation (8), a scalar λ∈(0,1) and a mapping P:E R →S n×n + is given, the following conditions 1 and 2 are equivalent. Condition 1: Equation (16) holds.
[0085] Condition 2: Mapping T: Ξ × E R →S n×n exists, and the following equations (17) and (18) hold.
[0086]
[0087] Proof that if condition 1 is satisfied, then condition 2 is satisfied: Let T be the mapping T(·;η k+1 ) = A (·;η k+1 ) T P (ηk+1 ) A (·;η k+1 ), the formulas (17) and (18) are established.
[0088] Proof that if condition 2 is satisfied, then condition 1 is satisfied: ξ in equation (18) * Niξ k The validity of equation (16) can be shown from equation (17) and equation (18) obtained by substituting the above equations and taking the expected values of both sides of equation (18).
[0089] The mappings P and T in equations (17) and (18) are limited to convex combination mappings expressed by the following equations (19) and (20), respectively.
[0090]
[0091] Also, the end point F of the distribution function F (r) The probability integral by is defined as the following equation (21).
[0092]
[0093] Lemma 2: Given the state feedback and scalar λ∈(0,1) shown in equation (8), the following condition 4 is a sufficient condition for condition 3.
[0094] Condition 3: Mapping T: Ξ × E R →S n×n , mapping P:E R →S n×n + There exists such that for any η∈E R Equations (17) and (18) hold for
[0095] Condition 4: Mapping T (r) :Ξ→S n×n , P (r) ∈S n×n + (r=1,...,R), S∈R 2n×n exists, and the following equations (22) and (23) hold.
[0096]
[0097] Distribution function F (r) Since (r=1, . . . , R) is time-invariant, the following equation (24) holds.
[0098]
[0099] By substituting the left side of equation (24) into equation (22), the following equation (25) is derived.
[0100]
[0101] On both sides of equations (25) and (23), (r) k η (r+) k+1 Multiply by r, r + By calculating the sum for = 1, . . . , R, equation (17) and the following equation (26) are obtained based on equations (5), (19), and (20).
[0102]
[0103] For each of both sides of equation (26), the matrix [−I, A(ξ * ;η k+1 ) T ] and multiplying each of both sides of equation (26) from the right by the transpose of the matrix, equation (18) is obtained.
[0104] From Theorem 1 and Lemma 1 and 2, by the conditional expressions (22) and (23) that do not depend on the time k, any η∈ε R From the conditions (22) and (23), the mapping T (r) Eliminate and set any matrix H∈R n×n Theorem 2 can be obtained by introducing a deterministic matrix that satisfies all of the following equations (27), (28), and (29).
[0105]
[0106] Theorem 2: Given the state feedback shown in equation (8), and the scalar λ∈(0,1),P (r) ∈S n×n + (r=1,...r),S 1 , S 2 ∈R n×n exists and the following equations (30) and (31) hold, the closed loop system shown by equation (10) is RIt is robust second-moment exponentially stable with respect to
[0107]
[0108] Proof of Theorem 2: From equation (30), the lower right block of the left side of equation (31) is positive definite, and therefore, from Schur's lemma, equation (31) is equivalent to equation (32).
[0109]
[0110] By using equations (27) to (29), equation (32) is further equivalent to equation (33) below.
[0111]
[0112] As in Lemma 1, if equation (33) holds, and if there exists a mapping T that satisfies equation (22) and the following equation (34), (r) :Ξ→S n×n It is equivalent to the existence of (r=1, . . . , R).
[0113]
[0114] According to Schur's lemma, equation (34) is S = [S T 1 , S T 2 ] T is equivalent to equation (23) under the above condition. That is, Theorem 2 is derived from Theorem 1 and Lemmas 1 and 2. As shown in Theorem 2, if a deterministic matrix can be properly constructed, the robust stability of a closed-loop system can be analyzed by solving the standard forms of finite-dimensional matrix inequalities (30) and (31). The deterministic matrix can be constructed using only the distribution information of the end points of the A matrix of the closed-loop system shown in equation (12).
[0115] II-E. Inequality Conditions for Controller Design Now, by extending Theorem 2, we can see that for any η∈ε R The matrix inequality conditions for designing a controller 100 that performs state feedback as shown in equation (8) and stabilizes a closed-loop system under the following condition will be described. First, a matrix as shown in equation (35) below is set to satisfy equations (36A) and (36B).
[0116]
[0117] Since the right-hand side of equation (36A) is a positive semi-definite matrix, the matrix of equation (35) can be easily constructed through numerical calculations such as singular value decomposition. Using these matrices, matrices (r = 1, ..., R) as shown in the following equations (37), (38), and (39) are further calculated.
[0118]
[0119] When calculating the matrices shown in equations (37) to (39), these matrices can be calculated by direct calculation using an arbitrary matrix H∈R n×n It is shown that all of the following equations (40), (41), (42), (43), (44), and (45) are satisfied under
[0120]
[0121] However, it is defined as in the following equation (46).
[0122]
[0123] Equations (46), (37), and (38) (r=1, . . . , R) and the end point K of the control gain of the state feedback shown as equation (8) (r) Using (r=1, . . . , R), matrices are defined as shown in the following equations (47) and (48).
[0124]
[0125] Taking the following equation (49), from equations (40) to (45), the matrix (r, r + = 1, ..., R) is an arbitrary matrix H∈R n×n It is confirmed that the equations (27) to (29) are satisfied.
[0126]
[0127] By using equations (46), (37), and (38) (r=1, . . . , R), the following Theorem 3 is derived.
[0128] Theorem 3: A s ∈R n×nis a Schur stable matrix. (r) ∈S n×n + , N (r) ∈R m×n (r=1,...,R), V∈R n×n exists and the following equations (50) and (51) hold, the closed loop system shown in equation (10) is R The end point K of the control gain of the state feedback of Equation (8) that is robust second-order moment exponentially stable with respect to (r) ∈R m×n In particular, V is regular and such a gain exists for K (r) = N (r) V -1 (r=1, . . . , R).
[0129]
[0130] Proof of Theorem 3: Q(r) is positive definite, and V is regular according to equation (50). Multiply each of both sides of equation (51) from the right by the matrix of equation (52) below, and multiply each of both sides of equation (51) from the left by the transpose of that matrix to perform a congruent transformation. Then, perform a variable transformation G=V -1 , P (r) = G T Q (r) G.K. (r) = N (r) After applying G, the following equation (53) is obtained.
[0131]
[0132] Therefore, the matrix (r, r) shown in equations (47) and (48) + = 1, ..., R), the following equation (54) is set, and then the equation (31) holds.
[0133]
[0134] Regarding equation (50), V is added to the right of each side. -1 Multiply by V from the left on each side -1 By a congruence transformation that multiplies the transpose matrix of 2 When is set as in equation (54), equation (30) holds. Therefore, Theorem 3 is proved from Theorem 2.
[0135] The conditional expressions (50) and (51) of Theorem 3 are finite-dimensional matrix inequalities of the standard form, as in Theorem 2. In particular, if the scalar λ is fixed, the conditional expressions (50) and (51) become linear matrix inequalities, so if a solution to the conditional expressions (50) and (51) exists, it is possible to easily search for it by numerical calculation (note that the matrix A s (The feature η k The time-varying parameter ξ, which utilizes the value of k We can design a controller 100 as shown in equation (8) that provides stabilizing state feedback that is adaptive to changes in the probability distribution of .
[0136] As stated in Theorem 3, N (r) Since the control gain K (r) also depends on r. Therefore, N (r) = N 0 (r=1,...,R) 0 By fixing r and solving the conditional expressions (50) and (51), a control gain K that does not depend on r can be obtained. The controller 100 that realizes state feedback corresponding to a constant control gain K can be obtained by using the feature quantity η k Such a controller 100 does not use information on the time-varying parameter ξ k This realizes robust stabilization without any mechanism for adapting to changes in the probability distribution of the
[0137] III. Verification of the effectiveness of adapting to distribution changes. For the same numerical example, (r) When designing to minimize the scalar λ in the case where r is not shared (with adaptation) and in the case where it is shared (without adaptation), it was confirmed that the minimum value of the scalar λ is smaller in the case with adaptation than in the case without adaptation (i.e., control performance is improved). There were also numerical examples in which improvements were seen in the actual response when controlled. In addition, the available feature η kAs long as the value of is accurate, it is theoretically guaranteed that the design results with adaptation will not be worse than those without adaptation. Furthermore, the above numerical example is a counterexample, and it has been denied that there are only cases where performance does not improve.
[0138] IV. Sequential Distribution Prediction As explained in III, at time k, the feature η k+1 If the value of is available, the feature η k+1 By utilizing the information of η, it is expected that the control performance will be improved. k+1 Since it is difficult to directly know the value of at time k, the time-varying parameter ξ k From the data up to time k, the feature value η k+1 In an embodiment, the time-varying parameter ξ k From the data up to time k, the predictive distribution PD k+1 and construct a predictive distribution PD k+1 and the mixture distribution shown by equation (5) 2 The connection weight (coefficient) that minimizes the distance is the feature value η in the distribution one step ahead. k+1 The predicted value of η k+1|k That is, for each time k, the predicted value η k+1|k is the characteristic quantity η of the controller 100 shown in Equation (8). k+1 is assigned to.
[0139] IV-A. Mixture Distribution Approximation In the following, the prediction distribution PD k+1 Assume that a discrete distribution such as an empirical distribution is used for , and its cumulative distribution function F s The explanation will be limited to the case where one of the following is given. Since it is not possible to strictly define the probability density function of a discrete distribution, for convenience, the density function that introduces the concept of probability density function is called f s and the corresponding probability measure (empirical measure) is defined as in the following equation (55).
[0140]
[0141] In formula (55), N s is a positive number (number of samples). ζ is a vector-valued variable of order Z. w i∈R(i=0,...,N s -1) are probability weights that sum to 1 and correspond to the importance of each sample. δ represents the Dirac delta function. w i = 1 / N s (∀i = 0, ..., N s −1), the density function f s is N s This is a standard empirical distribution that treats all ζ samples equally. i is not a target for optimization, but is arbitrarily determined in advance. s, w i Both of these may be time-varying.
[0142] discrete distribution F s is approximated by the mixture distribution shown in equation (55). Note that here, the discrete distribution is fixed to one and the time concept is omitted, so for convenience of explanation, the same mixture distribution as equation (55) is used as the above ζ and θ∈E R The weight vector θ = [θ (1) , …, θ (R) ] T and is expressed as the following equation (56).
[0143]
[0144] Also, the end point F of the mixture distribution shown in Equation (56) (r) Assuming absolute continuity in f, the corresponding density function is (r) In this case, any θ∈E R Then F(·;θ) also becomes absolutely continuous, and the corresponding density function f can be defined as follows:
[0145]
[0146] In the mixture distribution approximation (approximation formula), the function f s θ is optimized so that the density function f(·;θ) of Equation (57) is as close as possible to (·). The weight vector θ optimized in this way is used as the feature quantity η k Corresponds to.
[0147] Function f sThe degree of similarity between the density function f(·;θ) and the density function f(·;θ) can be determined by various distances between them (e.g., Kullback-Leibler divergence, L 1 Distance (Manhattan distance) or L 2 It is possible to define the distance (Euclidean distance) as L 2 The case where distance is used will be explained. s (·) and the density function f(·;θ) 2 Distance d 2 is defined as the following equation (58).
[0148]
[0149] The root on the right side of equation (58) is expanded as in equation (59) below.
[0150]
[0151] The third term on the right side of equation (59) does not depend on θ, so L 2 This does not affect the minimization of the distance. The integral included in the first term on the right side of equation (59) is calculated as shown in equation (60) below.
[0152]
[0153] The integral included in the second term on the right side of equation (59) is calculated as shown in equation (61) below.
[0154]
[0155] If the first and second terms on the right side of equation (59) are equation J, then L 2 Distance d 2 The minimization problem of can be reduced to the minimization problem of equation J, as shown in the following equation (62). Note that the notation θ≧0 in equation (62) indicates that each component of the vector θ is non-negative.
[0156]
[0157] Equation (62) is a constrained quadratic programming problem. By solving this problem, the empirical distribution f sThe weight vector θ of the mixture distribution F having the shortest L2 distance between is derived. This problem is one of the classes of optimization problems used in real-time processing, such as model predictive control. That is, even in actual control problems, it is possible to solve a constrained quadratic programming problem such as that shown in equation (62) sufficiently quickly (for example, in about several tens of milliseconds even if it is solved directly). The solution θ of the quadratic programming problem k As the feature η k By designing the end point matrix of the control gain K in advance on the premise of deriving the distribution F, it is possible to move high-load processing such as the calculation of stochastic integrals from online processing to offline processing. By performing offline processing in advance, the calculation load of online processing is significantly reduced, and therefore online processing for each control cycle (step) is sped up. As a result, it is possible to improve the real-time performance of adaptive control. s Although the explanation was given assuming that is a discrete distribution, even if it is not a discrete distribution, L 2 A metric can be defined and an optimization problem can be formulated.
[0158] IV-B. Application to sequential distribution prediction When one discrete distribution is given, the connection weights that make the mixture distribution closest to the given discrete distribution can be found by solving the constrained quadratic programming problem of equation (62). In the following, we will apply this to find the predicted value η of the connection weights of the mixture distribution. k+1|k Specifically, at time k, the time-varying parameter ξ k Suppose that sample values of are obtained in some way, and the predicted value η is calculated based on the series data of sample values up to time k. k+1|k In the following, a case where an empirical distribution constructed using a most recent sample value sequence is used as the predictive distribution will be described as a first embodiment, and a case where a discrete distribution constructed using a forgetting factor is used as the predictive distribution will be described as a second embodiment.
[0159] [First embodiment] Number of samples N s Determine and fix one. The current time is k = κ (≧ N s−1), the predetermined time interval (horizon) K(κ):=[κ−N s +1, κ] ∩N 0 The time-varying parameter ξ in k The sample value (k∈K(κ)) of i The sample values (i = 0, ..., N s -1). Specifically, the following equation (58) holds true at the current time k.
[0160]
[0161] Also, w i = 1 / N s (i=0,...,N s −1). ζ as in equation (63) i The probability measure of the empirical distribution consisting of the sample values of and the probability weights wi is expressed as in the following equation (64) and depends on the time k.
[0162]
[0163] As shown in equation (64), the empirical distribution depends on time k (is time-varying), so the L 2 The connection weight of the mixture distribution that minimizes the distance also depends on time k. Let the connection weight that depends on time k be θ k The connection weight θ k The optimal value of is derived from equations (55) to (62) as a solution to the following quadratic programming problem:
[0164]
[0165] The matrix H in equation (65) is the same as the matrix H in equation (60). k is expressed as ζ of the coefficient b in Equation (61) j By substituting the sample values of , it is expressed as the following equation (66).
[0166]
[0167] The constrained quadratic programming problem expressed as in equation (65) is solved at each time k to obtain the optimal connection weight θ k Under the assumption that the distribution will not change significantly even one step ahead, the feature ηk+1|k = θ k The control signal is used by the controller 100 to control the controlled object 400.
[0168] In the online calculation for the control of the controlled object 400 by the controller 100, in addition to solving the equation (65), the coefficient b k The coefficient b must be recalculated using the new sample values. k The following equation (67) holds true for
[0169]
[0170] When using the recurrence formula shown in equation (67), the coefficient b ik Since the behavior of is unstable, if the calculation is not precise, errors may accumulate and the control of the control target 400 may become unstable. Therefore, in the first embodiment, the following update processes (P1), (P2), and (P3) are performed.
[0171]
[0172] FIG. 12 shows the number of samples N s is 100. As shown in FIG. 12 , when the current time is k=99, 100 sample values from k=0 to 99 included in the time interval going back from the current time to time k=0 are used to predict the empirical distribution. When the current time is k=100, 100 sample values from k=1 to 100 included in the time interval going back from the current time to time k=1 are used to predict the empirical distribution. When the current time is k=101, 100 sample values from k=2 to 101 included in the time interval going back from the current time to time k=2 are used to predict the empirical distribution.
[0173] According to the update processes (P1) to (P3), even if a calculation error occurs somewhere, the calculation error does not accumulate, and therefore the control target 400 can be stably controlled.
[0174] [Embodiment 2] In embodiment 2, a case will be described in which a discrete distribution is constructed so that the time interval going back from the current time is not clearly defined, and the older sample values among the multiple sample values are gradually forgotten (the usage rate is reduced) using a forgetting coefficient. The forgetting coefficient is expressed as γ∈(0,1) (for example, 0.9). In addition, when the number of samples N s0 and probability weight w 0i (i=0,...,N s0 -1), an empirical distribution is given as an initial distribution based on equation (64). That is, under appropriate sample values, the probability measure of the empirical distribution at the initial time k = 0 is given as the following equation (68). Note that the initial distribution may be constructed in any way. For example, N s0 = 1, w 00 = 1, and the initial distribution may be constructed from only a single sample value.
[0175]
[0176] A probability measure f with an initial value expressed as in equation (68) s (ζ, k) dζ is updated according to the following equation (69) when time k is 1 or later.
[0177]
[0178] As shown in equation (69), the distribution f s (ζ, k) is the distribution f s The term obtained by multiplying (ζ, k-1) by the forgetting factor γ and the time-varying parameter ξ at the current time k k The equation (69) is the sum of the delta function (specific function) of Dirac, which has a peak at the sample value of , and a term obtained by multiplying the value obtained by subtracting the forgetting factor γ from 1. The equation (69) is the distribution f of the time (k-1) one step before the current time k. s (ζ, k-1), the distribution f s The proportion remaining (remembered) in (ζ, k) is limited to the forgetting coefficient γ, and the other proportions (1-γ) are the empirical distribution f s (ζ, k) is lost (forgotten). Also, equation (69) expresses that the empirical distribution fs Empirical distribution f missing from (ζ, k) s The ratio of (ζ, k-1) is the time-varying parameter ξ k This indicates that the delta function (specific function) of Dirac, which has a peak at the sample value of (1-γ), is replaced by the ratio of (1-γ).
[0179] When equations (68) and (69) are interpreted according to the description of equation (64), the number of samples N s is N at time k s0 +k, so it is time-varying. 0i However, the probability weight w 0i (i=0,...,N s0-1 ) is 1, the sum of the probability weights when a probability measure is given as in equation (69) is also 1 regardless of time k. When a discrete distribution (i.e., a predictive distribution) is expressed using such a probability measure, the connection weights of the mixture distribution are derived by solving the constrained quadratic programming problem shown in equation (65) as in the first embodiment using a finite horizon. In this case, the matrix H in equation (65) corresponds to the matrix H shown in equation (60). Also, the coefficient b in equation (65) k is the coefficient consisting of k = 0 using the initial distribution, 0 and is configured as the following equation (70) for k≧1.
[0180]
[0181] Equation (70) is a recurrence formula similar to equation (67), but since γ<1, the coefficient b ik Therefore, the behavior of equation (70) is stable. k Similarly, the influence of the initial distribution also decays as the time k increases, so even if the initial distribution is not appropriate, problems caused by the inappropriateness of the initial distribution will no longer occur after a sufficient amount of time has passed. k θ derived from equation (65) under k is the feature η under the assumption that the distribution does not change significantly even one step ahead. k+1|k are used to control the control target 400.
[0182] For convenience of explanation, the probability measures shown in equations (68) and (69) are explicitly stated. However, in the second embodiment, past sample values or (i) It is not necessary to store the values, etc., to memory. In addition, since it is the equation (70) that needs to be calculated online to construct the equation (67), the data that needs to be stored in memory as information related to the sample value is the coefficient b k In this way, in the case based on the forgetting factor γ, even if the number of samples is unbounded with respect to time k, the required memory capacity can be limited.
[0183] Below, we will compare the stability of adaptive stochastic control according to embodiment 2, the stability of robust stochastic control, and the stability of nominal stochastic control according to a comparative example, using Figures 13 to 15. Figure 13 is a diagram showing simulation results for adaptive stochastic control according to embodiment 2. Figure 14 is a diagram showing simulation results for robust stochastic control. Figure 15 is a diagram showing simulation results for nominal stochastic control according to a comparative example. In nominal stochastic control, the control gain is designed assuming that the probability distribution of transmission delay is known. Each of Figures 13 to 15 shows the results of 100 simulations superimposed on each other.
[0184] 5, in each of FIGS. 13 to 15, parameter x1 is the lateral position deviation x1 in FIG. 5. Parameter x2 is the first-order time differential of parameter x1, and is the velocity of the moving object 400A relative to the intersection point INP. Parameter x3 is the angle deviation x3 in FIG. 5. Parameter x4 is the first-order time differential of parameter x3, and is the angular velocity of the moving object 400A relative to the tangent line that contacts the target route TR at the intersection point INP. Parameter x5 is the steering angle of the steering wheel. Each of parameters x1, x3, and x5 at time k is a function of the state x k Or the control amount y k This is an example of the parameter u at time k. c is the control amount u kThe initial value of the parameter x1 is 3m, and the initial values of the other parameters are 0.
[0185] As shown in Figure 15, in nominal stochastic control, the parameters do not converge over time. On the other hand, as shown in Figures 13 and 14, in both adaptive stochastic control and robust stochastic control, the parameters converge to 0 over time. The convergence time is approximately 10 to 13 seconds for adaptive stochastic control and approximately 15 to 18 seconds for robust stochastic control, making adaptive stochastic control shorter than robust stochastic control. In other words, adaptive stochastic control stabilizes more quickly than robust stochastic control. Oscillations in the transient response are also suppressed more effectively in adaptive stochastic control than in robust stochastic control.
[0186] In the approximation of the predictive distribution in the first and second embodiments, prior information on the predictive distribution is uncertain, or even a prior distribution is not required. Meanwhile, the calculation speed of the approximation of the predictive distribution is at the same level as that of MPC (Model Predictive Control), so the time required to approximate the predictive distribution is relatively short. Approximately 100 data points are sufficient for a control simulation. For example, if the time-varying parameter is a communication delay, it takes approximately 50 ms to acquire one time-varying parameter, so approximately 5 to 10 seconds is sufficient to acquire 100 time-varying parameters. When approximating the predictive distribution using 100 time-varying parameters from the present time, the predictive distribution can be updated for each time-varying parameter after the 100th time-varying parameter, and the control is also updated for each time-varying parameter. For example, if the controlled object is a car, the car can be allowed to idle for 10 seconds after the start of control before starting the car. After that, if there is a change in the predictive distribution, the control adapts to that change every 50 ms, allowing for rapid response to that change.
[0187] Note that the feature amount η at each time k is calculated by a method other than those in the first and second embodiments. k+1|k For example, at each time instant k, the time-varying parameter ξ kIf the sample values of are not directly available, the time-varying parameter ξ can be calculated using a particle filter or an ensemble Kalman filter. k A distribution of predicted values can be constructed from a set of particles that predict one step ahead and the weight of each particle. Using this distribution as a predictive distribution, the feature quantity η k+1|k Alternatively, the method using the filter may be combined with the first or second embodiment to determine the feature quantity η k+1|k may be determined.
[0188] V. Numerical Example of Control for Adjusting Control Gain Based on Sequential Distribution Prediction Using the same example as in III, the adaptive stochastic control according to the first embodiment was compared with the adaptive control according to the second embodiment. The time-varying parameter ξ k It was found that the time-varying parameter ξ quickly adapts to changes in the probability distribution, and good results are likely to be obtained in this example. On the other hand, it was confirmed that if the predetermined time interval can be made sufficiently long (if the time to acquire sample values can be made long enough), the prediction accuracy of the first embodiment may be higher than that of the second embodiment in a situation where no changes in the probability distribution occur. In both the first and second embodiments, the time-varying parameter ξ k There is a trade-off between speeding up adaptation to changes in the probability distribution of and improving prediction accuracy. However, this trade-off is s Alternatively, it can be changed by adjusting the forgetting coefficient γ, etc. Therefore, by appropriately adjusting these depending on the problem actually being handled, it is possible to realize control that is more suited to the problem in question.
[0189] As described above, the control device, control system, control program, and control method according to the first and second embodiments can realize stable adaptive control.
[0190] The embodiments disclosed herein are intended to be combined as appropriate within the scope of compatibility. The embodiments disclosed herein should be considered to be illustrative and not restrictive in all respects. The scope of the present disclosure is defined by the claims, not the above description, and is intended to include all modifications within the meaning and scope of the claims.
[0191] 1 Steering wheel, 2 Steering shaft, 3 Electric motor, 4 Rack and pinion mechanism, 5 Tie rod, 6 Front knuckle, 7 Vehicle drive device, 9 Acceleration / deceleration control device, 10 Brake control device, 11 Brake, 12 Steering control device, 13 Pinion shaft, 14 Rack shaft, 15 Front wheels, 16 Rear wheels, 20 Vehicle speed sensor, 21 IMU sensor, 22 Steering angle sensor, 23 Steering torque sensor, 50 Bus, 51 Processing circuit, 52 Memory, 53 Input unit, 54 Output unit, 55 Communication unit, 100 Controller, 100A Server, 110, 110A Distribution prediction unit, 120, 120A Extraction unit, 130, 130A Control unit, 131 State estimation unit, 132 Control amount calculation unit, 140 Delay calculation unit, 151, 404 Receiving unit, 152, 405 Transmitting unit, 406 Time synchronization unit, 400 Control object, 400A Moving body, 401 Internal sensor, 402 Command value calculation unit, 403 Actuator, CS Control system, Dc Control data, Ds Sampled value, NW Network, Pg Control program, Pr1 to Pr6 Time interval, RCS Remote control system, RT1 to RT4 Routers, TR Target route.
Claims
1. A control device that outputs a first parameter that changes the state of a controlled object to the controlled object, comprising: a distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each sample value corresponding to at least one time including the current time; an extraction unit that extracts features from the predictive distribution; and a control unit that calculates the first parameter based on the features and the state.
2. The control device according to claim 1, wherein the extraction unit derives an approximation formula for the predictive distribution and extracts the feature quantity from the approximation formula.
3. The control device according to claim 2, wherein the extraction unit extracts the feature by minimizing the distance between the approximation formula and the predictive distribution.
4. The control device described in claim 3, wherein the extraction unit derives a mixed distribution that is the sum of a plurality of predetermined base distributions as the approximation formula, derives coefficients of the plurality of base distributions as a solution to a quadratic programming problem that minimizes the distance, and extracts the coefficients of the plurality of base distributions as the feature quantities, the first parameter is expressed as the product of a control gain and the state, and the control gain is expressed as the sum of a plurality of predetermined base gains that each have a coefficient of the plurality of base distributions.
5. A control device according to any one of claims 1 to 4, wherein the at least one time is included in a time period preceding the current time by a predetermined time interval or later.
6. A control device according to any one of claims 1 to 5, wherein the predictive distribution at the current time is the sum of a term obtained by multiplying the predictive distribution at the time one step before the current time by a forgetting factor, and a term obtained by multiplying a specific function having a peak in the sample value of the second parameter at the current time by a value obtained by subtracting the forgetting factor from 1.
7. A control device according to any one of claims 1 to 6, wherein the controlled object and the control device are connected via a network in which a transmission delay occurs, and the second parameter includes the transmission delay.
8. A control system comprising: a controlled object; and a control device that outputs a first parameter that changes the state of the controlled object to the controlled object, wherein the control device comprises: a distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes a characteristic of the controlled object, each sample value corresponding to at least one time including the current time; an extraction unit that extracts a feature from the predictive distribution; and a control unit that calculates the first parameter based on the feature and the state.
9. A control program that, when executed by a processor, causes the processor to perform processing for outputting a first parameter that changes the state of a controlled object to the controlled object, the processing including: constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristic of the controlled object, each sample value corresponding to at least one time including the current time; extracting a feature from the predictive distribution; and calculating the first parameter based on the feature and the state.
10. A control method for performing processing to output a first parameter that changes the state of a controlled object to the controlled object, the control method comprising: a step of constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each sample value corresponding to at least one time including the current time; a step of extracting features from the predictive distribution; and a step of calculating the first parameter based on the features and the state.
Citation Information
Patent Citations
Remote control system, remote control method, remote control program, travel control device and moving body
JP2023008586A
Remote control device, remote control method, remote control system and mobile object
JP7330398B1