Control device, control system, control program, and control method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2025-07-28
- Publication Date
- 2026-08-05
AI Technical Summary
Existing adaptive control systems fail to detect and adapt to continuous or discrete changes in the probability distribution of parameters affecting controlled objects, leading to instability in remote control due to factors like transmission delay, and assume known distributions, preventing effective adaptive control.
A control device and system that includes a distribution prediction unit, extraction unit, and control unit to construct and analyze predicted distributions from sample values, extracting features to calculate parameters that adjust to changing distributions, ensuring stable adaptive control.
Stable adaptive control is achieved by calculating parameters based on features extracted from predictive distributions, enabling robust control even with unknown or changing probability distributions.
Smart Images

Figure 00000029_0000 
Figure 00000029_0001 
Figure 00000029_0002
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a control device, control system, control program, and control method that perform adaptive control to the temporal changes in the probability distribution of parameters that change the characteristics of a controlled object. [Background technology]
[0002] Conventionally, configurations are known that perform adaptive control (control that adapts to the temporal changes in the probability distribution of parameters that change the characteristics of the controlled object). For example, International Publication No. 2024 / 009418 (Patent Document 1) discloses a remote control device that estimates the probability distribution of transmission delay using a transmission delay model based on the mode of transmission delay between a mobile object and a remote control device. According to this remote control device, the target behavior of the mobile object can be changed according to the transmission delay mode, thereby improving the security of remote control of the mobile object via a network.
[0003] Furthermore, Non-Patent Document 1 derives the Lyapunov inequality, which characterizes the stability of linear stochastic systems without limiting the class of stochastic processes. Non-Patent Document 2 discloses a configuration relating to a nonlinear stochastic system. Non-Patent Document 3 derives a conditional equation that can evaluate H2 performance, which is an evaluation index for control performance using the H2 norm. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] International Publication No. 2024 / 009418 [Non-patent literature]
[0005] [Non-Patent Document 1] Y. Hosoe and T. Hagiwara, "On Second-Moment Stability of Discrete-Time Linear Systems With General Stochastic Dynamics", IEEE Transactions on Automatic Control, vol. 67, no. 2, pp. 795-809, 2022. [Non-Patent Document 2] Y. Kawano and Y. Hosoe, "Contraction Analysis of Discrete-Time Stochastic Systems", IEEE Transactions on Automatic Control, vol. 69, no. 2, pp. 982-997, 2024. [Non-Patent Document 3] Y. Hosoe, T. Okamoto, and T. Hagiwara, "H2 Performance Analysis and Synthesis for Discrete-Time Linear Systems With Dynamics Determined by an iid Process", IEEE Control Systems Letters, vol. 7, pp. 751-756, 2023. [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In Patent Document 1, the remote control device, which is the controller, outputs a control input calculated using a control gain based on information obtained from the mobile object, to the mobile object, which is the object to be controlled. In order to achieve stable adaptive control regarding the temporal changes in the probability distribution of parameters such as transmission delay, which change the characteristics of the object to be controlled, it is necessary to detect the continuous or discrete changes in the probability distribution and for the control gain to follow those changes. However, Patent Document 1 does not disclose a configuration for detecting such changes in the probability distribution of transmission delay, and the relationship between the change in the probability distribution and the control gain is not realized. Furthermore, in Non-Patent Documents 1 to 3, it is assumed that the distribution of parameters within the object to be controlled is known, so adaptive control to the temporal changes in the distribution is not realized.
[0007] This was undertaken to solve the aforementioned problems, and its purpose is to achieve stable adaptive control. [Means for solving the problem]
[0008] A control device according to one aspect of this disclosure outputs a first parameter to the controlled object that changes the state of the controlled object. The control device comprises a distribution prediction unit, an extraction unit, and a control unit. The distribution prediction unit constructs a predicted distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time. The extraction unit extracts features from the predicted distribution. The control unit calculates the first parameter based on the features and the state of the controlled object.
[0009] A control system relating to another aspect of this disclosure comprises a controlled object and a control device. The control device outputs a first parameter to the controlled object that changes the state of the controlled object. The control device includes a distribution prediction unit, an extraction unit, and a control unit. The distribution prediction unit constructs a predicted distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time. The extraction unit extracts features from the predicted distribution. The control unit calculates the first parameter based on the features and the state of the controlled object.
[0010] A control program relating to another aspect of this disclosure, when executed by a processor, causes the processor to perform a process to output a first parameter that changes the state of the controlled object to the controlled object. This process includes constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time; extracting features from the predictive distribution; and calculating the first parameter based on the features and the state of the controlled object.
[0011] A control method relating to another aspect of this disclosure is a control method for performing processing to output a first parameter that changes the state of the controlled object to the controlled object. The control method includes the steps of: constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time; extracting features from the predictive distribution; and calculating the first parameter based on the features and the state of the controlled object. [Effects of the Invention]
[0012] According to the control device, control system, control program, and control method disclosed herein, stable adaptive control can be achieved by calculating the first parameter using features extracted from the predictive distribution of the second parameter. [Brief explanation of the drawing]
[0013] [Figure 1] This is a block diagram showing the configuration of the control system according to the embodiment. [Figure 2] This is a block diagram showing the configuration of the controller in Figure 1. [Figure 3] Figure 1 shows a flowchart illustrating an example of the processing flow performed by the controller. [Figure 4] This figure shows the configuration of a remote control system, which is an example of the control system shown in Figure 1. [Figure 5]This figure shows an example of a target path when the moving object in Figure 4 is a vehicle. [Figure 6] Figure 4 is a block diagram showing an example of the configuration of a mobile body. [Figure 7] This figure shows an example of a specific configuration when the moving object in Figure 6 is a vehicle. [Figure 8] Figure 4 is a block diagram showing an example of a server configuration. [Figure 9] This figure shows an example of the time variation of sample values of transmission delay measured by the delay calculation unit in Figure 8. [Figure 10] This figure shows an example of the configuration of the control unit shown in Figure 8. [Figure 11] This figure shows an example of the server hardware configuration shown in Figure 8. [Figure 12] This diagram schematically shows the time intervals referenced when the sample size is 100. [Figure 13] This figure shows the simulation results of adaptive probability control according to Embodiment 1. [Figure 14] This figure shows the simulation results of robust stochastic control according to Embodiment 1. [Figure 15] This figure shows the simulation results of nominal stochastic control related to the comparative example. [Modes for carrying out the invention]
[0014] The embodiments of this disclosure will be described in detail below with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals, and their descriptions will not be repeated in principle. Also, R,R + ,N0 represents the sets of real numbers, positive numbers, and non-negative integers, respectively. n ,R m×n S represents an n-th order real vector and an m×n real sequence, respectively. n×n ,S n×n + These represent the sets of n×n real symmetric matrices and positive definite matrices, respectively. n This represents the identity matrix. Identity matrix I nWhen the size n is clear, the subscript n is omitted. ||·|| represents the Euclidean norm. row i (·) represents the i-th row of the matrix. The row expansion row(·) of an m-row matrix is [row1(·),…,row m (·)]. E[·] represents the expected value. For a real square matrix M, He(M) represents M + M T (T is the transpose). diag(·) represents a (block) diagonal matrix.
[0015] I. Overview of the Embodiment FIG. 1 is a block diagram showing the configuration of a control system CS according to an embodiment. As shown in FIG. 1, the control system CS includes a controller 100 (control device) and a controlled object 400. The control system CS predicts the probability distribution of a time-varying parameter ξ k that changes the dynamic characteristics (particularly the time transition of the state x k ) of the controlled object 400 from the data obtained at each moment, and controls the controlled object 400 so as to sequentially adapt to the prediction result. The time-varying parameter ξ k is a parameter that changes due to external or internal factors that cannot be controlled by the controller 100.
[0016] The controller 100 outputs an operation amount u k (first parameter) as an input to the controlled object 400. The operation amount u k changes the state x k of the controlled object 400. The state x k is also affected by the time-varying parameter ξ k (second parameter). The controlled object 400 outputs a control amount y k corresponding to the operation amount u k and the state x k . The controller 100 measures or estimates the state x k , the control amount y k , and the time-varying parameter ξ k , and calculates the operation amount u k+1 from these parameters. The state v k of the controller 100 changes according to the control amount y k . Note that the state xk If it is possible to directly measure the controlled quantity y k state x k It may be considered identical to state x. k and time-varying parameter ξ k If it is not possible to directly measure each of these, then they can be measured using an ensemble Kalman filter or particle filter, etc., with the manipulated quantity u k , controlled variable y k This may also be estimated using other available information.
[0017] The control system CS sequentially predicts the probability distribution at time k+1, one step ahead, using data up to the current time k, and reflects the prediction result in the control of the controlled object 400. In such sequential control, state x k+1 state x k and the control amount u k and the time-varying parameter ξ k It can be expressed using and the controlled variable y k state x k and the control amount u k and the time-varying parameter ξ k This can be expressed using and . Among the relationships between these parameters, the general relationship, including both nonlinear and linear systems, can be expressed as shown in equation (1) below (see Non-Patent Document 2).
[0018]
number
[0019] Equation (1) can be further generalized as a nonlinear system with a probabilistic mapping, as shown in equation (2) below.
[0020]
number
[0021] In equation (2), π op,k and ρ op,kThis represents a time-varying conditional probability. The relationship between the two sides, indicated by "~", shows that the left side of equation (2) follows the corresponding distribution.
[0022] state v k+1 and state v k and the controlled variable y k and the time-varying parameter ξ k Predicted distribution PD k+1 Feature η k+1 The general relationship between and the manipulated variable u k and state v k and feature η k+1 The general relationship between these two can include both nonlinear and linear systems, and can therefore be expressed as shown in equation (3) below. Note that the predictive distribution is the predicted probability distribution of the future using multiple available sample values.
[0023]
number
[0024] Equation (3) can also be further generalized as a nonlinear system with a probabilistic mapping, as shown in equation (4) below.
[0025]
number
[0026] In equation (4), π K,k and ρ K,k This represents a time-varying conditional probability. The relationship between the two sides "~" indicates that the left side of equation (4) follows the corresponding distribution. The expression of conditional probabilities in equations (2) and (4) has a similar form to the expression used to implement probabilistic policies in reinforcement learning. Note that equations (1) to (4) above, regardless of whether they are deterministic or probabilistic, can be used to solve optimization problems and sequentially apply the manipulated variable u, as in model predictive control. k This also includes a representation of a controller 100 that determines such a manipulated variable u. kBy considering the algorithm that determines as a deterministic or probabilistic function, it is possible to express the controller 100 using equations (1) to (4).
[0027] II. Control Theory II-A. Non-stationary stochastic systems In the following, the non-stationary stochastic systems of linear systems included in equations (1) to (4) (see Non-Patent Document 1) will be explained with reference to equations (5) to (9). ξ=(ξ k Let )(k∈N0) be a temporally independent stochastic process of order Z (where Z is a natural number). Feature η k If is a time-varying parameter that determines the probability measure, then the time-varying parameter ξ k The cumulative distribution function F(ξ k η k ) is expressed as shown in equation (5) below, the feature η k A set of predetermined basis distributions F, each with a coefficient representing one of its components. (r) (ξ k It is expressed as a convex combination (linear combination) of ).
[0028]
number
[0029] basis distribution F (r) (ξ k ) satisfies the definition of the cumulative distribution function. The cumulative distribution is F (r) (ξ k The base Ξ of the random vector given by ) (r) (⊂R Z ) is defined as shown in equation (6) below.
[0030]
number
[0031] From the definition of equation (6), the time-varying parameter ξ k The base is given by the set Ξ or a subset thereof. Feature η k Since it is time-varying, the feature η kThe corresponding distribution is also time-varying. Therefore, the time-varying parameter ξ with such a distribution k is generally a non-stationary stochastic process and not an i.i.d. (independent and identically distributed) process. Assuming a probability system dependent on such a time-varying parameter ξ k , the relationship between the state x k+1 and the state x k and the control variable u k is expressed as in the following equation (7).
[0032]
Number
[0033] The state x k is represented as an n-dimensional (n is a real number) vector. Assume that the initial state x0 of the state x k is deterministic. The control variable u k is represented as an m-dimensional (m is a real number) vector. The coefficients A OP (ξ k ), B OP (ξ k ) have dimensions that match the state x k and the control variable u k respectively. The functions A OP , B OP that determine the coefficients are matrix-valued Borel measurable functions with the domain Ξ. The control variable u k in equation (7) is expressed as the product of the control gain K(η k ) of the controller 100 and the state x k as in the following equation (8). That is, equation (8) represents state feedback where the state x k is fed back to the control variable u k through the control gain K.
[0034]
Number
[0035] The control gain K is an m×n matrix and is expressed as a convex combination of a plurality of predetermined basis gains K (r) as shown in the following equation (9).
[0036]
Equation
[0040] Figure 3 is a flowchart showing an example of the processing flow performed by the controller 100 in Figure 1. The processing shown in Figure 3 is called every step (sampling time) by a main routine (not shown) that comprehensively performs adaptive control on the controlled object 400. Hereafter, steps will simply be referred to as S.
[0041] As shown in Figure 3, in S101, the controller 100 sets a time-varying parameter ξ corresponding to at least one time point. k At least one sample value is obtained, and the process proceeds to S102. In S102, the controller 100 obtains the time-varying parameter ξ from multiple sample values. k Future predicted distribution PD k The controller predicts the distribution PD and proceeds to S103. In S103, the controller 100 predicts the distribution PD k From the features η k The data is extracted and the process proceeds to S104. In S104, the controller 100 processes the feature quantity η k and state x k Based on this, the manipulated variable u k The controller calculates the manipulated variable u and proceeds to S105. In S105, the controller 100 controls the manipulated variable u k The output is sent to the controlled device 400, and the processing is returned to the main routine.
[0042] Figure 4 shows the configuration of a remote control system RCS, which is an example of the control system CS in Figure 1. As shown in Figure 4, the remote control system RCS comprises a server 100A, a mobile device 400A, and routers RT1, RT2, RT3, and RT4. In the remote control system RCS, adaptive control (e.g., autonomous driving) is performed on the mobile device 400A by the server 100A. The mobile device 400A includes, for example, automobiles, robots, and drones. The server 100A and the mobile device 400A are examples of the controller 100 and controlled object 400 in Figure 1, respectively.
[0043] Routers RT1 through RT4 are each connected to each other via the network NW. Mobile unit 400A is connected to routers RT1 and RT2. Server 100A is connected to routers RT3 and RT4. The operation amount u output from server 100A k This is transmitted to the mobile unit 400A via router RT4, network NW, and router RT1. The control quantity y output from the mobile unit 400A is... k This is transmitted to server 100A via router RT2, network NW, and router RT3. A difference (transmission delay) occurs between the transmission time and the reception time in communication between server 100A and mobile device 400A. The probability distribution of this transmission delay changes over time. That is, the transmission delay is related to the time-varying parameter ξ in Figure 1. k This is one example. Therefore, below, we will discuss the time-varying parameter ξ in Figure 1. k The case where transmission delay is included will be explained. Note that the transmission delay may be represented by at least one of the following: the transmission delay amount, the average value of the transmission delay over a predetermined time interval, the variance of the transmission delay, the maximum value of the transmission delay, or the minimum value of the transmission delay.
[0044] Figure 5 shows an example of a target path TR when the moving object 400A in Figure 4 is a vehicle. As shown in Figure 5, the target path TR is defined in an absolute coordinate system having mutually orthogonal X and Y axes. The moving object 400A has mutually orthogonal x b Axis and y b A relative coordinate system with axes is set up. bThe axis represents the direction of travel of the moving object 400A. The target path TR is y at intersection INP. b It intersects with the axis. The lateral position deviation x1 is the distance between the intersection point INP and the moving object 400A. The angular deviation x3 is the distance between the tangent line that touches the target path TR at intersection point INP and x b It is the angle between the axis and the other axis.
[0045] Referring also to Figure 4, when a target path TR is given, the server 100A controls the amount u of the mobile body 400A to follow the target path TR. k Calculate (for example, steering angle). If the target path TR includes a series of target speeds, control quantities u control the accelerator and brakes to achieve that target speed. k Perform the calculation.
[0046] Figure 6 is a block diagram showing an example of the configuration of the mobile unit 400A in Figure 4. As shown in Figure 6, the mobile unit 400A includes an internal sensor 401, a command value calculation unit 402, an actuator 403, a receiving unit 404, a transmitting unit 405, and a time synchronization unit 406.
[0047] The internal sensor 401 controls the amount y by, for example, an IMU (Inertial Measurement Unit) sensor, a speed sensor, an acceleration sensor, a steering angle sensor, or a steering torque sensor. k The sensor detects internal information including the following. The internal sensor 401 outputs this internal information as mobile information to the router RT2 via the transmission unit 405.
[0048] The command value calculation unit 402 calculates the control quantity u calculated in the server 100A. k This is acquired via the receiving unit 404. The command value calculation unit 402 calculates the controlled quantity u k Convert the control variable u into an actuator command value for actuator 403. k If the target steering angle is u, the command value calculation unit 402 calculates the controlled amount u kThis is converted into a current command value or a voltage command value, etc., for the electric power steering (EPS). The actuator 403 consists of a motor, etc., that actually operates the moving body 400A.
[0049] The time synchronization unit 406 works in conjunction with the time synchronization unit in server 100A to synchronize the timing of data transmission and reception. The time synchronization unit in server 100A will be explained later using Figure 8.
[0050] Figure 7 shows an example of a specific configuration when the moving body 400A in Figure 6 is a vehicle. The steering wheel 1 is installed for the driver to operate the vehicle and is engaged with the steering shaft 2. The steering shaft 2 is engaged with the pinion shaft 13 of the rack and pinion mechanism 4. The rack shaft 14 of the rack and pinion mechanism 4 is reciprocally movable in accordance with the rotation of the pinion shaft 13, and front knuckles 6 are connected to both its left and right ends via tie rods 5. The front knuckles 6 rotatably support the front wheels 15 as steering wheels and are steerably supported by the vehicle frame.
[0051] The torque generated by the driver operating the steering wheel 1 rotates the steering shaft 2. The rack and pinion mechanism 4 moves the rack shaft 14 left and right in response to the rotation of the steering shaft 2. The movement of the rack shaft 14 causes the front knuckle 6 to rotate around a kingpin axis (not shown), thereby causing the front wheels 15 to steer left and right. The driver can change the amount of lateral movement of the vehicle by operating the steering wheel 1 when the vehicle moves forward and backward. In the case of non-passenger mobile devices such as fully autonomous vehicles and drones, components for driver operation such as a steering wheel are not required.
[0052] The mobile unit 400A is equipped with internal sensors 401 for recognizing the movement state of the mobile unit 400A, including a vehicle speed sensor 20, an IMU sensor 21, a steering angle sensor 22, and a steering torque sensor 23. The detected values from these sensors are output to the command value calculation unit 402. The command value calculation unit 402 outputs actuator command values to the acceleration / deceleration control device 9 and the steering control device 12.
[0053] The mobile unit 400A is equipped with actuators such as an electric motor 3 for lateral movement of the mobile unit 400A, a vehicle drive unit 7 for controlling longitudinal movement of the mobile unit 400A, and a brake control unit 10. The acceleration / deceleration control unit 9 controls the vehicle drive unit 7 and the brake control unit 10. The steering control unit 12 controls the electric motor 3.
[0054] The electric motor 3 is generally composed of a motor and gears. The electric motor 3 can rotate the steering shaft 2 freely by applying torque to it. The electric motor 3 can steer the front wheels 15 freely, independently of the driver's steering wheel operation.
[0055] The vehicle drive unit 7 is an actuator for driving the mobile body 400A in the forward and backward directions. The vehicle drive unit 7 rotates the front wheels 15 and rear wheels 16 via a transmission and shaft (not shown) using a driving force obtained from a drive source such as an engine or motor. This allows the vehicle drive unit 7 to freely control the driving force of the mobile body 400A.
[0056] On the other hand, the brake control device 10 is an actuator for braking the mobile body 400A, and controls the amount of braking force of the brakes 11 installed on the front wheels 15 and rear wheels 16 of the mobile body 400A. Conventional brakes generate braking force by using hydraulic pressure to press pads against disc rotors that rotate together with the front wheels 15 and rear wheels 16.
[0057] The aforementioned internal sensors and other devices form a network using CAN (Controller Area Network) or LAN (Local Area Network) within the mobile unit 400A. Each device within the mobile unit 400A shown in Figure 7 can acquire information from its respective source via this network. Furthermore, the internal sensors can send and receive data from each other via this network. Note that even when the mobile unit is not a vehicle, the system has a corresponding configuration of actuators, internal sensors, and command value calculation unit.
[0058] Figure 8 is a block diagram showing an example configuration of server 100A in Figure 4. Referring to both Figure 8 and Figure 2, server 100A comprises a distribution prediction unit 110A, an extraction unit 120A, a control unit 130A, a delay calculation unit 140, a receiving unit 151, and a transmitting unit 152. Distribution prediction unit 110A, extraction unit 120A, and control unit 130A are examples of the distribution prediction unit 110, extraction unit 120, and control unit 130 in Figure 2, respectively.
[0059] Information from router RT3 is output to the control unit 130A and delay calculation unit 140 via the receiving unit 151. The transmitting unit 152 outputs the information received from the control unit 130A to router RT4.
[0060] The delay calculation unit 140 uses the information received from the receiving unit 151 to calculate the total delay, which consists of the transmission delay occurring between the mobile unit 400A and the server 100A, and the calculation delay required for processing within the server 100A, and outputs a sample value of the said delay to the distribution prediction unit 110A and the control unit 130A.
[0061] The distribution prediction unit 110A predicts the transmission delay distribution PD from multiple sample values of delay calculated for each sampling time. k+1 Predict the predicted distribution PD k+1 The output is sent to the extraction unit 120A. The extraction unit 120A outputs the predicted distribution PD as the solution to the quadratic programming problem. k+1 From the features η k+1 Extract the feature η k+1is output to the control unit 130A. The control unit 130A uses the feature quantity η k+1 to apply it to a predetermined relational expression (8) between the manipulated variable u k and the state x k to calculate the manipulated variable u k and outputs the manipulated variable u k to the transmitter 152. The state x k is estimated using at least one of filters such as an ensemble Kalman filter or a particle filter, based on the control quantity u k and the control quantity y k .
[0062] FIG. 9 is a diagram showing an example of the time change of the sample values of the transmission delay measured by the delay calculation unit 140 in FIG. 8. In FIG. 9, the time interval Pr1 represents the time interval from time t1 to t2 (t1 < t2). The time interval Pr2 represents the time interval from time t2 to t3 (t2 < t3). The time interval Pr3 represents the time interval from time t3 to t4 (t3 < t4). The time interval Pr4 represents the time interval from time t4 to t5 (t4 < t5). The time interval Pr5 represents the time interval from time t5 to t6 (t5 < t6). The time interval Pr6 represents the time interval after time t6.
[0063] As shown in FIG. 9, when a dedicated line or the like is not used, the transmission delay generally rarely takes a constant value and often takes varying values. The transmission delays in the time intervals Pr1, Pr3, Pr4, and Pr6 vary in a similar manner to each other. On the other hand, the variation in the transmission delay in the time intervals Pr2 and Pr5 is larger than the variation in the transmission delay in the time intervals Pr1, Pr3, Pr4, and Pr6. The phenomenon in which the variation in the transmission delay differs depending on the time interval is caused by the time change of the probability distribution of the transmission delay.
[0064] FIG. 10 is a diagram showing an example of the configuration of the control unit 130A in FIG. 8. As shown in FIG. 10, the control unit 130A includes a state estimator 131 and a control quantity calculator 132. The state estimator 131 receives the time-varying parameter ξ from the delay calculation unit 140 kThe sample value is obtained, and the control variable u is calculated from the control variable calculation unit 132. k The system acquires mobile object information, surrounding information, and map data from the receiving unit 151. Surrounding information includes, for example, the time, the radio wave conditions around the mobile object, the presence or absence of structures around the mobile object and the distance of structures from the mobile object, conductors and obstacles around the mobile object, radio wave strength, weather, and traffic. Map data includes, for example, the road shape around the mobile object and the location and shape of surrounding structures. The state estimation unit 131 uses at least one filter, such as an ensemble Kalman filter or a particle filter, to determine the time-varying parameter ξ k Sample value, controlled variable u k State x from mobile object information, surrounding information, and map data k We estimate the state x k This is output to the control variable calculation unit 132.
[0065] The control variable calculation unit 132 calculates the feature quantity η from the extraction unit 120A. k+1 It acquires the feature quantity η, and also acquires mobile object information, surrounding information, and map data from the receiving unit 151. The control quantity calculation unit 132 calculates the feature quantity η k+1 The manipulated variable u k and state x k By applying this to the predetermined relational expression (8) between the two, the manipulated variable u k Calculate the manipulated variable u k This is output to the transmission unit 152.
[0066] Figure 11 shows an example of the hardware configuration of server 100A shown in Figure 8. As shown in Figure 11, server 100A comprises a processing circuit 51, memory 52, input unit 53, output unit 54, and communication unit 55. These components are connected to each other by a bus 50.
[0067] The processing circuit 51 includes, for example, a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), NPU (Neural Processing Unit), or DSP (Digital Signal Processor). The processing circuit 51 executes the program stored in the memory 52 and realizes the various functions of the server 100A. The processing circuit 51 may be configured as dedicated hardware. If the processing circuit 51 is dedicated hardware, it may be configured as, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.
[0068] Memory 52 includes, for example, non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), HDD (Hard Disk Drive), magnetic disk, flexible disk, optical disk, compact disk, minidisc, DVD (Digital Versatile Disc), and their drive devices, or any storage medium to be used in the future.
[0069] Memory 52 stores a control program Pg for implementing adaptive control of the mobile unit 400A, sample values Ds of transmission delay measured for each sampling time, and control data Dc necessary for the adaptive control. The processing circuit 51 that executes the control program Pg implements the distribution prediction units 110A, 120A, and control unit 130A shown in Figure 8. Although not shown in Figure 11, memory 52 may also store the OS (Operating System), data necessary for the execution of each program, and the processing results of each program.
[0070] The input unit 53 receives user input such as GUI (Graphical User Interface) operations, gestures, CUI (Character User Interface) input, command input, or voice input. The input unit 53 includes a mouse, keyboard, touch panel, camera, and microphone. The output unit 54 outputs the processing results of the processing circuit 51 to the user. The output unit 54 includes, for example, a touch panel, display, and speaker. The communication unit 55 includes a NIC (Network Interface Card) and enables data transmission and reception between routers RT3, RT4 and the processing circuit 51.
[0071] II-B. Stabilization problem Assuming a stochastic system represented by equation (7) and a controller 100 that performs feedback control represented by equation (8), the closed-loop system can be expressed as shown in equations (10) to (12) below.
[0072]
number
[0073] Given a state feedback represented by equation (8), the sequence η = (η k Under (k∈N0), a=a(η)∈R + And if there exists a λ=λ(η)∈(0,1) such that equation (13) holds, then the closed-loop system represented by equation (10) is said to be exponentially stable in the second moment under that series.
[0074]
number
[0075] In control by controller 100, the feature quantity η is calculated at each time step k. k+1 It is assumed that it is available. On the other hand, the feature η of all futures at the prior time (initial time k=0) k The sequence η = (η k Since predicting (k∈N0) is practically difficult, it is not assumed that the sequence is known. Therefore, the feature η k Regardless of what value it may take in the future, it is necessary to ensure that the stability of the closed-loop system is not compromised. From equation (9), the feature η at each time step k k is set E R Since it belongs to the set ε, although the value of the sequence η is unknown, it is represented by the set ε shown in equation (14) below. R It belongs to the category.
[0076]
number
[0077] Any sequence η∈ε R If a closed-loop system is exponentially stable under the conditions ε R This is referred to as robust second-moment exponential stability, or simply robust stability. Below, we will use the feature η such that the closed-loop system is robustly stable. k We derive inequality conditions for designing a controller 100 that provides state feedback, allowing the control method to be adjusted sequentially according to the value of .
[0078] II-C. Lyapunov Inequality Time-varying parameter ξ k Predicted distribution PD k The feature η that determines the feature k The time-varying parameter ξ that is explicitly defined. k The probability integral with respect to is defined as shown in equation (15) below.
[0079]
number
[0080] Theorem 1: Given the state feedback shown in equation (8), under the sequence η, there is a scalar λ∈(0,1) and a mapping P:E R →S n×n + If such that equation (16) holds, then the closed-loop system shown in equation (10) is second-moment exponentially stable under that sequence (this can be proven using the results of Non-Patent Document 1).
[0081]
number
[0082] In equation (16), P(η k ) is the Lyapunov matrix, and equation (16) is the Lyapunov inequality. If the sequence η is known, the stability of a closed-loop system can be evaluated by searching for a scalar λ and a mapping P that satisfy the Lyapunov inequality. However, as mentioned above, the assumption that the sequence η is known is not realistic, and since the Lyapunov inequality (16) can be solved infinitely simultaneously with respect to time k, it is difficult to analyze stability using the Lyapunov inequality (16) directly.
[0083] To address the difficulties in such stability analysis, we will derive practical inequality conditions that allow us to evaluate the stability of a closed-loop system, assuming that the state feedback shown in equation (8) is given. Subsequently, by further extending these inequality conditions, we will derive inequality conditions for designing the state feedback shown in equation (8) such that the closed-loop system becomes robustly stable.
[0084] II-D. Inequality conditions for analysis Lemma 1: State feedback shown in equation (8), scalar λ∈(0,1) under sequence η and mapping P:ER →S n×n + Given the conditions, the following conditions 1 and 2 are equivalent. Condition 1: Equation (16) holds.
[0085] Condition 2: Mapping T: Ξ×E R →S n×n There exists such that the following equations (17) and (18) hold.
[0086]
number
[0087] Proof that if condition 1 is satisfied, then condition 2 is also satisfied: Let the mapping T be T(·;η) k+1 )=A(·;η k+1 ) T P(η k+1 )A(·;η k+1 By constructing it as shown above, equations (17) and (18) hold true.
[0088] Proof that if condition 2 is satisfied, then condition 1 is also satisfied: ξ in equation (18) * ni ξ k By substituting this value and taking the expected value of both sides of equation (18), we can prove that equation (16) holds true from equation (17).
[0089] The mappings P and T in equations (17) and (18) are limited to the convex-joint mappings expressed in equations (19) and (20), respectively.
[0090]
number
[0091] Also, the endpoints F of the distribution function F (r) The stochastic integral due to is defined as shown in equation (21) below.
[0092]
number
[0093] Lemma 2: Given the state feedback shown in equation (8) and the scalar λ∈(0,1), the following condition 4 is a sufficient condition for condition 3.
[0094] Condition 3: Mapping T: Ξ × E R →S n×n , mapping P:E R →S n×n + There exists such that for any η ∈ E R For this, equations (17) and (18) hold.
[0095] Condition 4: Mapping T (r) :Ξ→S n×n ,P (r) ∈S n×n + (r=1,…,R), S∈R 2n×n There exists such that the following equations (22) and (23) hold.
[0096]
number
[0097] Distribution function F (r) Since (r=1,…,R) is time-invariant, the following equation (24) holds.
[0098]
number
[0099] By substituting the left-hand side of equation (24) into equation (22), the following equation (25) is derived.
[0100]
number
[0101] η (r) k η (r+) k+1 Multiply by r,r+ By calculating the sum for =1, ..., R, equation (17) and the following equation (26) are obtained based on equations (5), (19), and (20).
[0102]
number
[0103] For each side of equation (26), from left to right, we have the matrix [-I,A(ξ * η k+1 ) T By multiplying by ] and multiplying each side of equation (26) from the right by the transpose of the matrix, we obtain equation (18).
[0104] By Theorem 1 and Lemmas 1 and 2, the time-independent conditions (22) and (23) allow any η ∈ ε R The stability of the closed-loop system can be guaranteed under the following conditions: From equations (22) and (23), the mapping T (r) Eliminate any matrix H∈R n×n Theorem 2 can be obtained by introducing a deterministic matrix that satisfies all of the following equations (27), (28), and (29).
[0105]
number
[0106] Theorem 2: Given the state feedback shown in equation (8), scalar λ∈(0,1), P (r) ∈S n×n + (r=1,…r), S1,S2∈R n×n If such that equations (30) and (31) below hold, then the closed-loop system shown by equation (10) is ε R It is robust and exponentially stable in terms of the second moment.
[0107]
number
[0108] Proof of Theorem 2: From equation (30), the lower right block on the left side of equation (31) is positive definite, so by Schur's lemma, equation (31) is equivalent to equation (32).
[0109]
number
[0110] By using equations (27) through (29), equation (32) is further equivalent to equation (33).
[0111]
number
[0112] Similar to Lemma 1, equation (33) must hold, and the mapping T must satisfy equation (22) and the following equation (34). (r) :Ξ→S n×n The existence of (r=1,…,R) is equivalent.
[0113]
number
[0114] From Schur's lemma, equation (34) is S = [S T 1,S T 2] T Under this condition, equation (23) is equivalent. That is, Theorem 2 is derived from Theorem 1 and Lemmas 1 and 2. As shown in Theorem 2, if a deterministic matrix can be appropriately constructed, the robust stability of the closed-loop system can be analyzed by solving the standard forms of finite-dimensional matrix inequalities (30) and (31). The deterministic matrix can be constructed using only the distribution information of the endpoints of matrix A of the closed-loop system shown in equation (12).
[0115] II-E. Inequality conditions for controller design Below, by extending Theorem 2, for any η∈εR The matrix inequality conditions for designing a controller 100 that performs state feedback as shown in Equation (8) to stabilize the closed-loop system under R will be described. First, a matrix as shown in the following Equation (35) is set to satisfy Equations (36A) and (36B).
[0116]
Equation
[0117] Since the right side of Equation (36A) is a positive semi-definite matrix, such a matrix of Equation (35) can be easily constructed through numerical calculations such as singular value decomposition. Using these matrices, matrices (r = 1,..., R) as shown in the following Equations (37), (38), and (39) are further calculated.
[0118]
Equation
[0119] When calculating the matrices shown in Equations (37) to (39), it is shown that these matrices satisfy all of the following Equations (40), (41), (42), (43), (44), and (45) under an arbitrary matrix H ∈ R n×n However, it is defined as in the following Equation (46).
[0120]
Equation
[0121] That is, it is defined as in the following Equation (46).
[0122]
Equation
[0123] Equations (46), (37), (38) (r = 1,..., R), and the endpoint K of the control gain of the state feedback shown as Equation (8) (r)Using (r=1,…,R), the matrix is defined as shown in equations (47) and (48) below.
[0124]
number
[0125] Taking the following equation (49), from equations (40) to (45), the matrix (r,r) shown in equations (47) and (48) + (=1,…,R) is any matrix H∈R n×n It can be confirmed that equations (27) to (29) are satisfied.
[0126]
number
[0127] Using equations (46), (37), and (38) (r=1,…,R), theorem 3 can be derived as follows.
[0128] Theorem 3:A s ∈R n×n Let be a Schur-stable matrix. Scalar λ∈(0,1),Q (r) ∈S n×n + ,N (r) ∈R m×n (r=1,…,R),V∈R n×n If such that exists and equations (50) and (51) below hold, then the closed-loop system shown in equation (10) is ε R The endpoint K of the state feedback control gain in equation (8) is such that it is robust and exponentially stable with respect to the second moment of the moment. (r) ∈R m×n There exists a (r=1,…,R) such that V is holomorphic, and such a gain is K (r) =N (r) V -1 It is given by (r=1,…,R).
[0129]
number
[0130] Proof of Theorem 3: Since Q(r) is positive definite and from Equation (50), V is regular. Multiply each side of Equation (51) from the right by the matrix of the following Equation (52), multiply each side of Equation (51) from the left by the transpose matrix of the said matrix for congruence transformation, and further perform a variable transformation G = V -1 , P (r) = G T Q (r) G, K (r) = N (r) When G is applied, the following Equation (53) is obtained.
[0131]
Number
[0132] Therefore, under the matrices (r, r + = 1,..., R) shown in Equations (47) and (48), Equation (31) holds when set as the following Equation (54).
[0133]
Number
[0135] The conditional equations (50) and (51) of Theorem 3 are in the form of standard finite-dimensional matrix inequalities similar to the case of Theorem 2. In particular, if the scalar λ is fixed, the conditional equations (50) and (51) become linear matrix inequalities. Therefore, when solutions to the conditional equations (50) and (51) exist, they can be easily searched by numerical calculation (note that the matrix A s may be fixed as a zero matrix, etc.). Through the search for such solutions, the value of the feature η k is utilized, and the time-varying parameter ξ kA controller 100, as shown in equation (8), can be designed to implement a stabilization state feedback that can adapt to changes in the probability distribution.
[0136] Furthermore, as described in Theorem 3, N (r) Because it is r-dependent, the designed control gain K (r) It also depends on r. Therefore, N (r) By fixing a common variable N0, such as =N0(r=1,…,R), and solving the conditional equations (50) and (51), a control gain K that does not depend on r can be obtained. The controller 100 that realizes state feedback corresponding to the constant control gain K has a feature η k This information is not used. Such a controller 100 uses the time-varying parameter ξ k This achieves robust stabilization without an adaptation mechanism to changes in the probability distribution.
[0137] III. Verification of the effectiveness of adapting to changes in distribution For the same numerical example, N (r) When designing to minimize scalar λ with respect to r, both with and without adaptation (where r is not common), it was confirmed that the minimum value of scalar λ is smaller with adaptation than without adaptation (i.e., control performance is improved). There were also numerical examples showing improvement in the actual response when controlled. Note that the available feature quantity is η. k As long as the value is accurate, it is theoretically guaranteed that the design results will not be worse with the adaptation compared to without the adaptation. Furthermore, the possibility that the above numerical examples serve as counterexamples and that only cases where performance does not improve exist has been ruled out.
[0138] IV. Sequential Distribution Prediction As explained in III, the feature η at time k k+1 If the value of is available, feature η k+1 By utilizing this information, improvements in control performance can be expected. However, in reality, the feature η k+1 Since it is difficult to directly know the value of at time k, the time-varying parameter ξ kFrom the data up to time k, the feature η k+1 The value of is predicted. In this embodiment, the time-varying parameter ξ k Predicted distribution PD from data up to time k k+1 It constitutes the predictive distribution PD k+1 The connection weight (coefficient) that minimizes the L2 distance between the given distribution and the mixture distribution shown by equation (5) is the feature η in the distribution one step ahead. k+1 Predicted value η k+1|k Let's assume that the predicted value η is obtained at each time k. k+1|k The characteristic quantity η of the controller 100 shown in equation (8) is k+1 It is substituted into.
[0139] IV-A. Mixture distribution approximation Below, we will discuss the predicted distribution PD. k+1 Assuming that a discrete distribution such as an empirical distribution is used, the cumulative distribution function F s We will explain the situation by limiting it to the case where one of the following is given. Since the probability density function of a discrete distribution cannot be strictly defined, we will use the density function in which the concept of probability density function is introduced for convenience as f s We denote this as and define the corresponding probability measure (empirical measure) as shown in equation (55) below.
[0140]
number
[0141] In equation (55), N s w is a positive number (sample size). ζ is a vector-valued variable of degree Z. i ∈R(i=0,…,N s -1) represents the probability weights, which sum to 1, and correspond to the importance of each sample. δ represents the Dirac delta function. i = 1 / N s (∀i=0,…,N s By setting -1), the density function f in equation (55) s is, N s This is a standard empirical distribution that treats each sample of ζ equally. Note that the probability weights w are also considered. iThis is not subject to optimization, but is determined arbitrarily in advance. Also, although the concept of time is omitted here, in sequential prediction, N s, w i Any of these may be time-varying.
[0142] discrete distribution F s This is approximated by the mixture distribution shown in equation (55). Note that here we have fixed the discrete distribution to one and omitted the concept of time, so for the sake of explanation, the same mixture distribution as in equation (55) is used for ζ and θ∈E above. R Weight vector θ = [θ] that satisfies this condition (1) ,…,θ (R) ] T Using and , it can be expressed as follows in equation (56).
[0143]
number
[0144] Furthermore, the endpoint F of the mixed distribution shown in equation (56) (r) Assuming absolute continuity, the corresponding density function is f (r) This is expressed as follows. In this case, any θ∈E R Thus, F(·;θ) is also absolutely continuous, and the corresponding density function f can be defined as shown in equation (57) below.
[0145]
number
[0146] In the mixture distribution approximation (approximation formula), the function f s θ is optimized so that the density function f(·;θ) in equation (57) is as close as possible to (·). The weight vector θ thus optimized becomes the feature η k It corresponds to.
[0147] function f sThe degree to which the density function f(·;θ) is close (similar) to (·) can be defined by various distances between them (for example, Kullback-Leibler divergence, L1 distance (Manhattan distance), or L2 distance (Euclidean distance)). Below, we will explain the case where L2 distance is used as the distance. Function f s The L2 distance d2 between (·) and the density function f(·;θ) is defined by the following equation (58).
[0148]
number
[0149] The expression inside the square root on the right-hand side of equation (58) is expanded as shown in equation (59) below.
[0150]
number
[0151] The third term on the right-hand side of equation (59) does not depend on θ and therefore does not affect the minimization of the L2 distance. The integral included in the first term on the right-hand side of equation (59) is calculated as shown in equation (60) below.
[0152]
number
[0153] The integral included in the second term on the right-hand side of equation (59) is calculated as shown in equation (61) below.
[0154]
number
[0155] If we denote the first and second terms on the right-hand side of equation (59) as equation J, then the problem of minimizing the L2 distance d2 can be reduced to the problem of minimizing equation J, as shown in equation (62) below. Note that the notation θ≧0 in equation (62) indicates that each component of the vector θ is non-negative.
[0156]
number
[0157] Equation (62) is a constrained quadratic programming problem. By solving this problem, the empirical distribution f s The weight vector θ of the mixture distribution F with the closest L2 distance between it and is derived. This problem is one of the classes of optimization problems used in real-time processing, such as model predictive control. In other words, even in actual control problems, it is possible to solve constrained quadratic programming problems like the one shown in equation (62) quite quickly (even if solved directly, it would take, for example, several tens of milliseconds). Solution θ of the quadratic programming problem k Feature η k By designing the endpoint matrix of the control gain K in advance, assuming that the following will be derived, computationally intensive processes such as the calculation of stochastic integrals can be moved from online processing to offline processing. Performing offline processing in advance significantly reduces the computational load of online processing, thus speeding up online processing for each control cycle (step). As a result, the real-time performance in adaptive control can be improved. Note that the distribution F is used here. s Although the explanation assumed that the distribution was discrete, the L2 distance can still be defined and the optimization problem can be formulated even if the distribution is not discrete.
[0158] IV-B. Applications to successive distribution prediction Given a discrete distribution, the bond weights that make the mixture distribution closest to that discrete distribution can be found by solving the constrained quadratic programming problem in equation (62). Below, we apply this to predict the bond weights η of the mixture distribution. k+1|kWe will describe a configuration that sequentially calculates the time-varying parameter ξ at time k. Specifically, at time k, k Assuming that sample values are obtained by some method, the predicted value η is based on the series of sample values up to time k. k+1|k Let's consider the case where we decide on the prediction distribution. Below, we will describe the case where the prediction distribution is an empirical distribution constructed from the most recent sample value series as Embodiment 1, and the case where the prediction distribution is a discrete distribution constructed using the forgetting coefficient as Embodiment 2.
[0159] [Embodiment 1] Sample size N s Choose and fix one value. The current time is k = κ (≧N s -1) When the predetermined time interval (horizon) K(κ):=[κ-N s Time-varying parameter ξ in [+1,κ]∩N0 k The sample value (k∈K(κ)) is given by ζ in equation (55) i Sample values (i=0,…,N s This corresponds to -1). Specifically, at the current time k, the following equation (58) holds true.
[0160]
number
[0161] Also, w i = 1 / N s (i=0,…,N s Let -1). Equation (63) i The probability measure of the empirical distribution, which consists of sample values and probability weights wi, is expressed by the following equation (64) and depends on time k.
[0162]
number
[0163] As shown in equation (64), since the empirical distribution is time-dependent (time-varying) due to time k, the combined weights of the mixture distribution that minimize the L2 distance between it and the empirical distribution also depend on time k. The time-dependent combined weights are given by θ. k This is expressed as follows: Bond weight θ k The optimal value of is derived from equations (55) to (62) as the solution to the following quadratic programming problem.
[0164]
number
[0165] The matrix H in equation (65) is the same as the matrix H in equation (60). The coefficient b in equation (65) k This is the coefficient b of equation (61) of equation (63) = ζ j By substituting this into the sample values, it can be expressed as shown in equation (66) below.
[0166]
number
[0167] Solving the constrained quadratic programming problem expressed as in equation (65) at each time k to find the optimal bond weights θ k This is required. Under the assumption that the distribution does not change significantly even one step ahead, the feature η k+1|k =θ k It is used for controlling the controlled object 400 by the controller 100.
[0168] In the online calculation for controlling the controlled object 400 by the controller 100, in addition to solving equation (65), the coefficient b of equation (66) is also calculated. k The coefficient b needs to be recalculated using new sample values. k The following equation (67) holds true for this.
[0169]
number
[0170] When using the recurrence relation shown in equation (67), the coefficient b ik Because its behavior is unstable, if the calculations are not precise, errors can accumulate and the control of the controlled object 400 may become unstable. Therefore, in Embodiment 1, the following update processes (P1), (P2), and (P3) are performed.
[0171]
number
[0172] Figure 12 shows the sample size N. s This diagram schematically shows the time intervals referenced when k is 100. As shown in Figure 12, when the current time is k=99, the 100 sample values from k=0 to 99 included in the time interval going back from the current time to time k=0 are used to predict the empirical distribution. When the current time is k=100, the 100 sample values from k=1 to 100 included in the time interval going back from the current time to time k=1 are used to predict the empirical distribution. When the current time is k=101, the 100 sample values from k=2 to 101 included in the time interval going back from the current time to time k=2 are used to predict the empirical distribution.
[0173] According to the update process (P1)~(P3), even if a calculation error occurs somewhere, the calculation error will not accumulate. Therefore, stable control can be performed on the controlled object 400.
[0174] [Embodiment 2] In Embodiment 2, we describe a case where the time interval going back from the current time is not clearly defined, and the discrete distribution is constructed so that older sample values are gradually forgotten (the usage rate decreases) among multiple sample values using a forgetting coefficient. The forgetting coefficient is expressed as γ∈(0,1) (for example, 0.9). Also, the sample size N s0 and probability weight w 0i (i=0,…,N s0Under (-1), assume that an empirical distribution is given as the initial distribution based on equation (64). That is, under appropriate sample values, assume that the probability measure of the empirical distribution at initial time k=0 is given by the following equation (68). Note that the initial distribution can be constructed in any way. For example, N s0 =1,w 00 The initial distribution may be constructed using only a single sample value, by setting = 1.
[0175]
number
[0176] A probability measure f has an initial value expressed as shown in equation (68). s (ζ,k)dζ is updated according to the following equation (69) when time k is 1 or later.
[0177]
number
[0178] As shown in equation (69), the distribution of the current time k is f s (ζ,k) is the distribution f of time (k-1) one step before the current time k. s The term obtained by multiplying (ζ,k-1) by the forgetting coefficient γ, and the time-varying parameter ξ of the current time k. k It is the sum of a term obtained by multiplying the Dirac delta function (specific function) which has a peak in the sample values by a value obtained by subtracting the forgetting coefficient γ from 1. Equation (69) is the distribution f at time (k-1) one step before the current time k. s Distribution f of the current time k within (ζ,k-1) s The proportion that remains (is remembered) at (ζ,k) is limited to the forgetting coefficient γ, and the proportion of the remaining (1-γ) is the empirical distribution f at the current time k. s This represents the loss (forgetting) from (ζ,k). Also, equation (69) shows the empirical distribution f at the current time k. s The empirical distribution f lost from (ζ,k) s The ratio of (ζ,k-1) is the time-varying parameter ξ at the current time k.k This represents the proportion of (1-γ) of the Dirac delta function (specific function) that has a peak in the sample values.
[0179] Interpreting equations (68) and (69) according to the description in equation (64), the sample size N is s At time k, N s0 It becomes +k, so it is time-varying. Probability weight w 0i However, since the information is forgotten at a rate equal to the forgetting rate γ, it is not uniform across the sample. 0i (i=0,…,N s0-1 Since the sum of ) is 1, the sum of the probability weights when a probability measure is given as in equation (69) is also 1 regardless of time k. When a discrete distribution (i.e., a predictive distribution) is represented by such a probability measure, the combined weights of the mixture distribution are derived by solving the constrained quadratic programming problem shown in equation (65), similar to Embodiment 1 using a finite horizon. In this case, the matrix H in equation (65) corresponds to the matrix H shown in equation (60). Also, the coefficient b in equation (65) k The coefficients are constructed using the initial distribution, with the initial value b0 being the coefficients constructed when k=0, and for k≧1, the following equation (70) is used.
[0180]
number
[0181] Equation (70) is a recurrence relation, similar to equation (67), but since γ < 1, the coefficient b ik The behavior is stable. Therefore, equation (70) is given by coefficient b k Even if used for updating, errors will not continue to accumulate. Similarly, the effect of the initial distribution also decays as time k increases, so even if the initial distribution is not appropriate, problems caused by the inappropriateness of the initial distribution will no longer occur after a sufficient amount of time has passed. Such a coefficient b k The θ derived from equation (65) below k This assumes that the distribution does not change significantly even one step ahead, and the feature η k+1|k It is used for controlling the controlled object 400.
[0182] For the sake of explanation, the probability measures shown in equations (68) and (69) are explicitly shown, but in Embodiment 2, past sample values or f (i) The values obtained by applying this do not need to be stored in memory. Also, since equation (70) is the one that needs to be calculated online in order to construct equation (67), the data that needs to be stored in memory as information related to the sample value is the coefficient b from the previous step. k This is the value of γ. In this way, when based on the forgetting coefficient γ, the required memory capacity can be limited even if the number of samples is unbounded with respect to time k.
[0183] In the following sections, we will compare the stability of adaptive stochastic control according to Embodiment 2, the stability of robust stochastic control, and the stability of nominal stochastic control according to the comparative example using Figures 13 to 15. Figure 13 shows the simulation results of adaptive stochastic control according to Embodiment 2. Figure 14 shows the simulation results of robust stochastic control. Figure 15 shows the simulation results of nominal stochastic control according to the comparative example. In nominal stochastic control, the control gain is designed assuming that the probability distribution of transmission delay is known. In each of Figures 13 to 15, the results of 100 simulations are superimposed.
[0184] Referring also to Figure 5, in each of Figures 13 to 15, parameter x1 is the lateral position deviation x1 in Figure 5. Parameter x2 is the first time derivative of parameter x1, and is the velocity of the moving object 400A with respect to intersection INP. Parameter x3 is the angular deviation x3 in Figure 5. Parameter x4 is the first time derivative of parameter x3, and is the angular velocity of the moving object 400A with respect to the tangent line that touches the target path TR at intersection INP. Parameter x5 is the steering angle of the handle. At time k, each of parameters x1, x3, and x5 is in state x k Or the controlled variable y k This is an example. The parameter u at time k c This is the control amount u, which is the target steering angle. kThis is the voltage command value for the electric power steering, converted from [the original value]. The initial value of parameter x1 is 3m, and the initial values of the other parameters are 0.
[0185] As shown in Figure 15, in nominal stochastic control, the parameters do not converge over time. On the other hand, as shown in Figures 13 and 14, in both adaptive stochastic control and robust stochastic control, the parameters converge to 0 over time. Regarding the time to convergence, adaptive stochastic control takes about 10 to 13 seconds, while robust stochastic control takes about 15 to 18 seconds, meaning that adaptive stochastic control is faster than robust stochastic control. In other words, adaptive stochastic control stabilizes faster than robust stochastic control. Oscillation in the transient response is also suppressed more effectively in adaptive stochastic control than in robust stochastic control.
[0186] In the approximation of the predictive distribution in Embodiments 1 and 2, prior information about the predictive distribution is uncertain, or no prior distribution is required. On the other hand, the calculation speed of the predictive distribution approximation is at the same level as MPC (Model Predictive Control), so the time required for predictive distribution approximation is relatively short. For control simulation, about 100 data points are sufficient. For example, if the time-varying parameter is communication delay, it takes about 50 ms to acquire one time-varying parameter, so 5 to 10 seconds is sufficient to acquire 100 time-varying parameters. When approximating the predictive distribution with the past 100 time-varying parameters from the present, after the 100th time-varying parameter, the predictive distribution can be updated for each time-varying parameter, and the control is also updated for each such time-varying parameter. For example, if the controlled object is a car, the car can be idled for 10 seconds from the start of control before starting the car. After that, if there is a change in the predictive distribution, the control will adapt to the change every 50 ms, making it possible to respond to the change at high speed.
[0187] Furthermore, the feature quantity η at each time point k can be obtained by methods other than those in Embodiments 1 and 2. k+1|k The time-varying parameter ξ may be calculated at each time k. For example, at each time k, the time-varying parameter ξk If the sample value cannot be obtained directly, the time-varying parameter ξ can be obtained using a particle filter or an ensemble Kalman filter, etc. k A distribution of predicted values can be constructed from a set of particles that predict one step ahead and the weight of each particle. By solving a quadratic programming problem similar to that in Embodiments 1 and 2 using this distribution as the prediction distribution, the feature η can be obtained. k+1|k Alternatively, the method of using the filter may be combined with Embodiment 1 or 2 to determine the feature quantity η k+1|k You may decide that.
[0188] V. Numerical example of control that adjusts the control gain based on successive distribution prediction Using the same example as in III, we compared the adaptive stochastic control according to Embodiment 1 with the adaptive control according to Embodiment 2, and found that Embodiment 2 was superior in terms of the time-varying parameter ξ k It was found that the adaptation to changes in the probability distribution is rapid, and good results are easily obtained in this example. On the other hand, it was confirmed that if the predetermined time interval can be made sufficiently long (if the time for acquiring sample values can be made sufficiently long), the prediction accuracy may be higher in Embodiment 1 than in Embodiment 2 when no change in the probability distribution occurs. In both Embodiments 1 and 2, the time-varying parameter ξ k There is a trade-off between quickly adapting to changes in the probability distribution and improving prediction accuracy. However, such a trade-off occurs when the sample size N s Alternatively, these can be altered by adjusting the forgetting coefficient γ, etc. Therefore, by appropriately adjusting these parameters according to the specific problem being addressed, a control system better suited to that problem can be achieved.
[0189] As described above, stable adaptive control can be achieved according to the control device, control system, control program, and control method according to Embodiments 1 and 2.
[0190] Each embodiment disclosed herein is intended to be implemented in appropriate combinations, to the extent that they do not contradict each other. The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of this disclosure is indicated by the claims and not by the foregoing description, and all modifications are intended to be in the sense and scope equivalent to the claims. [Explanation of Symbols]
[0191] 1 Steering wheel, 2 Steering shaft, 3 Electric motor, 4 Rack and pinion mechanism, 5 Tie rod, 6 Front knuckle, 7 Vehicle drive system, 9 Acceleration / deceleration control device, 10 Brake control device, 11 Brake, 12 Steering control device, 13 Pinion shaft, 14 Rack shaft, 15 Front wheel, 16 Rear wheel, 20 Vehicle speed sensor, 21 IMU sensor, 22 Steering angle sensor, 23 Steering torque sensor, 50 Bus, 51 Processing circuit, 52 Memory, 53 Input unit, 54 Output unit, 55 Communication unit, 100 Controller, 100A Server, 110,110A Distribution prediction unit, 120,120A Extraction unit, 130,130A Control unit, 131 State estimation unit, 132 Control amount calculation unit, 140 Delay calculation unit, 151,404 Receiving unit, 152,405 Transmitter, 406 Time synchronization unit, 400 Control object, 400A Mobile object, 401 Internal sensor, 402 Command value calculation unit, 403 Actuator, CS Control system, Dc Control data, Ds Sample value, NW Network, Pg Control program, Pr1~Pr6 Time interval, RCS Remote control system, RT1~RT4 Router, TR Target path.
Claims
1. A control device that outputs a first parameter to the controlled object that changes the state of the controlled object, A distribution prediction unit that constructs a predictive distribution, which is a future probability distribution, from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, A control device comprising a control unit that calculates the first parameter based on the aforementioned feature quantity and the aforementioned state.
2. A control device that outputs a first parameter to a controlled object that changes the state of the controlled object, A distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, The system comprises a control unit that calculates the first parameter based on the aforementioned feature quantity and the aforementioned state, The extraction unit is a control device that derives an approximation formula for the predicted distribution and extracts the feature quantities from the approximation formula.
3. The control device according to claim 2, wherein the extraction unit extracts the feature quantity by minimizing the distance between the approximation formula and the predicted distribution.
4. The extraction unit derives a mixture distribution, which is the sum of a predetermined number of basis distributions, as the approximation formula, derives the coefficients of the multiple basis distributions as the solution to the quadratic programming problem that minimizes the distance, and extracts the coefficients of the multiple basis distributions as the feature quantities. The first parameter is expressed as the product of the control gain and the state, The control device according to claim 3, wherein the control gain is expressed as the sum of a predetermined number of base gains, each having a coefficient of the plurality of base distributions.
5. The control device according to claim 1, wherein the at least one time is included in the time obtained by going back a predetermined time interval from the current time.
6. A control device that outputs a first parameter to a controlled object that changes the state of the controlled object, A distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, The system comprises a control unit that calculates the first parameter based on the aforementioned feature quantity and the aforementioned state, A control device in which the predicted distribution for the current time is the sum of a term obtained by multiplying the predicted distribution for the time one step prior to the current time by a forgetting coefficient, and a term obtained by multiplying a specific function having a peak in the sample value of the second parameter at the current time by a value obtained by subtracting the forgetting coefficient from 1.
7. The controlled object and the control device are connected via a network where transmission delays occur. The control device according to any one of claims 1 to 6, wherein the second parameter includes the transmission delay.
8. The controlled object and, The system includes a control device that outputs a first parameter to the controlled object that changes the state of the controlled object, The control device is A distribution prediction unit that constructs a predictive distribution, which is a future probability distribution, from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, A control system including a control unit that calculates the first parameter based on the aforementioned feature quantities and the aforementioned state.
9. A controlled object and The system includes a control device that outputs a first parameter to the controlled object that changes the state of the controlled object, The control device is A distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, Includes a control unit that calculates the first parameter based on the feature quantity and the state, The extraction unit is a control system that derives an approximation formula for the predicted distribution and extracts the feature quantities from the approximation formula.
10. A controlled object and The system includes a control device that outputs a first parameter to the controlled object that changes the state of the controlled object, The control device is A distribution prediction unit that constructs a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time, An extraction unit that extracts features from the aforementioned predictive distribution, Includes a control unit that calculates the first parameter based on the feature quantity and the state, A control system in which the predicted distribution for the current time is the sum of a term obtained by multiplying the predicted distribution for the time one step prior to the current time by a forgetting coefficient, and a term obtained by multiplying a specific function having a peak in the sample value of the second parameter at the current time by a value obtained by subtracting the forgetting coefficient from 1.
11. A control program, when executed by a processor, causes the processor to perform processing to output a first parameter to the controlled object that changes the state of the controlled object, The aforementioned process is, A predictive distribution, which is a future probability distribution, is constructed from at least one sample value of a second parameter that changes the characteristics of the controlled object, corresponding to at least one time, including the current time. Extracting features from the aforementioned predictive distribution, A control program that includes calculating the first parameter based on the aforementioned feature quantities and the aforementioned state.
12. A control program, when executed by a processor, causes the processor to perform processing to output a first parameter that changes the state of the controlled object to the controlled object, The aforementioned process is, A predictive distribution is constructed from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time. Extracting features from the aforementioned predictive distribution, This includes calculating the first parameter based on the aforementioned feature quantities and the aforementioned state, The extraction process involves a control program that derives an approximation formula for the predicted distribution and extracts the feature quantities from the approximation formula.
13. A control program, when executed by a processor, causes the processor to perform processing to output a first parameter that changes the state of the controlled object to the controlled object, The aforementioned process is, A predictive distribution is constructed from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time. Extracting features from the aforementioned predictive distribution, This includes calculating the first parameter based on the aforementioned feature quantities and the aforementioned state, A control program in which the predicted distribution for the current time is the sum of a term obtained by multiplying the predicted distribution for the time one step prior to the current time by a forgetting coefficient, and a term obtained by multiplying a specific function having a peak in the sample value of the second parameter at the current time by a value obtained by subtracting the forgetting coefficient from 1.
14. A control method that performs processing to output a first parameter to the controlled object that changes the state of the controlled object, A step of constructing a predictive distribution, which is a future probability distribution, from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time. The steps include: extracting features from the aforementioned predictive distribution, A control method comprising the step of calculating the first parameter based on the aforementioned feature quantities and the aforementioned state.
15. A control method for performing a process to output a first parameter that changes the state of the controlled object to the controlled object, A step of constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time; The steps include: extracting features from the aforementioned predictive distribution, The step includes calculating the first parameter based on the aforementioned feature quantities and the aforementioned state, The extraction step is a control method that derives an approximation formula for the predictive distribution and extracts the features from the approximation formula.
16. A control method for performing a process to output a first parameter that changes the state of the controlled object to the controlled object, A step of constructing a predictive distribution from at least one sample value of a second parameter that changes the characteristics of the controlled object, each corresponding to at least one time, including the current time; The steps include: extracting features from the aforementioned predictive distribution, The step includes calculating the first parameter based on the aforementioned feature quantities and the aforementioned state, A control method in which the predicted distribution for the current time is the sum of a term obtained by multiplying the predicted distribution for the time one step prior to the current time by a forgetting coefficient, and a term obtained by multiplying a specific function having a peak in the sample value of the second parameter at the current time by a value obtained by subtracting the forgetting coefficient from 1.