Information processing systems, information processing methods, and programs

A virtual control system in a virtual space is used to simulate the control of a controlled object, allowing for the technical problem of existing technologies have not addressed or efficiently solve the challenges of existing systems, the technical solution addresses the technical problem of optimizing controllers in real-world environments using reinforcement learning, reducing the effort required to adapt them to the real environment by simulating the control method of existing systems, the technical solution of using a virtual control system, the technical solution of solving the technical problem of existing systems, the technical solution of solving the technical problem of optimizing the control system in a virtual environment, the technical efficacy of reducing the effort required to adapt the control system to the real environment by simulating in a virtual space.

JP2026079158APending Publication Date: 2026-05-15KK TOYOTA CHUO KENKYUSHO
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KK TOYOTA CHUO KENKYUSHO
Filing Date
2024-10-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Optimizing controllers in real-world environments using reinforcement learning is costly due to the need for shutting down operations and creating dedicated environments for training.

Method used

A virtual control system is generated in a virtual space corresponding to a real space, using a first and second virtual controller to simulate the controlled object, where the first controller performs PID control and the second controller performs reinforcement learning, with a mixer combining their signals for feedback control.

Benefits of technology

This approach allows reinforcement learning to be performed on real controllers in operation, reducing the effort required to adapt them to the real environment by simulating in a virtual space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079158000001_ABST
    Figure 2026079158000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system, method, and program for performing feedback control over a virtual control target. [Solution] The method using an information processing system involves an acquisition unit acquiring information about a control object in real space and information about a real controller equipped with the control object, the real controller controlling the autonomous operation of the control object to generate a virtual control system in a virtual space corresponding to the real space that controls the operation of a virtual control object that simulates the control object, the virtual control system generating a first control signal that can control the control object with a first virtual controller, and generating a second control signal that can control the control object by performing reinforcement learning using a pre-set reward with a second virtual controller, the mixer outputting an input signal to be input to the virtual control object by mixing the first control signal and the second control signal, and the virtual control object performing feedback control to the virtual control object by inputting the input signal output from the mixer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 discloses a motor control device for achieving both advanced control and stability. The motor control device comprises a first control circuit, a second control circuit, a determination circuit, and a command circuit. The first control circuit outputs a rule-based first control value from a command value of angular velocity and a measured value of angular velocity. The second control circuit outputs a learned model second control value from the command value of angular velocity and the measured value of angular velocity. The determination circuit determines the state based on at least the second control value. The command circuit obtains and outputs a control command value from the first control value and the second control value based on the result determined by the determination circuit. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2024-44031 [Overview of the project] [Problems that the invention aims to solve]

[0004] However, optimizing existing controllers operating in real-world environments using reinforcement learning incurs significant costs, such as shutting down their operation in the real environment and preparing a dedicated environment for reinforcement learning. [Means for solving the problem]

[0005] According to one aspect of the present invention, an information processing system comprises at least one processor capable of executing a program so as to perform the following steps: an acquisition step acquires information about a controlled object in real space and information about a real controller provided by the controlled object, the real controller being configured to control the autonomous operation of the controlled object; and a generation step generates a virtual control system in a virtual space corresponding to the real space, based on the acquired information about the controlled object and the information about the real controller, which controls the operation of a virtual controlled object that simulates the controlled object, the virtual control system comprising a first virtual controller, a second virtual controller, and a mixer, and the first virtual controller An information processing system is provided in which a device generates a first control signal that can control a virtual controlled object based on the operation results of the virtual controlled object, in order to simulate the control of a controlled object by a real controller; a second virtual controller generates a second control signal that can control the controlled object by performing reinforcement learning based on the behavior of the virtual controlled object using a preset reward; a mixer is configured to output an input signal that is input to the virtual controlled object by mixing the first control signal and the second control signal; and in the target control step, feedback control is performed on the virtual controlled object by inputting the input signal output from the mixer to the virtual controlled object.

[0006] With this configuration, for example, reinforcement learning can be performed on a real controller of a controlled object that is already in operation in a real environment, through simulation in a virtual space. Therefore, the effort required to implement a controller with reinforcement learning results suitable for the operating environment of the controlled object in the real environment can be reduced. [Brief explanation of the drawing]

[0007] [Figure 1] This is a diagram showing the configuration of Information Processing System 1. [Figure 2] This is a block diagram showing the hardware configuration of the information processing device 2. [Figure 3] This is a block diagram showing the hardware configuration of user terminal 3. [Figure 4]This figure shows an example configuration of the real space RS that is the subject of this information processing, and the virtual space VS that corresponds to the real space RS. [Figure 5] This flowchart shows an example of the information processing flow performed in Information Processing System 1. [Figure 6] This figure shows an example configuration of virtual control system 5. [Figure 7] This figure shows an example configuration of a second virtual controller 53 that can accept high-dimensional features as input. [Figure 8] This figure shows an example configuration of a second virtual controller 53 that generates a second control signal s26 based solely on the first measured value s22. [Modes for carrying out the invention]

[0008] Embodiments of the present invention will be described below with reference to the drawings. The various features shown in the embodiments below can be combined with each other.

[0009] Incidentally, the program for implementing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or it may be provided as a downloadable medium from an external server, or it may be provided so that the program is launched on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0010] Furthermore, in various information processing according to one embodiment, an input and an output corresponding to the input can be realized. Here, as long as an output is obtained as a result of the input, the form of the information referenced in such information processing (hereinafter referred to as "reference information") is not limited. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression equation constructed by a statistical method), or a pre-trained model that has learned the correlation between input and output in advance, or a large-scale language model that can output a desired result by inputting a prompt.

[0011] Also, in one embodiment, the "unit" may include, for example, hardware resources implemented by a circuit in a broad sense and information processing of software that can be specifically realized by these hardware resources. Also, in one embodiment, various information is handled, and this information is represented, for example, by a physical value of a signal value representing voltage or current, the level of a signal value as a binary bit aggregate composed of 0 or 1, or a quantum superposition (so-called qubit), and communication and calculation can be executed on a circuit in a broad sense.

[0012] Furthermore, a circuit in a broad sense is a circuit realized by appropriately combining at least a circuit (Circuit), circuitry (Circuitry), a processor (Processor), and a memory (Memory), etc. Also, the processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (for example, a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0013] 1. Hardware Configuration In this section, the hardware configuration will be described.

[0014] <Information Processing System 1> FIG. 1 is a configuration diagram showing an information processing system 1. The information processing system 1 includes an information processing apparatus 2 and a user terminal 3. The information processing apparatus 2 and the user terminal 3 are configured to be communicable through a telecommunication line. In one embodiment, the information processing system 1 consists of one or more devices or components. For example, if it consists only of the information processing apparatus 2, the information processing system 1 can be the information processing apparatus 2. Hereinafter, these components will be described.

[0015] <Information processing apparatus 2> FIG. 2 is a block diagram showing the hardware configuration of the information processing apparatus 2. The information processing apparatus 2 includes a communication unit 21, a storage unit 22, and a processor 23, and these components are electrically connected via a communication bus 20 inside the information processing apparatus 2. Each component will be further described.

[0016] The communication unit 21 preferably uses wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), and wired LAN network communication. However, wireless LAN network communication, mobile communication such as 3G / LTE / 5G, and BLUETOOTH (registered trademark) communication may be included as needed. That is, it is more preferable to implement it as a set of these multiple communication means. That is, the information processing apparatus 2 may communicate various information from the outside via the communication unit 21 and the network.

[0017] The storage unit 22 stores various information defined as described above. This can be implemented as a storage device such as a solid state drive (SSD) that stores various programs related to the information processing apparatus 2 executed by the processor 23, or as a memory such as a random access memory (RAM) that stores temporarily necessary information (arguments, arrays, etc.) related to the calculation of the program. The storage unit 22 stores various programs, variables, etc. related to the information processing apparatus 2 executed by the processor 23.

[0018] The processor 23 performs processing and control of the overall operation related to the information processing device 2. The processor 23 is, for example, a central processing unit (CPU) not shown. The processor 23 realizes various functions related to the information processing device 2 by reading predetermined programs stored in the memory unit 22. That is, information processing by software stored in the memory unit 22 can be concretely realized by the processor 23, which is an example of hardware, and executed as each functional unit included in the processor 23. These will be described in more detail in the next section. Note that the processor 23 is not limited to being a single unit, and may be implemented with multiple processors 23 for each function, or a combination thereof.

[0019] The processor 23 is configured as an acquisition unit to acquire information from the user terminal 3 or other devices. The processor 23 is configured to acquire various information by reading various information stored in the storage area, which is at least a part of the memory unit 22, and writing the read information to the work area, which is at least a part of the memory unit 22. The storage area is, for example, the area of ​​the memory unit 22 that is implemented as a storage device such as an SSD. The work area is, for example, the area that is implemented as memory such as RAM. The acquisition by the processor 23 includes acquiring the output results of each functional unit included in the processor 23.

[0020] The processor 23 is configured as a display processing unit to display various types of information. This information can be presented to the user via the display unit 34 of the user terminal 3 or another device. In such a case, for example, the processor 23 controls the display unit 34 of the user terminal 3 to display visual information such as screens, images including still images or videos, icons, and messages. The processor 23 may also generate only rendering information for displaying the visual information on the user terminal 3. The processor 23 may also present the outputted information to the user without going through the user terminal 3 or another device user.

[0021] <User Terminal 3> Figure 3 is a block diagram showing the hardware configuration of the user terminal 3. The user terminal 3 comprises a communication unit 31, a storage unit 32, a processor 33, a display unit 34, and an HMI device 35, and these components are electrically connected within the user terminal 3 via a communication bus 30. The descriptions of the communication unit 31, storage unit 32, and processor 33 are the same as the descriptions of each part in the information processing device 2, so they are omitted here.

[0022] The display unit 34 may be included in the user terminal 3 housing or it may be an external component. The display unit 34 displays a graphical user interface (GUI) screen that can be operated by the user. This is preferably done by using different display devices such as a CRT display, liquid crystal display, organic EL display, and plasma display, depending on the type of user terminal 3.

[0023] The HMI device 35 is a human-machine interface device. The HMI device 35 may be included in the housing of the user terminal 3 or it may be an external device. For example, the HMI device 35 may be implemented as a touch panel integrated with the display unit 34. If it is a touch panel, the user can input tap operations, swipe operations, etc. Of course, instead of a touch panel, a switch button, mouse, QWERTY keyboard, voice recognition device, gesture detection device, gaze detection device, biosignal detection device, imaging device, etc. may be used. In other words, the HMI device 35 receives operation input made by the user. In response, the HMI device 35 transmits a signal corresponding to the operation input to the processor 33 via the communication bus 30. The processor 33 can perform predetermined controls and calculations as needed. The HMI device 35 can also be said to include an input unit configured to accept input from the user.

[0024] 3. Regarding information processing This section describes the information processing performed in the aforementioned information processing system 1.

[0025] 3.1. Regarding the space subject to information processing This information processing can be performed to optimize the operation of a controlled object 100, which is configured to operate autonomously, by reproducing the operation of the real-space RS in a virtual space VS using reinforcement learning, while also reproducing various constraints such as the purpose of use and layout of the real-space RS. Through simulation in the virtual space VS, parameters that can optimize the operation of the controlled object 100 can be searched for. Therefore, first, an example of the system (real-space RS) that is the target of this information processing and the virtual space VS corresponding to the real-space RS will be described. Figure 4 shows an example configuration of the real-space RS that is the target of this information processing and the virtual space VS corresponding to the real-space RS.

[0026] Real-space RS is a system in which tangible objects can operate and may contain spatial information. Note that the operation of tangible objects is not necessarily limited to mechanical operation, but may include any physical system capable of transmitting energy such as electricity or information. Real-space RS includes the controlled object 100. The controlled object 100 is configured to perform autonomous operations such as automatic driving in real-space RS. In one embodiment, the controlled object 100 is a mobile body that can move within real-space RS. In this embodiment, the controlled object 100 is a forklift. The controlled object 100 as a forklift comprises a passenger compartment 101, a cargo handling device 102, and a control device 103. The passenger compartment 101 is configured to allow a driver to ride on it and is equipped with electrical equipment for moving the forklift. The cargo handling device 102 is configured, for example, to be located in front of the passenger compartment 101, enabling the handling of cargo placed in front of it. Specifically, the cargo handling device 102 is configured to handle objects loaded on pallets, etc., together with the pallets by moving a pair of forks up and down. The control device 103 is configured to control the autonomous operation of the controlled object 100. The control device 103 can be implemented using any information processing device, such as an integrated circuit or a microcomputer. Here, the control device 103 is treated as controlling the local operation of the controlled object 100. The control device 103 can function as a real controller configured to control the autonomous operation of the controlled object 100. The real controller can be configured to perform control based on a linear model, such as PID control, which is an example of classical control. The specific form of the forklift is arbitrary as long as it can move autonomously, i.e., semi-automatically or fully automatically, and it may be a reach type or a counterbalanced type, and may be manned or unmanned.

[0027] For example, the controlled object 100 can move according to a travel plan from its current location P1 to its destination G1 based on a pre-set target path R1. The control device 103 controls the motion of the controlled object 100 (for example, the throttle speed and steering angle of the controlled object 100's wheels) to move the current location P1 to a predetermined position on the target path R1, based on the discrepancy (deviation) between the target path R1 defined in the travel plan and the current location P1.

[0028] The virtual space VS is a space modeled to reproduce the real space RS. In this embodiment, the virtual space VS is generated as a space on which simulations relating to spatial behavior, such as physical calculations, can be performed. The virtual space VS can be modeled based on constraints imposed on the real space RS (for example, constraints on the movable area, constraints on the performance of the controlled object 100, constraints on the control system for autonomous control of the controlled object 100, etc.). Modeling may be performed by the user operating the user terminal 3, or it may be performed automatically by inputting the constraints into the user terminal 3. When the simulation is performed in a cloud environment, the generated virtual space VS can be constructed on the information processing device 2.

[0029] In this embodiment, the virtual space VS includes a virtual control object 200. The virtual control object 200 is configured to simulate the control object 100. For example, the virtual control object 200 is a movable object modeled to simulate the constraints on the operation of the control object 100 (e.g., size, shape, weight, acceleration method, speed, etc.). In this embodiment, the virtual control object 200 is a three-dimensional model of the control object 100 as a forklift, and includes a virtual passenger compartment 201 that simulates the function of the passenger compartment 101, and a virtual cargo handling device 202 that simulates the function of the cargo handling device 102. Furthermore, even if the configuration corresponding to the control device 103 in the virtual space VS is virtually provided by the virtual control object 200, it may also be realized by the simulator that constructs the virtual space VS directly controlling the behavior of the virtual control object 200. In this embodiment, the simulator is treated as being executed by the processor 23 of the information processing device 2. The virtual controlled object 200, like the controlled object 100, is controlled to move from its current location P2 in the virtual space VS to its destination G2 in the virtual space VS, along the target path R2 in the virtual space VS that corresponds to the target path R1 in the real space RS.

[0030] 3.2. Information Processing Flow Next, an example of the information processing flow performed in Information Processing System 1 will be described. Figure 5 is a flowchart showing an example of the information processing flow performed in Information Processing System 1. Note that this information processing may include any exception handling not shown. Exception handling includes interrupting the information processing or omitting each process. The selection or input performed in this information processing may be based on user operation or may be performed automatically without user operation.

[0031] [Step S1] First, in step S1, the processor 23, as an acquisition unit, acquires information about the controlled object 100 in real space RS and information about the actual controller of the controlled object 100. The information about the controlled object 100 may include the performance, physical properties, mechanical characteristics, and operational constraints of the controlled object 100 as described above. The information about the actual controller is information about the control mode of the controlled object 100 by the control device 103. In this embodiment, the actual controller is assumed to be a PID controller, and the information about the actual controller may include parameters used when performing PID control (e.g., proportional gain, integral time, derivative time, etc.). This information can be acquired, for example, by being input to the user terminal 3 in response to user operation, and then transmitted from the user terminal 3 to the information processing device 2.

[0032] [Step S2] Next, in step S2, the processor 23 generates a virtual control target 200 based on the information about the control target 100 obtained in step S1 and places it in the virtual space VS. If a virtual control target 200 has already been generated, the processor 23 only needs to place that virtual control target 200 in the virtual space VS.

[0033] [Step S3] Next, in step S3, the processor 23, as a generation unit, generates a virtual control system 5 (see Figure 6) in the virtual space VS corresponding to the real space RS, based on the information about the controlled object 100 and the information about the actual controller 42 acquired in step S1, to control the operation of the virtual controlled object 200. The details of the virtual control system 5 will be described in detail in the following sections.

[0034] [Step S4] Next, in step S4, the processor 23 sets a route plan for the virtual controlled object 200. This route plan may be associated with a route plan that has been pre-set for the controlled object 100, or it may be a route plan that is arbitrarily set by the user. The route plan may include, for example, a starting point D1, a destination G1 in real space RS, and a target route R1 from the starting point D1 to the destination G1. The target route R1 may be approximately defined using a certain curve. This curve may be a clothoid curve, a Bézier curve, a spline curve, or any other grid-based or graph-based route planning method.

[0035] [Step S5] Next, in step S5, the processor 23 controls the virtual control object 200 based on the latest state of the virtual control object 200 and the path plan. The latest state of the virtual control object 200 can be represented by low-dimensional state quantities such as the latest position (coordinates), orientation, and velocity of the virtual control object 200 in the virtual space VS, and the state of the virtual cargo handling device 202 (e.g., fork height, tilt angle, etc.), or by high-dimensional state quantities such as image data in the virtual space VS relating to the virtual control object 200. Image data as a high-dimensional state quantity may include, for example, a subjective image showing the subjective viewpoint of the virtual control object 200, or an objective image of the virtual control object 200 taken from a third-person viewpoint.

[0036] Low-dimensional state variables are those with a number of parameters less than or equal to the degrees of freedom describing the behavior (in this case, the mode of movement) of the virtual controlled object 200. For example, in a 3D space, three position coordinates (3D vectors), an angle representing the orientation of the virtual controlled object 200 (scalar), the velocity of the virtual controlled object 200 (3D vector), a throttle signal (scalar), and a steering signal (scalar) are all low-dimensional state variables. High-dimensional state variables, on the other hand, are a set of feature quantities for each pixel of an image, and their degrees of freedom can exceed 100. In other words, the behavior of the virtual controlled object 200 in response to an input signal can be represented using low-dimensional state variables with a number of parameters less than or equal to the degrees of freedom describing the motion of the virtual controlled object 200, and high-dimensional state variables with a number of parameters greater than those degrees of freedom. High-dimensional state variables can include not only image data, but also any state variables that can be described by a large number of feature quantities, such as observational data related to vibration, tactile data, etc.

[0037] [Step S6] Next, in step S6, the processor 23 measures the behavior of the virtual controlled object 200 that was controlled (moved) in step S5. The parameters measured may include low-dimensional state quantities such as the latest position, orientation, and velocity of the virtual controlled object 200 in the virtual space VS, and the state of the virtual cargo handling device 202 (e.g., fork height, tilt angle, etc.), as well as high-dimensional state quantities such as image data of the virtual controlled object 200 in the virtual space VS.

[0038] [Step S7] Next, in step S7, the processor 23 performs reinforcement learning in the virtual control system 5 based on the measurement results in step S6. Parameters input during reinforcement learning include, for example, the changes (deviations) in the measured low-dimensional state variables and the features contained in the measured image data. The specific reinforcement learning algorithm can be any algorithm, such as dynamic programming, Monte Carlo method, time-difference learning, or nearest neighbor policy optimization. Preferably, the reinforcement learning is deep reinforcement learning using a neural network model.

[0039] [Step S8] Next, in step S8, the processor 23 updates the internal parameters of the virtual control system 5 based on the results of reinforcement learning.

[0040] [Step S9] Subsequently, in step S9, the processor 23 determines whether the optimization of the internal parameters is complete. Whether the optimization is complete may be determined depending on whether predetermined convergence conditions are met. Convergence conditions may be met, for example, when the amount of change in the internal parameters is less than or equal to a predetermined value, or when a predetermined number of learning iterations have been performed. If it is determined in step S9 that the optimization is not complete, the process returns to step S4, and the processor 23 resets the path plan and performs reinforcement learning again.

[0041] [Step S10] On the other hand, if it is determined in step S9 that optimization is complete, in step S10 the processor 23 outputs information regarding the final internal parameters of the virtual control system. Based on this information regarding internal parameters, the processor 23 may implement the control system in real space RS corresponding to the virtual control system 5 in real space on the control device 103 or other device. The processor 23 may present this information to the user via the display unit 34 and present to the user a method for implementing the control system in real space RS.

[0042] Subsequently, information processing system 1 terminates this information processing.

[0043] 3.3. Examples of Virtual Control System Configurations Next, we will describe an example configuration of the virtual control system 5 generated in step S3. Figure 6 shows an example configuration of the virtual control system 5.

[0044] The virtual control system 5 according to this embodiment simulates the actual control system 4 that controls the autonomous operation of the controlled object 100, and is configured to optimize the control of the virtual controlled object 200 by performing reinforcement learning based on the behavior of the virtual controlled object 200 in the virtual space VS.

[0045] The actual control system 4 is a control system for controlling the behavior of the controlled object 100. The actual control system 4 can be implemented, for example, by a control device 103 or other devices in real space RS. For example, the actual control system 4 comprises a route planner 41 and an actual controller 42. The route planner 41 generates a target value s11 for moving the controlled object 100 along a travel path, based on planning information regarding the travel plan of the controlled object 100. The target value s11 may include information about the origin D1, destination G1, and target route R1 of the controlled object 100.

[0046] The actual controller 42 is configured to control the behavior of the controlled object 100 based on the set movement plan. Here, the actual controller 42 inputs the deviation of measured values ​​s12, such as the current location P1 of the controlled object 100, from the target path R1 set by the path planner 41 as a deviation signal s13. This input then outputs an actual input signal s14 that includes a command for the front wheel rotation angular velocity related to longitudinal control of the forklift as the controlled object 100 (hereinafter referred to as a throttle command) and a command for the steering angle of the rear wheels corresponding to lateral control (hereinafter referred to as a steering command). Based on the actual input signal s14 output from the actual controller 42, the controlled object 100 executes PID control of its steering wheels and attempts to move the controlled object 100 along the target path R1. The state of the controlled object 100 after movement (for example, current location P1, speed, direction, etc.) is measured by various sensors provided on the controlled object 100 and output as measured values ​​s12. The measured value s12 is converted into a deviation signal s13 by taking the deviation from the target value s11, and is input again to the actual controller 42. In this way, the actual controller 42 performs feedback control regarding the behavior of the controlled object 100 and controls the steering wheels of the controlled object 100 so that the controlled object 100 moves along the target path R1. Note that the actual controller 42 performs control based on low-dimensional state quantities output as the measured value s12, and does not accept input of high-dimensional state quantities such as image data features.

[0047] In contrast, the virtual control system 5, which corresponds to the actual control system 4, is configured to perform feedback control on the virtual control target 200 and comprises a virtual path planner 51, a first virtual controller 52, a second virtual controller 53, and a mixer 54. The virtual control target 200 moves based on the input signal s3, which will be described later. The measurement results of its behavior can be measured as a first measurement value s22 corresponding to a low-dimensional state quantity and a second measurement value s25 corresponding to a high-dimensional state quantity. The first measurement value s22 is a signal corresponding to the measurement value s12 related to the control target 100 and may include the current location P1 (spatial coordinates), orientation, speed, throttle command, steering command, etc., of the virtual control target 200 after movement. On the other hand, the second measurement value s25 can be defined as a set of feature quantities of the subjective image (or objective image) of the virtual control target 200.

[0048] The virtual route planner 51 is configured to simulate the route planner 41 in the virtual space VS. Based on the movement plan of the virtual control target 200, the virtual route planner 51 outputs target values ​​s21 for parameters that indicate the state of the virtual control target 200 (e.g., speed, direction of travel, fork height position, etc.).

[0049] The first virtual controller 52 generates a first control signal s24 that can control the controlled object 100 based on the operation result of the virtual controlled object 200, in order to simulate the control of the controlled object 100 by the actual controller 42. In this embodiment, the first virtual controller 52 is configured to perform PID control in the same way as the actual controller 42. For example, similar to the actual controller 42, the first virtual controller 52 is configured to output the first control signal s24 by inputting a low-dimensional state variable. The low-dimensional state variable may include a parameter relating to at least one of the position, velocity, and acceleration of the virtual controlled object 200. In this embodiment, the first virtual controller 52 is configured to generate a first control signal s24 based on the deviation between the virtual movement plan, which is the movement plan of the virtual controlled object 200 set by the virtual path planner 51, and the behavior of the virtual controlled object 200. In this embodiment, the virtual path planner 51 outputs a target value s21 to the first virtual controller 52 based on the input movement plan, and the first virtual controller 52 is configured to output a first control signal s24 based on the target value s21 and a first measured value s22, according to a deviation signal s23 which is the difference between them. The first control signal s24 is obtained by performing predetermined signal processing, such as calculating a PID gain on the deviation signal s23. The virtual movement plan may include a starting point D2, a destination G2, a target path R2 from the starting point D2 to the destination G2 in the virtual space VS.

[0050] The second virtual controller 53 generates a second control signal s26 capable of controlling the control target 100 by performing reinforcement learning based on the behavior of the virtual control target 200 using a pre-set reward. For example, the second virtual controller 53 is configured to perform reinforcement learning based on low-dimensional state variables and high-dimensional state variables. With such a configuration, high-dimensional state variables that are difficult for the first virtual controller 52 to process can be processed using the second virtual controller 53. Therefore, the control system can be optimized according to more complex situations that cannot be grasped by low-dimensional state variables alone. The high-dimensional state variables may include image data in the virtual space VS that shows the behavior of the virtual control target 200 (for example, subjective and objective images of the virtual control target 200 as described above). With such a configuration, a more precise virtual control system 5 can be constructed based on visual information in the virtual space VS. The parameters included in the second control signal s26 are of the same type as those in the first control signal s24 and are configured to be linearly combinable.

[0051] The specific form of the reward used in reinforcement learning is arbitrary, but for example, the second virtual controller 53 is configured to perform reinforcement learning using a reward set to reduce the deviation between the virtual movement plan and the behavior of the virtual controlled object 200. With such a configuration, for example, reinforcement learning can be performed in the virtual space VS for control of the movement of a moving object, which is difficult to perform reinforcement learning in the real space RS, thus optimizing the control of the movement of the moving object more efficiently. For example, the reward is set to be higher the smaller the difference between the position on the planned target path R1 and the actual current location P1 of the virtual controlled object 200. Specifically, the reward for reinforcement learning is the sum of the reciprocals of the parameter deviations used to control the virtual controlled object 200, for example, the weighted sum of the reciprocals of the deviations related to the longitudinal control (throttle command) and the lateral control (steering command) as described above. Note that in addition to the reward based on the deviation from the planned path, any reward may be considered depending on the task. For example, rewards may be considered for not moving (or moving) an object (a pallet in the case of a forklift) at the target position, rewards based on the distance to obstacles, and rewards based on the time taken to reach the target position. The first virtual controller 52 in this embodiment is configured to perform reinforcement learning based on a neural network model. An example of the configuration of the second virtual controller 53 will be described later.

[0052] The mixer 54 is configured to output an input signal s3 to be input to the virtual control target 200 by mixing the first control signal s24 and the second control signal s26. For example, the mixer 54 outputs a weighted linear sum of the first control signal s24 and the second control signal s26 as the input signal s3 to the virtual control target 200. The mixing ratio of the first control signal s24 and the second control signal s26 (for example, the weight of each signal) can be set as appropriate. The processor 23 may perform reinforcement learning starting from a state where the weight of the first control signal s24 is greater than the weight of the second control signal s26, and after convergence, sequentially increase the weight of the second control signal s26 relative to the first control signal s24 to perform reinforcement learning. This allows for the gradual introduction of optimization elements through reinforcement learning from a state where the weight of the first control signal s24 is large and close to the actual behavior of the real control system 4, thereby obtaining a virtual control system 5 that is less likely to contradict the realistic behavior of the real control system 4 and has high interpretability. Furthermore, the control signals s24 and s26 from the first virtual controller 52 and the second virtual controller 53 can be mixed using a weighted sum of the two signals, or by using an arbitrary function such as the geometric mean. In order to maintain intuitive interpretability by providing a correction amount based on the first virtual controller 52, it is desirable to use a simple countable sum or a weighted sum.

[0053] The processor 23 inputs the input signal s3 output from the mixer 54 to the virtual control target 200, thereby performing feedback control on the virtual control target 200. With this configuration, for example, reinforcement learning for the actual controller 42 of the control target 100, which is already in operation in the real space RS, can be performed through simulation in the virtual space VS. Therefore, the effort required to implement a controller with reinforcement learning results suitable for the operating environment of the control target 100 in the real space RS can be reduced.

[0054] The virtual control system 5 may further include a gain adjuster 55. The gain adjuster 55 is configured to adjust the second control signal s26 output from the second virtual controller 53 based on a set gain. As a result, the gain adjuster 55 outputs a gain adjustment signal s27, which is the second control signal s26 with the adjusted gain, to the mixer 54. In this case, the processor 23, as a gain setting unit, may set the gain within a range that stabilizes the control of the virtual control system 5 using the second control signal output from the second virtual controller 53 based on the state equation of the virtual control system 5 in the virtual space VS. The mixer 54 may also be configured to mix the first control signal with the adjusted second control signal. With such a configuration, the operation of the virtual control target 200 in the virtual space VS can be stabilized, and reinforcement learning can be executed more smoothly. Specific embodiments of gain adjustment will be described later.

[0055] The virtual control system 5 may further include a parameter estimator 56 that estimates the control parameters of the virtual controlled object 200. The parameter estimator 56 estimates state variables in a model describing the behavior of the virtual controlled object 200 based on low-dimensional and high-dimensional state variables, and outputs the estimation result s28 to the second virtual controller 53. Such state variables include the mass (and even the mass distribution) and the center of gravity of the virtual controlled object 200. The center of gravity is the position when a virtual cargo is loaded onto the virtual cargo handling device 202. The parameter estimator 56 estimates these state variables from the input first measurement value s22 and second measurement value s25, for example, using a Kalman filter based on a dynamic system model of a forklift as the controlled object 100. Furthermore, if the mass distribution of the model of the virtual control target 200 is predetermined, the processor 23 may set an initial value for the center of gravity of the virtual control target 200 based on that mass distribution and use it as the initial value when the parameter estimator 56 estimates the parameters.

[0056] If the parameter estimator 56 is used, the second virtual controller 53 may further generate a second control signal s26 based on the estimated state variables. With this configuration, reinforcement learning can be performed by taking into account information about the behavior of the virtual controlled object 200, which is represented by state variables that are difficult to observe directly. More specifically, a virtual control system 5 can be obtained in which the forklift, as the controlled object 100, adapts to changes in the overall mass and center of gravity position determined by the mass and position of the transported goods. Therefore, while using the existing actual controller 42, an optimized actual control system 4 can be reconstructed that takes such transported goods into consideration.

[0057] The virtual control system 5 may further include a randomizer 57. The randomizer 57 is configured to introduce random noise to the parameters input to the first virtual controller 52 or the second virtual controller 53. This configuration can increase the robustness of the virtual control system 5 against system uncertainty. In this embodiment, the randomizer 57 is configured to introduce random noise to the parameters input to the first virtual controller 52 and the second virtual controller 53, respectively. The random noise may include, for example, changing the values ​​of each parameter included in the first measurement value s22 within a predetermined range, or modulating an image to obtain feature quantities included in the second measurement value s25. The image modulation method is arbitrary, but may include, for example, changing the color, shape, or texture of an object in the virtual space VS, changing the background of the virtual space VS, changing the lighting conditions in the virtual space VS, or changing weather conditions. In particular, by introducing such image modulation, the robustness of operation can be increased, for example, when the controlled object 100 is equipped with an image classifier and controls various devices based on the image classifier. Furthermore, the randomizer 57 may be configured to introduce random noise by adding a random mask or object translation to the image data. The randomizer 57 may also add a time delay to the signals input to the first virtual controller 52 and the second virtual controller 53. Additionally, the randomizer 57 may introduce random noise to the estimation result s28 by the parameter estimator 56. Preferably, the noise introduced by the randomizer 57 is configured to introduce random noise individually to the parameters input to the first virtual controller 52 and the second virtual controller 53, respectively. The randomizer 57 may address measurement errors in the real environment or delays that cannot be fully modeled in the controlled object 100. Therefore, when reproducing the virtual control system 5 optimized by reinforcement learning in the actual control system 4, the mechanism corresponding to the randomizer 57 may be omitted.

[0058] 3.4. Example Configuration of the Second Virtual Controller Next, an example configuration of the second virtual controller 53 included in the virtual control system 5 described above will be explained. Figure 7 shows an example configuration of the second virtual controller 53 that can input high-dimensional features. As shown in Figure 7, the second virtual controller 53 is configured to output a throttle command and a steering command as a second control signal s26 by inputting a first measurement value s22 corresponding to a low-dimensional feature and two types of subjective images corresponding to two types of high-dimensional features. The first measurement value s22 is the moving speed v of the controlled object 100. f or, a which indicates the previous input signal s3 (throttle command and steering command). old These may include the following. In this embodiment, the first measured value s22 includes five feature quantities.

[0059] The two types of subjective images are, for example, subjective images captured by two cameras (left camera and right camera) mounted on the left and right sides of the virtual control target 200, each capable of capturing the front view. Note that it is not necessary for the cameras to actually be mounted in the virtual space VS; the processor 23 only needs to capture an image of the front of the control target 100 from a reference position on the simulator. The second virtual controller 53 comprises an image feature extractor 531, a feature combiner 532, and an input signal determination unit 533.

[0060] The image feature extractor 531 is configured to extract features from each image data, particularly those corresponding to the input high-dimensional features. For example, the image feature extractor 531 may be configured to extract image features using a convolutional neural network such as ResNet, which has been pre-trained using the ImageNet dataset and has fixed internal parameters. Alternatively, the image feature extractor 531 may be a neural network such as Vision Transformer, or a classical feature extractor such as Scale-Invariant Feature Transform (SIFT) or Histgrams of Oriented Gradients (HOG). Furthermore, the image feature extractor 531 may use a pre-trained machine learning model as is, or one that has been fine-tuned during reinforcement learning. When utilizing a pre-trained machine learning model, the model may be a foundational model trained using diverse data such as ImageNet, or a model pre-trained using data measured in the actual environment of application or data created in a simulator environment. In this embodiment, the image feature extractor 531 extracts 512 features from each image data.

[0061] The feature combiner 532 is configured to combine features obtained from signals input to the second virtual controller 53. The manner in which the features are combined is arbitrary, but for example, the feature combiner combines the features obtained or extracted from the first measurement value s22 and the second measurement value s25 in series as a single vector. In this embodiment, the feature combiner 532 combines 512 × 2 features extracted from two image data using the image feature extractor 531 and the five features obtained as the first measurement value s22 in series to generate a 1029-dimensional feature vector.

[0062] The input signal determination unit 533 determines the next input signal s3 using a neural network from the features combined by the feature coupler 532, and outputs it to the virtual control target 200 as a throttle command and a steering command. The reinforcement learning performed by the virtual control system 5 is configured to optimize the second virtual controller 53 by updating the internal parameters of the input signal determination unit 533 (for example, the parameters of the neural network). The neural network in this embodiment is configured to include a fully connected layer using an Exponential Linear Unit (ELU) as the activation function. The specific configuration of the neural network is not limited to this and is arbitrary; a sigmoid function or a ReLU function may also be used as the activation function.

[0063] Furthermore, the second virtual controller 53 is not limited to one that performs reinforcement learning based on data corresponding to high-dimensional features such as image data. Figure 8 shows an example configuration of the second virtual controller 53 that generates a second control signal s26 based only on the first measurement value s22. As shown in Figure 8, in this case the second virtual controller 53 does not include an image feature extractor 531 and a feature coupler 532, but only an input signal determination unit 533. In this case the input signal determination unit 533 accepts the set of features included in the acquired first measurement value s22 (for example, the relative position in the x direction, relative position in the y direction, velocity in the x direction, velocity in the y direction, angular velocity of the control object 100, and the previous input signal s3 on a 2D xy plane to which the control object 100 can move) as a single feature vector and generates a second control signal s26 corresponding to the input.

[0064] 3.5. An example of gain adjustment Next, we will explain an example of how to adjust the gain using the gain adjuster 55. First, we assume that the state x of the controlled object 100 is modeled as a linear system as follows.

[0065]

number

[0066] x k This is a state vector that shows the state of the controlled object 100 at step k, and u k Here, A and B represent the parameters for controlling the virtual controlled object 200 (in the example above, the throttle command and the steering command). A and B are coefficients that characterize the linear system. We also assume that the reference state r is generated from the following linear system.

[0067]

number

[0068] Here, A r Assuming that k is Surreal stable, the reference state converges to 0 in the limit of positive infinity of k. Therefore, we can choose a coordinate system such that the final target state is at the origin. In this case, let e = xr be the deviation of the state x of the controlled object 100 from the reference state r, and define its extended state as follows.

[0069]

number

[0070] In this case, the state equation for the extended state e is written as follows:

[0071]

number

[0072] We consider applying the following linear state feedback to the actual control system 4 described by the state equations described above.

[0073]

number

[0074] Thus, the state equation is described as follows.

[0075]

Equation

[0076] The value of the l2 gain from v to z of the system described by such a state equation (in a linear time-invariant system, it is the H ∞ norm) is obtained by solving a linear matrix inequality (LMI). Furthermore, it is also possible to obtain an optimal feedback gain K that minimizes the l2 gain.

[0077] Here, if an additional input signal by a neural network is v = φ(z), according to the small gain theorem, it can be seen that the closed-loop system is stable if the following inequality holds.

[0078]

Equation

[0082] Furthermore, if the above inequality is not satisfied by the output from the second virtual controller 53 and the preset gain, the processor 23 may satisfy the inequality (i.e., stabilization condition) by multiplying the gain by a constant within the range that satisfies the inequality. Alternatively, if the stabilization condition is not satisfied, the processor 23 may set the gain to zero and not adjust the output from the second virtual controller 53. In this case, the parameter estimator 56 can function as a switching controller that determines the stabilization of the system.

[0083] The controlled system 100 may be modeled as a known deterministic system, or it may be modeled taking randomness into account. In this case, the robust stability of the control system can be guaranteed by using a small gain condition with respect to the worst-case gain.

[0084] [others] The above embodiments may be implemented as appropriate based on, for example, the following embodiments.

[0085] The state variables estimated by the parameter estimator 56 are not limited to the mass or center of gravity of the forklift that handles the transported goods. For example, if the virtual control system 5 is a system that performs mechanical control of the virtual controlled object 200, the estimated state variables may be the moment of inertia, the coefficient of air resistance, the coefficient of friction with the road surface, etc. Also, if the virtual control system 5 is a system that performs electrical control of the virtual controlled object 200, the estimated state variables may be parameters that represent the electrical characteristics included in the virtual controlled object 200, such as resistance, inductance, capacitance, etc. Furthermore, the controlled object of the virtual control system 5 is not limited to the mechanical or electrical control of the virtual controlled object 200, but may be any object that can be described as a control system, such as thermal control.

[0086] Furthermore, the method for estimating state variables by the parameter estimator 56 is not limited to the method using a Kalman filter, but can use any algorithm that can be used in modern control, such as the least squares method or the maximum likelihood estimation method. Also, the method for estimating state variables by the parameter estimator 56 may use an estimation model based on machine learning. In the virtual space VS, when performing simulations, the mass and center of gravity position, which change according to the handling conditions of the transported object by the controlled object 100, can be directly obtained from the simulator running on the information processing device 2. Therefore, the values ​​of state variables such as mass and center of gravity position obtained directly from the simulator can be used as ground truth data to perform supervised learning of the estimation model. Consequently, it becomes easier to implement a virtual control system 5 and a real control system 4 that reflect more nonlinear and complex situations.

[0087] The dynamic system model of the controlled object 100 is not limited to a globally linearized linear model; a nonlinear model may be used, which is locally linearized around the initial state or the set target path R1. For example, if the controlled object 100 is a vehicle, the turning angle can be approximated to a linear model by assuming that the roll angle and pitch angle are small, and if it is an aircraft. Furthermore, nonlinear models that can be considered as linear systems through nonlinear transformations, such as the Koopman operator model or the Hammerstein-Wiener model, can be used as nonlinear models.

[0088] The camera used to obtain image data corresponding to high-dimensional features may be attached to the controlled object 100 itself (i.e., capable of capturing subjective images) or attached to the outside of the controlled object 100 (for example, on a wall or ceiling) (i.e., capable of capturing objective images).

[0089] The data corresponding to high-dimensional features is not limited to image data; it can also be point cloud data measured by LiDAR, etc.

[0090] The second virtual controller 53 is not limited to one that performs reinforcement learning using high-dimensional state variables, but may also perform reinforcement learning using only low-dimensional state variables.

[0091] The actual controller 42 (the first virtual controller 52 which mimics the actual controller 42) is not limited to a PID controller, but may also be a classical controller or modern controller based on a dynamic system model, or a model predictive controller (MPC), etc. Furthermore, the actual controller 42 may be implemented using a path-following control method that takes into account the non-holonomic nature of the controlled object 100.

[0092] The simulator executed by the information processing device 2 is not limited to a physical simulator, but may also be a numerical simulator for dynamic system models. Furthermore, the method for identifying dynamic system models may be based on machine learning.

[0093] The controlled object 100 is not limited to forklifts, but may also be a mobile object that moves on the ground, such as the vehicles mentioned above (especially autonomous vehicles) or industrial vehicles, or a mobile object that moves in the air, such as aircraft, helicopters, or drones (i.e., flying objects), or a mobile object that moves on or underwater, such as ships or submarines. Of course, the controlled object 100 is not limited to mobile objects, but may also be a stationary object, such as an industrial robot that performs work in place.

[0094] The information processing device 2 may be on-premise or in a cloud-based configuration. In the case of a cloud-based information processing device 2, for example, the above-mentioned functions and processing may be provided in the form of SaaS (Software as a Service) or cloud computing.

[0095] In the above embodiment, the information processing device 2 performed various storage and control functions, but instead of the information processing device 2, multiple external devices may be used. That is, various information and programs may be stored in a distributed manner across multiple external devices using blockchain technology or the like.

[0096] The above embodiment is not limited to the information processing system 1, but may also be an information processing method or an information processing program. The information processing method includes each step of the information processing system 1. The information processing program causes at least one computer to execute each step of the information processing system 1.

[0097] The above-mentioned information processing system 1, etc., may be provided in any of the following embodiments.

[0098] (1) An information processing system comprising at least one processor capable of executing a program such that the following steps are performed, wherein in the acquisition step, information relating to a controlled object in real space and information relating to a real controller provided by the controlled object are acquired, the real controller is configured to control the autonomous operation of the controlled object, and in the generation step, based on the acquired information relating to the controlled object and the information relating to the real controller, a virtual control system is generated in a virtual space corresponding to the real space to control the operation of a virtual controlled object that simulates the controlled object, the virtual control system comprising a first virtual controller, a second virtual controller, and a mixer, the first virtual controller is configured to control the real controller An information processing system comprising: generating a first control signal capable of controlling the virtual controlled object based on the operation results of the virtual controlled object to simulate the control of the controlled object by the control device; generating a second control signal capable of controlling the controlled object by performing reinforcement learning based on the behavior of the virtual controlled object using a preset reward; a mixer configured to output an input signal to be input to the virtual controlled object by mixing the first control signal and the second control signal; and in the target control step, performing feedback control to the virtual controlled object by inputting the input signal output from the mixer to the virtual controlled object.

[0099] With this configuration, for example, reinforcement learning can be performed on a real controller of a controlled object that is already in operation in the real world, through simulation in a virtual space. Therefore, the effort required to implement a controller with reinforcement learning results suitable for the operating environment of the controlled object in the real world can be reduced.

[0100] (2) An information processing system as described in (1) above, wherein the controlled object is a moving body that moves in the real space, and in the gain setting step, the gain is set within a range in which the control of the virtual control system is stabilized using a second control signal output from the second virtual controller based on the state equation of the virtual control system in the virtual space, the virtual control system further comprises a gain adjuster, the gain adjuster is configured to adjust the second control signal output from the second virtual controller based on the set gain, and the mixer is configured to mix the first control signal and the adjusted second control signal.

[0101] This configuration allows for the stabilization of the virtual control target's behavior in the virtual space, thereby enabling smoother reinforcement learning.

[0102] (3) An information processing system according to (1) or (2) above, wherein the behavior of the virtual controlled object in response to the input signal is configured to be expressible using a low-dimensional state quantity with a number of parameters less than or equal to the degrees of freedom describing the behavior of the virtual controlled object, and a high-dimensional state quantity with a number of parameters greater than the degrees of freedom, the first virtual controller is configured to output the first control signal by inputting the low-dimensional state quantity, and the second virtual controller is configured to perform reinforcement learning based on the low-dimensional state quantity and the high-dimensional state quantity.

[0103] With this configuration, high-dimensional state variables that are difficult to process with the first virtual controller can be processed using the second virtual controller.

[0104] (4) An information processing system as described in (3) above, wherein the low-dimensional state quantity includes a parameter relating to at least one of the position, velocity, and acceleration of the virtual controlled object, and the high-dimensional state quantity includes image data in the virtual space that shows the behavior of the virtual controlled object.

[0105] With this configuration, a more precise virtual control system can be constructed based on visual information within the virtual space.

[0106] (5) An information processing system according to (3) or (4) above, wherein the virtual control system further comprises a parameter estimator for estimating the control parameters of the virtual controlled object, the parameter estimator estimates state variables in a model describing the behavior of the virtual controlled object based on the low-dimensional state variables and the high-dimensional state variables, and the second virtual controller further generates the second control signal based on the estimated state variables.

[0107] With this configuration, reinforcement learning can be performed while taking into account information about the behavior of a virtual controlled object, which is represented by state variables that are difficult to observe directly.

[0108] (6) An information processing system according to any one of (1) to (5) above, wherein the virtual control system further comprises a randomizer, and the randomizer is configured to introduce random noise to the parameters input to the first virtual controller or the second virtual controller.

[0109] This configuration enhances the robustness of the virtual control system against system uncertainty.

[0110] (7) An information processing system according to any one of (1) to (6) above, wherein the controlled object is a moving object that can move in the real space according to a movement plan from its current location to a destination based on a pre-set target path, the real controller is configured to control the behavior of the controlled object based on the set movement plan, in the generation step, a virtual movement plan is generated which is the movement plan of the virtual controlled object in the virtual space based on the movement plan for the controlled object in the real space, the first virtual controller is configured to generate a first control signal based on the deviation between the virtual movement plan and the behavior of the virtual controlled object, and the second virtual controller is configured to perform the reinforcement learning using the reward which is set to reduce the deviation between the virtual movement plan and the behavior of the virtual controlled object.

[0111] With this configuration, for example, reinforcement learning can be performed in a virtual space to control the movement of a moving object, which is difficult to do in real space. This allows for more efficient optimization of the control of the movement of the moving object.

[0112] (8) An information processing method comprising each step of an information processing system described in any one of (1) to (7) above.

[0113] (9) A program that causes at least one computer to perform each step of the information processing system described in any one of (1) to (7) above. Of course, this is not always the case.

[0114] Finally, while various embodiments relating to this disclosure have been described, these are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be implemented in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents. [Explanation of Symbols]

[0115] 1: Information Processing System 2: Information Processing Device 20: Communications bus 21: Communications Department 22: Storage section 23: Processor 3: User terminal 30: Communications bus 31: Communications Department 32: Storage section 33: Processor 34:Display section 35: HMI devices 4: Actual control system 41: Route Planner 42: Actual Controller 5: Virtual control system 51: Virtual Route Planner 52: First virtual controller 53: Second virtual controller 531: Image Feature Extractor 532: Feature Coupler 533: Input signal determination unit 54: Mixer 55: Gain Adjuster 56: Parameter Estimator 57: Randomizer 100: Controlled object 101: Boarding area 102: Cargo handling equipment 103: Control device 200: Virtual control target 201: Virtual Boarding Section 202: Virtual cargo handling equipment D1: Departure point D2: Departure point G1: Destination G2: Destination K: Feedback Gain P1: Current location P2:Current location R1: Target route R2: Target path RS: Real space VS: Virtual Space s11: Target value s12: Measured value s13: Deviation signal s14: Actual input signal s21: Target value s22: First measurement value s23: Deviation signal s24: First control signal s25: Second measurement value s26: Second control signal s27: Gain adjustment signal s28: Estimation result s3: Input signal

Claims

1. An information processing system, The system comprises at least one processor capable of executing a program so that each of the following steps is performed, In the acquisition step, information about the controlled object in real space and information about the actual controller provided by the controlled object are acquired, and the actual controller is configured to control the autonomous operation of the controlled object. In the generation step, based on the acquired information about the controlled object and the information about the actual controller, a virtual control system is generated in the virtual space corresponding to the real space, which controls the operation of a virtual controlled object that simulates the controlled object. The virtual control system comprises a first virtual controller, a second virtual controller, and a mixer. The first virtual controller generates a first control signal capable of controlling the controlled object based on the operation result of the virtual controlled object, in order to simulate the control of the controlled object by the actual controller. The second virtual controller generates a second control signal capable of controlling the virtual controlled object by performing reinforcement learning based on the behavior of the virtual controlled object using a pre-set reward. The mixer is configured to output an input signal to be input to the virtual control target by mixing the first control signal and the second control signal. In the target control step, the information processing system performs feedback control on the virtual control target by inputting the input signal output from the mixer to the virtual control target.

2. In the information processing system described in claim 1, The controlled object is a moving body that moves in the real space, In the gain setting step, the gain is set within a range in which the control of the virtual control system is stabilized using the second control signal output from the second virtual controller based on the state equation of the virtual control system in the virtual space. The virtual control system further includes a gain adjuster, The gain adjuster is configured to adjust the second control signal output from the second virtual controller based on the set gain. The mixer is configured to mix the first control signal with the adjusted second control signal, in an information processing system.

3. In the information processing system described in claim 1, The behavior of the virtual control target in response to the input signal is configured to be expressible using a low-dimensional state quantity with a number of parameters less than or equal to the degrees of freedom describing the behavior of the virtual control target, and a high-dimensional state quantity with a number of parameters greater than the degrees of freedom. The first virtual controller is configured to output the first control signal by inputting the low-dimensional state quantities. The second virtual controller is an information processing system configured to perform reinforcement learning based on the low-dimensional state variables and the high-dimensional state variables.

4. In the information processing system described in claim 3, The low-dimensional state quantity includes a parameter relating to at least one of the position, velocity, and acceleration of the virtual controlled object. The aforementioned high-dimensional state quantities include image data in the virtual space that represents the behavior of the virtual controlled object, and is part of an information processing system.

5. In the information processing system described in claim 3, The virtual control system further includes a parameter estimator for estimating the control parameters of the virtual control target, The parameter estimator estimates the state variables in the model describing the behavior of the virtual controlled object based on the low-dimensional state variables and the high-dimensional state variables. The second virtual controller is an information processing system that generates the second control signal based on the estimated state variables.

6. In the information processing system described in claim 1, The virtual control system further includes a randomizer, An information processing system in which the randomizer is configured to introduce random noise into the parameters input to the first virtual controller or the second virtual controller.

7. In the information processing system described in claim 1, The controlled object is a mobile body capable of moving through the real space according to a movement plan from its current location to its destination based on a pre-set target path. The actual controller is configured to control the behavior of the controlled object based on the set movement plan. In the generation step, a virtual movement plan is generated, which is the movement plan of the virtual controlled object in the virtual space, based on the movement plan of the controlled object in the real space. The first virtual controller is configured to generate the first control signal based on the deviation between the virtual movement plan and the behavior of the virtual controlled object. The second virtual controller is an information processing system configured to perform reinforcement learning using the reward, which is set to reduce the deviation between the virtual movement plan and the behavior of the virtual controlled object.

8. Information processing method, A method comprising each step of the information processing system described in any one of claims 1 to 7.

9. It is a program, A program that causes at least one computer to perform each step of the information processing system described in any one of claims 1 to 7.