Control device, control method, and program
The control device employs deep reinforcement learning to adapt flying device control to individual users, ensuring stable flight regardless of physique or user presence, addressing the inefficiencies of conventional methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- JAPAN AEROSPACE EXPLORATION AGENCY
- Filing Date
- 2022-05-30
- Publication Date
- 2026-05-08
AI Technical Summary
Conventional control methods for wearable flying devices fail to adequately adjust to individual user physiques, requiring frequent reconfiguration and incurring significant time and economic costs.
A control device that uses deep reinforcement learning to learn user-specific control methods, integrating state and operation data to optimize thrust and wing adjustments, enabling stable flight regardless of user physique or presence.
Enables stable flight control across varying user physiques and flight modes, reducing the need for frequent reconfiguration and minimizing operational costs.
Smart Images

Figure 0007855224000001 
Figure 0007855224000002 
Figure 0007855224000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a control device, a control method, and a program. [Background technology]
[0002] Wearable flying devices (flight equipment) that use the thrust of jets or rockets to allow users to fly are known. Such flying devices are also called portable personal air mobility systems. On the other hand, technologies that use deep reinforcement learning to control robots are known (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] XB Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 3803-3810. [Overview of the project] [Problems that the invention aims to solve]
[0004] Because each person has a different physique, if the flight device is not a large device like a helicopter, but rather a device like a suit that is relatively more affected by differences in human physique, then the control method of the flight device needs to be adjusted according to the user wearing it. However, conventional technology has not been able to adequately adjust the control method of the flight device according to the user. Furthermore, the control method had to be readjusted every time the user changed, which resulted in significant time and economic costs.
[0005] The present invention has been made in consideration of such circumstances, and one of its objectives is to provide a control device, a control method, and a program that can suitably control a flying device regardless of the user.
Means for Solving the Problem
[0006] One aspect of the present invention is a control device for controlling a flying device that can be worn by a user. The flying device acquires state data regarding the state of the flying device and operation data regarding the operation of the flying device, inputs the acquired attitude data and operation data to a model learned using deep reinforcement learning, and controls the flying device based on the output result of the model into which the state data and operation data have been input, and includes a processing unit.
Effects of the Invention
[0007] According to one aspect of the present invention, the flying device can be suitably controlled regardless of the user's physique or the presence or absence of the user.
Brief Description of the Drawings
[0008] [Figure 1] It is a diagram for explaining the usage scene of the flying device 1 according to the embodiment. [Figure 2] [[ID=二十七]]It is a diagram showing a configuration example of the flying device 1 according to the embodiment. [Figure 3] It is a diagram showing a configuration example of the control device 100 according to the embodiment [Figure 4] It is a flowchart showing the flow of a series of processes of the processing unit 170. [Figure 5] It is a diagram showing an example of the deep reinforcement learning model MDL.
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments of the control device, control method, and program of the present invention will be described with reference to the drawings.
[0010] [Scenario of using flying devices] Figure 1 is a diagram illustrating a usage scenario of the flying device 1 according to an embodiment. As shown in the figure, the flying device 1 is worn by user U. The flying device 1, worn by user U, can fly under the user U's control or fly autonomously like an autopilot. For example, the flying device 1 is used to travel from departure point A to destination B. After user U, wearing the flying device 1, travels from departure point A to destination B, and then detaches the flying device 1 and lands at destination B, the flying device 1 may continue to hover around destination B until user U puts it on again, or it may return from destination B to departure point A by autonomous flight. The flying device 1 may be used not only by a predetermined single user, but also by an unspecified number of users.
[0011] For example, the flying device 1 may be used by a mountain rescue team to travel by air from a headquarters base (departure point A) located at the foot of a mountain to a rescue site (destination B) on a mountain trail. In this case, after the first rescue team member arrives at destination B, the flying device 1 is detached and the team member lands at destination B. Then, the flying device 1 returns alone to departure point A, allowing the second rescue team member to put on the flying device 1 and head to the rescue site. By repeating this process, multiple rescue team members can be dispatched to destination B using a single flying device 1. Alternatively, after the rescue team members arrive at destination B, the flying device 1 may be detached and the team member lands at destination B. Then, the flying device 1 may travel alone to departure point A or refueling point C, and after refueling at departure point A or refueling point C, the flying device 1 returns alone to destination B. In this case, even if the aircraft only carries enough fuel for a one-way trip from departure point A to destination B, and can only perform a manned flight on the outbound leg, it is possible to perform a manned return flight from destination B to departure point A by refueling using the aircraft 1 alone along the way. In this way, the flight range can also be extended.
[0012] In addition to the above-mentioned uses, the flying device 1 may also be used to transfer a person in need of rescue on the ground to a helicopter waiting in the air. Furthermore, the flying device 1 is not limited to the ground and may also be used at sea. For example, the flying device 1 may be used to transfer a person in distress at sea to a helicopter in the air or a ship at sea.
[0013] [Configuration of Flying Device] FIG. 2 is a diagram showing a configuration example of the flying device 1 according to the embodiment. As shown in the figure, the flying device 1 includes, for example, a thrust device 10, wings 20, a detachable unit 30, and a control device 100.
[0014] Σ shown in FIG. 2 W represents one of the inertial coordinate systems, the Earth-fixed coordinate system Σ W and O W represents the origin of the Earth-fixed coordinate system Σ W The X W axis represents true north, the Y W axis represents east, and the Z W axis represents vertically downward. Also, when the principal inertial axes are defined as the body-fixed coordinate system, the X B axis in the figure represents the principal inertial axis of the aircraft when the center of gravity of the flying device 1 is the origin, the Z B axis represents the downward direction of the aircraft, and the Y B axis represents the direction on the right side of the traveling direction of the aircraft. In other words, the X B axis represents the roll axis, the Z B axis represents the yaw axis, and the Y B axis represents the pitch axis.
[0015] The thrust device 10 generates thrust for the flying device 1 using fuel 11. For the thrust device 10, for example, a known jet engine may be suitably used. Hereinafter, as an example, it will be described assuming that a jet engine with variable thrust is applied to the thrust device 10. A thrust deflection mechanism (for example, a thrust vectoring mechanism having paddles, nozzles, rings, etc.) for switching the direction of the jet flow generated by the duct fan is provided at the jet nozzle of the jet engine, and these thrust deflection mechanisms are controlled by the control device 100.
[0016] The wings 20 maintain the attitude of the flight device 1 and change its direction of flight. The change of direction by the wings 20 may be performed by the user U operating the user interface 120 described later, by the control device 100, or by the cooperation of the user U and the control device 100.
[0017] In this embodiment, the wing 20 is equipped with a link mechanism and can be folded like a bird's feather. The wingspan mentioned above is the wingspan when the wing 20 is extended. Because the wing 20 can be folded, it has the following functions. That is, during high-speed flight, the wing 20 can be folded to reduce air resistance, and during low-speed flight and takeoff / landing, the wing 20 can be fully extended to obtain aerodynamic force. Also, when the flight device 1 is not in use, folding the wing 20 may contribute to maneuverability during transport. Furthermore, the wing 20 may have a structure that allows it to be deployed and stored by having an extendable structure instead of folding. Alternatively, it may be a flat plate shape (i.e., a fixed wing) without a foldable structure. In addition, the wing 20 according to this embodiment is equipped with various actuators in addition to the link mechanism mentioned above, and the roll axis X shown in Figure 2 B , yaw axis Z B , pitch axis Y B It is assumed that it can rotate around. Details will be described later.
[0018] The flying device 1 may also be a wingsuit with fabric stretched between the hands and feet instead of having wings 20, or it may be a fixed-wing device as described above.
[0019] The attachment / detachment section 30 is a component for the user U to attach the flight device 1, and this component has a structure that allows the user U to easily attach and detach it. For example, the attachment / detachment section 30 may have a structure that includes a structure for hanging on the shoulder like a typical backpack and a fastener for fixing it to the user U. Alternatively, each user U may be equipped with an attachment member having a shape corresponding to the attachment / detachment section 30 in advance, and the user U and the attachment / detachment section 30 may be appropriately fixed via the attachment member equipped on the user U.
[0020] The control device 100 controls the thrust of the thruster 10 and the direction of that thrust. Furthermore, the control device 100 adjusts the attitude of the flight device 1 and changes its direction of flight by controlling the shape and orientation of the wings 20.
[0021] [Control device configuration] Figure 3 is a diagram showing an example configuration of the control device 100 according to the embodiment. As shown in the figure, the control device 100 includes, for example, a communication interface 110, a user interface 120, a sensor 130, a power supply 140, a storage unit 150, an actuator 160, and a processing unit 170.
[0022] The communication interface 110 communicates wirelessly with an external device via a network, such as a WAN (Wide Area Network). The external device may be, for example, a remote controller capable of remotely controlling the aircraft 1. For example, the communication interface 110 may receive commands from the external device instructing the aircraft 1 on the attitude and speed of the target it should take. This allows a skilled operator to remotely control the aircraft when the user U's piloting skills are insufficient and autonomous flight by the control unit 230 is impossible.
[0023] Furthermore, the communication interface 110 may receive information from an external device to inform the in-flight user U that destination B has changed, or it may receive information to inform user U of more detailed information about destination B.
[0024] Furthermore, the communication interface 110 may transmit information to an external device. For example, the communication interface 110 may transmit detailed information about the rescue site (such as coordinates and altitude) to an external device.
[0025] The user interface 120 includes an input interface 120a and an output interface 120b. For example, the input interface 120a may be a joystick, handle, button, switch, microphone, etc. The output interface 120b may be a display, speaker, etc. For example, user U may adjust the thrust and direction of the thruster 10, or adjust the shape and direction of the wings 20, by operating the joystick, etc., of the input interface 120a. Alternatively, user U may adjust the thrust and direction of the thruster 10, or adjust the shape and direction of the wings 20, by speaking the desired speed, altitude, attitude, etc., of the aircraft 1 into the microphone of the input interface 120a.
[0026] Sensor 130 is, for example, an inertial measurement device. The inertial measurement device includes, for example, a three-axis accelerometer and a three-axis gyroscope. The inertial measurement device outputs the detected values detected by the three-axis accelerometer and the three-axis gyroscope to the processing unit 170. The detected values from the inertial measurement device include, for example, acceleration and / or angular velocity in the horizontal, vertical, and depth directions, and velocity (rate) in the pitch, roll, and yaw axes. Sensor 130 may further include radar, finders, sonar, GPS (Global Positioning System) receivers, etc.
[0027] The power supply 140 is, for example, a rechargeable battery such as a lithium-ion battery. The power supply 140 supplies power to components such as the actuator 160 and the processing unit 170. The power supply 140 may also include a solar panel or the like.
[0028] Furthermore, the actuator 160 and the processing unit 170 may, instead of using the power supplied from the power supply 140, or in addition to using the power generated by the jet engine of the thrust device 10.
[0029] The storage unit 150 is implemented using a storage device such as an HDD (Hard Disk Drive), flash memory, EEPROM (Electrically Erasable Programmable Read Only Memory), ROM (Read Only Memory), or RAM (Random Access Memory). In addition to various programs such as firmware and application programs, the storage unit 150 stores calculation results from the processing unit 170 as logs. The storage unit 150 also stores model information 152. The model information 152 may be installed into the storage unit 150 from an external device via a network, for example, or from a portable storage medium connected to the drive device of the control device 100. The model information 152 will be described later.
[0030] The actuator 160 includes, for example, a thrust actuator 162, a sweep actuator 164, and a fold actuator 168.
[0031] The thrust actuator 162 drives the thrust device 10 to provide thrust to the flight device 1 and change the direction of that thrust. The sweep actuator 164 controls the yaw axis Z B The wings 20 rotate around it.
[0032] The processing unit 170 is implemented, for example, by a CPU (Central Processing Unit) or GPU (Graphics Processing Unit) executing a program stored in the memory unit 150. The processing unit 170 may also be implemented by hardware such as an LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), or FPGA (Field-Programmable Gate Array), or by the collaboration of software and hardware.
[0033] The processing unit 170 controls the thrust actuator 162 based on some or all of the following: (i) user U input operations to the input interface 120a, (ii) detection results from the sensor 130, and (iii) remote control commands received by the communication interface 110 from an external device. This controls the thrust of the thrust device 10 and the direction of that thrust. For example, the control device 100 controls the thrust actuator 162 to adjust the thrust by controlling the rotation speed of the duct fan of the jet engine of the thrust device 10, or to adjust the thrust direction by controlling the thrust deflection mechanism of the jet engine.
[0034] Furthermore, if the wing 20 is a variable-geometry wing, the control device 100 controls the sweep actuator 164 and the fold actuator 168 based on some or all of (i) to (iii). This controls the shape and orientation of the wing 20. The shape and orientation of the wing 20 are examples of the "variable-geometry wing control amount".
[0035] [Processing flow of the processing unit] The following describes the sequence of processes in the processing unit 170 using a flowchart. Figure 4 is a flowchart of the sequence of processes in the processing unit 170. The processes in this flowchart may be repeated, for example, at a predetermined interval.
[0036] First, the processing unit 170 processes a state variable s that indicates the state of the environment surrounding the flying device 1 at the current time t. t Obtain (step S100). State variable s t This includes, for example, at least one (preferably all) of the attitude, position, velocity, and angular velocity of the flying device 1 at the current time t. For example, state variables s t The angle included may be the angle around the pitch axis (hereinafter referred to as the pitch angle). Also, the state variable s t The angular velocity included may be the angular velocity of the pitch angle. Furthermore, the state variable s tThis may include the thrust of the thruster 10 and its direction at the current time t, and the shape and orientation of the wing 20 at the current time t. Attitude, position, velocity, and angular velocity at the current time t are examples of "state data," one or all of which are examples of "state data." The thrust of the thruster 10 and its direction at the current time t, and the shape and orientation of the wing 20 at the current time t are examples of "operation data."
[0037] For example, the processing unit 170 obtains attitude, position, velocity, and angular velocity from the sensor 130 as state variables s t It will be acquired as follows.
[0038] Furthermore, when user U instructs the thrust or direction of the thrust device 10 via the input interface 120a, the processing unit 170 records the user U's input operation to the input interface 120a as a state variable s t It may also be added.
[0039] Next, the processing unit 170 reads the model information 152 from the memory unit 150 and uses the deep reinforcement learning model MDL defined by the model information 152 to process the state variable s t Therefore, the optimal action (action variable) that the flying device 1 can take at the next time t+1 is a t+1 Determine (step S102).
[0040] Behavior (behavioral variable) a in this embodiment t+1 This refers to the actions required to accomplish the desired task, and may include, for example, the thrust and direction of the thruster 10 necessary to accomplish the task, and may also include the shape and orientation of the wing 20. The desired task may be a variety of tasks, such as keeping the flight device 1 hovering at a certain altitude, smoothly transitioning from horizontal flight to a hovering position, or flying straight even in strong winds.
[0041] Figure 5 shows an example of a deep reinforcement learning model MDL. The deep reinforcement learning model MDL according to this embodiment is a neural network that utilizes deep reinforcement learning. As shown in the figure, for example, the deep reinforcement learning model MDL may be a recurrent neural network in which part of the intermediate layer (hidden layer) is an LSTM (Long Short-Term Memory). The deep reinforcement learning model MDL is learned by randomly setting the dynamics of the flying device 1, such as its weight, center of gravity, and moment of inertia, and the system response delay, using Domain-Randomization.
[0042] When learning is performed using domain-randomization (where the dynamics of flying device 1 are randomized), the LSTM of the deep reinforcement learning model MDL stores a time series that reflects the randomly set dynamics of flying device 1. In this way, by incorporating an LSTM into the neural network, learning using domain-randomization is optimally performed.
[0043] For example, if the deep reinforcement learning algorithm used to train the deep reinforcement learning model MDL is value-based, the deep reinforcement learning model MDL may be trained using a DQN (Deep Q-Network). DQN is a type of reinforcement learning called Q-learning, where the state s of a certain environment at a given time t is used. t Under certain action a t The action-value function Q(s) represents the value of making a choice as a function. t a t This method involves training a neural network with an approximation function based on the value-based approach. In other words, the deep reinforcement learning model MDL, trained using a value-based method, learns one or more possible actions (action variables) that the flying device 1 can take at the current time t. t Among these, the action (behavioral variable) that maximizes the value (Q value) is a t It is acceptable for the system to be trained to output the following.
[0044] Q-learning learns the weights and biases of the deep reinforcement learning model MDL by, for example, giving a high reward when the wing 20 or thruster 10 is in an ideal state. For example, when the aircraft 1 is in a 90-degree pitch-up attitude above a designated point and its speed is at a speed that can be considered stationary, the reward may be high. On the other hand, when the aircraft 1 is in contact with the ground or trees, or deviates from the designated altitude, the reward may be low (e.g., zero).
[0045] Furthermore, for example, if the deep reinforcement learning algorithm used to train the deep reinforcement learning model MDL is policy-based, the deep reinforcement learning model MDL may be trained using methods such as policy gradients.
[0046] Furthermore, for example, if the deep reinforcement learning algorithm used to train the deep reinforcement learning model MDL is an Actor-Critic algorithm that combines value and policy, then while training the Actors (Action Units) included in the deep reinforcement learning model MDL, the Critics (Evaluators) that evaluate the policies may be trained simultaneously. The deep reinforcement learning model MDL exemplified in Figure 5 is a model trained using an Actor-Critic algorithm such as PPO (Proximal Policy Optimization), where the upper layer is trained to output policies and the lower layer is trained to output value.
[0047] The model information 152 defining such a deep reinforcement learning model MDL includes various information, such as connection information describing how the units in each of the multiple layers constituting the neural network are connected to each other, and connection coefficients assigned to the data input and output between connected units. Connection information includes, for example, the number of units in each layer, information specifying the type of unit to which each unit is connected, the activation function that realizes each unit, and information such as gates provided between units in the hidden layer. The activation function that realizes a unit may be, for example, a normalized linear function (ReLU function), a sigmoid function, a step function, or other functions. Gates selectively pass or weight data transmitted between units according to the value returned by the activation function (e.g., 1 or 0). Connection coefficients include, for example, the weights assigned to the output data when data is output from a unit in one layer to a unit in a deeper layer in the hidden layer of the neural network. Connection coefficients may also include bias components unique to each layer. Furthermore, the model information 152 may include information specifying the type of activation function for each gate in the LSTM, as well as recurrent weights and peep-hole weights.
[0048] For example, when the processing unit 170 obtains at least one of the attitude, position, velocity, and angular velocity of the flight device 1 at the current time t, and the thrust and direction of the thrust device 10 at the current time t, it stores them in state variables s t This is input to the deep reinforcement learning model MDL. State variable s t The deep reinforcement learning model MDL, upon receiving the input, outputs the optimal thrust and direction of the thruster 10 at the next time t+1. As described above, the deep reinforcement learning model MDL may be trained to output, in addition to or instead of, the thrust and direction that the thruster 10 should output at the next time t+1, the shape and orientation that the wing 20 should take at the next time t+1.
[0049] Let's return to the explanation of the flowchart in Figure 4. Next, the processing unit 170 determines the action (action variable) that the flying device 1 should take, using the deep reinforcement learning model MDL. t+1 In other words, based on the thrust that the thrust device 10 should output and its direction at the next time t+1, and the shape and orientation that the wing 20 should take at the next time t+1, control commands are generated to control the actuator 160 of the flight device 1 (step S104).
[0050] For example, the processing unit 170 uses the deep reinforcement learning model MDL to determine the behavioral variable a t+1 Based on the thrust and direction of the thrust device 10 output as a control command for the thrust actuator 162, the processing unit 170 may generate a control command for the thrust actuator 162. In addition, the processing unit 170 generates a control command for the action variable a t+1 Based on the shape and orientation of the wing 20 output, control commands for the sweep actuator 164 and the fold actuator 168 may be generated.
[0051] Next, the processing unit 170 controls the actuator 160 based on the generated control command (step S106). This achieves the desired task, and as a result the state of the environment surrounding the flying device 1 changes, and the state variable representing that state becomes s t ga s t+1 It changes to this.
[0052] The processing unit 170 processes the state variable s t ga s t+1 As a result of this change, the state variable s at time t+1 t+1 The state variable s at time t+1 is retrieved again. Then, the processing unit 170 retrieves the state variable s at time t+1. t+1 However, control commands are continuously issued to the target actuator 160 so that the desired task continues to be achieved by the flying device 1. This completes the processing of this flowchart.
[0053] According to the first embodiment described above, the processing unit 170 of the control device 100 determines at least one (preferably all) of the attitude, position, velocity, and angular velocity of the flight device 1 at the current time t, and the thrust and direction of the thrust device 10 at the current time t, using state variables s t The data is obtained as follows. At this time, the processing unit 170 obtains, in addition to or instead of, the thrust of the thrust device 10 and its direction at the current time t, the shape and direction of the wing 20 at the current time t, using the state variable s t It may be acquired as such.
[0054] The processing unit 170 processes the state variable s t When obtained, the state variable s is applied to the deep reinforcement learning model MDL, which has been pre-trained by deep reinforcement learning. t Input the following. Processing unit 170 processes the state variable s t The action variable a at the next time step t+1, output by the deep reinforcement learning model MDL in response to the input. t+1 The flying device 1 is controlled based on the following: The state variables s include the attitude, position, velocity, and angular velocity of the flying device 1 at the current time t, and the thrust and direction of the thruster 10 at the current time t. t Because the flight device 1 is controlled using the deep reinforcement learning model MDL, which is trained using deep reinforcement learning based on the above, in the case of manned flight, even if there is variation in the physical size (weight, height, etc.) of the user U wearing the flight device 1, the flight device 1 can be controlled appropriately regardless of the user U's physical size. Furthermore, even if the user detaches from the flight device 1 during flight and the system switches from manned flight to unmanned flight, the flight device 1 can still be controlled appropriately.
[0055] For example, as mentioned above, if a mountain rescue team is heading to a rescue site (destination B) on a mountain trail by air while wearing the flying device 1, it is assumed that after the first rescuer arrives at destination B, they will detach the flying device 1 and land at destination B, and then the flying device 1 will return to the departure point A on its own, allowing the second rescuer to wear the flying device 1 and head to the rescue site. In such a case, if the physiques of the first and second rescuers differ significantly, it would be difficult to use the same flying device 1 with conventional technology. In contrast, in this embodiment, since the recurrent neural network has an LSTM layer, time-series memory becomes possible, and tuning according to the user's physique becomes possible from the history of control output and state variables. As a result, the flying device 1 can be kept flying stably in the same way whether the user U, who is heavy, is wearing the flying device 1 or the user U, who is light.
[0056] Furthermore, for example, if the first rescue worker detaches the flying device 1 and lands at destination B after arriving at destination B, the load on the flying device 1 will decrease rapidly. In such a case, conventional technology makes it difficult to keep the flying device 1 flying stably. In contrast, in this embodiment, deep reinforcement learning is performed considering the dynamics and response delay variability of the flying device 1, rather than the user's physique. In other words, deep reinforcement learning is performed using domain randomization. Therefore, even if user U detaches from the flying device 1 and the flying device 1 remains on its own, the flying device 1 can continue to fly stably, just as it did when user U was wearing the flying device 1.
[0057] Although embodiments for carrying out the present invention have been described above using examples, the present invention is not limited in any way to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention. [Explanation of Symbols]
[0058] 1...Flight device, 10...Thrust device, 20...Wings, 30...Detachable parts, 100...Control device, 110...Communication interface, 120...User interface, 130...Sensor, 140...Power supply, 150...Memory unit, 160...Actuator, 170...Processing unit
Claims
1. A control device for controlling a user-worn flying device, A user interface on which the user's operations on the aforementioned flying device are input, The aforementioned flying device includes a thrust source that generates thrust, The system comprises a processing unit for controlling the aforementioned flying device, The aforementioned processing unit, The system acquires state data relating to the state of the aircraft and operation data relating to the user's operation of the aircraft input to the user interface, which includes the thrust of the thrust source and the direction of the thrust. The acquired state data and operation data are input to a model trained using deep reinforcement learning. Based on the thrust and direction of the thrust output by the model in response to the input of the state data and operation data, the attitude of the flight device is controlled. Control device.
2. The aforementioned model is a neural network trained by domain randomization. The control device according to claim 1.
3. The aforementioned model is a recurrent neural network that includes a memory layer. The control device according to claim 1 or 2.
4. The thrust source includes a jet engine, The state data includes at least one of the attitude, position, speed, and angular velocity of the aircraft. The aforementioned operation data includes the thrust of the jet engine and the direction of the thrust. The processing unit controls the attitude of the aircraft based on the thrust and the direction of the thrust output by the model. The control device according to claim 1 or 2.
5. The aforementioned flight device further includes variable wings, The aforementioned operation data further includes the amount of control of the variable wing, The processing unit controls the attitude of the flight device based on the thrust, the direction of the thrust, and the amount of control of the variable wing output by the model. The control device according to claim 4.
6. A control method for controlling a user-worn flying device using a computer, The acquisition of state data relating to the state of the aircraft and operation data relating to the user's operation of the aircraft input into the user interface, which includes the thrust of a thrust source that generates thrust for the aircraft and the direction of the thrust. The acquired state data and operation data are input to a model trained using deep reinforcement learning. Based on the thrust and direction of the thrust output by the model in response to the input of the state data and operation data, the attitude of the flight device is controlled. A control method including
7. A program that causes a computer to execute commands to control a user-worn flying device, The acquisition of state data relating to the state of the aircraft and operation data relating to the user's operation of the aircraft input into the user interface, which includes the thrust of a thrust source that generates thrust for the aircraft and the direction of the thrust. The acquired state data and operation data are input to a model trained using deep reinforcement learning. Based on the thrust and direction of the thrust output by the model in response to the input of the state data and operation data, the attitude of the flight device is controlled. A program that includes this.
Citation Information
Patent Citations
Morphing wing, flight controller, flight control method, and program
JP2021030950A
Control device, learning device, control method, learning method, and program
JP2021049841A
Flight apparatus and operation method
JP2023009892A
Aircraft attachable to the body of a pilot
US4253625A
Lift system intended for free-falling persons
US6685135B2