Flapping drone, autonomous flight control method for flapping drone, and autonomous flight controller performing same
By employing wing deformation data from a strain sensor for feature extraction and reinforcement learning, the autonomous flight control method allows flapping drones to navigate complex environments autonomously, addressing the limitations of existing control methods.
Patent Information
- Application Number
- PCT/KR2024/015665
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-26
AI Technical Summary
Existing control methods for flapping drones struggle to respond quickly to external flow changes, limiting their ability to fly autonomously in various environments, especially in windy conditions, without relying on external sensors.
The implementation of an autonomous flight control method that utilizes wing deformation data from a strain sensor to extract external and internal features of the drone, enabling reinforcement learning to determine an optimal flight policy that allows the drone to reach target positions without external sensors.
This solution enables flapping drones to perform autonomous flight stably and respond to external flow changes, achieving efficient navigation in complex environments with reduced computational costs and control complexity.
Smart Images

Figure KR2024015665_26062025_PF_FP_ABST
Abstract
Description
Flapping drone, autonomous flight control method for flapping drone, and autonomous flight controller performing the same
[0001] The present invention relates to a flapping drone, a method for controlling autonomous flight of a flapping drone, and an autonomous flight controller for performing the same. The present invention relates to a flapping drone capable of autonomous flight that can rapidly respond to external flow changes and stably fly autonomously, a method for controlling autonomous flight of a flapping drone, and an autonomous flight controller for performing the same.
[0002] In recent years, the drone industry has seen exponential growth in its market size, with applications spanning not only military missions but also agriculture, cargo transport, and leisure. Currently, widely used commercial drones can be broadly categorized into fixed-wing and rotary-wing types.
[0003] First, fixed-wing drones are advantageous for long-distance flight using their own power and drag, but they have limitations in attitude control such as hovering and sharp turns.
[0004] On the other hand, rotary wing drones can hover in the air by generating lift through the rotational movement of their wings, but they have the disadvantage of causing injury when they come into contact with people or objects, and generating noise when their wings rotate.
[0005] Meanwhile, insect-sized flapping-wing Micro Air Vehicles (FWMAVs) have the advantages of greater agility and maneuverability than fixed-wing and rotary-wing drones, lower noise, less impact when in contact with objects, making them easier to control in narrow spaces, and being safe to fly near people.
[0006] These advantages highlight the potential of flapping drones, especially in areas and environments where conventional commercial drones are difficult to use, such as disaster relief exploration in confined terrain and environmental monitoring such as forests or jungles.
[0007] However, controlling a flapping drone in windy conditions is difficult due to limitations in the control devices and control technology.
[0008] Specifically, the flapping drone performs various maneuvers by independently controlling the stroke amplitude, pitch angle, and phase delay of the wings, where the movement of the wings creates vortices in the fluid flow and diversifies the wing shape.
[0009] Therefore, complex interactions between the fluid and the structure (wing) due to fluid dynamics, aeroelasticity, and flight dynamics occur, causing the wing to deform nonlinearly.
[0010] For this reason, even simplified computational models require significant computational resources and each simulation takes a long time to converge.
[0011] Research has been reported on controlling flapping drones, but the aircraft is controlled through feedback control based on position and direction information using separate external devices such as motion capture cameras within a limited space.
[0012] However, this type of control method has the disadvantage of being unable to quickly respond to external flow changes and therefore can only fly in limited spaces. Therefore, the development of a control system that can compensate for this problem is required.
[0013] The present invention has been made to solve the above problems, and the purpose of the present invention is to provide a flapping drone that enables autonomous flight using only wing deformation information of the flapping drone without using separate external devices such as sensors such as vision sensors, accelerometers, and gyro sensors and motion capture cameras, an autonomous flight control method for the flapping drone, and an autonomous flight controller that performs the same.
[0014] In order to achieve the above object, a method for controlling autonomous flight of a flapping drone in an autonomous flight controller for controlling autonomous flight of a flapping drone according to an embodiment of the present invention comprises the steps of: collecting wing deformation data according to the flight of the flapping drone from a strain sensor attached to the wing of the flapping drone; extracting external features and internal features of the flapping drone from the wing deformation data; finding an autonomous flight policy that allows the flapping drone to reach a target position using the external features and internal features, and performing reinforcement learning to update the autonomous flight policy that provides the highest reward; and controlling the flapping drone according to an autonomous flight behavior determined by the autonomous flight policy.
[0015] And the external characteristic may include at least one of a wind direction and a wind speed when the flapping drone is in flight, and the internal characteristic may include at least one of a posture and a wing flapping speed of the flapping drone.
[0016] In addition, in the step of extracting the above features, the external features can be extracted using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to the wind direction and wind speed.
[0017] And in the step of performing the above reinforcement learning, the reinforcement learning can be performed according to a problem definition set in advance to generate observation information including the extracted external features and configure status information of the flapping drone according to the generated observation information.
[0018] In addition, in the step of performing the above reinforcement learning, the behavior of the flapping drone can be learned according to the following mathematical formula.
[0019] [Mathematical formula]
[0020]
[0021]
[0022] And the autonomous flight control method of the flapping drone may further include a step of collecting, in a replay buffer, observation information, which is the wing deformation data, and a transition including the behavior of the flapping drone at a certain time after the step of controlling the flapping drone.
[0023] Additionally, in the step of performing the reinforcement learning, the autonomous flight policy can be updated based on data collected in the replay buffer corresponding to the current state information of the flapping drone.
[0024] Meanwhile, according to an embodiment of the present invention for achieving the above object, a flapping drone capable of autonomous flight includes: a driving device that provides thrust to the flapping drone and controls a flight direction; a strain sensor attached to a wing of the flapping drone; and an autonomous flight controller that collects wing deformation data according to flight from the strain sensor, extracts external and internal features of the flapping drone from the wing deformation data, finds an autonomous flight policy that allows the flapping drone to reach a target location using the external and internal features, performs reinforcement learning to find the autonomous flight policy that provides the highest reward, and controls the driving device so that the flapping drone can fly according to an autonomous flight behavior determined by the autonomous flight policy.
[0025] Meanwhile, in order to achieve the above object, an autonomous flight controller for controlling a driving device provided in a flapping drone according to an embodiment of the present invention includes a collection unit for collecting wing deformation data according to a flight of the flapping drone from a strain sensor attached to the wing of the flapping drone; a feature extraction unit for extracting external features and internal features of the flapping drone from the wing deformation data; an update unit for finding an autonomous flight policy that allows the flapping drone to reach a target position using the external features and internal features, and performing reinforcement learning to update the autonomous flight policy that provides the highest reward; and a flight control unit for controlling the driving device so that the flapping drone can fly according to an autonomous flight behavior determined by the autonomous flight policy.
[0026] And the external characteristic may include at least one of a wind direction and a wind speed when the flapping drone is in flight, and the internal characteristic may include at least one of a posture and a wing flapping speed of the flapping drone.
[0027] In addition, the feature extraction unit can extract the external features using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to the wind direction and wind speed.
[0028] And the above update unit can perform the reinforcement learning according to a problem definition set in advance to generate observation information including the extracted external features and configure status information of the flapping drone according to the generated observation information.
[0029] In addition, the above update unit can learn the behavior of the flapping drone according to the following mathematical formula.
[0030] [Mathematical formula]
[0031]
[0032]
[0033] And the autonomous flight controller may further include a replay buffer that collects observation information, which is wing deformation data at a certain time after controlling the driving device, and a transition including the behavior of the flapping drone.
[0034] Additionally, the update unit can update the autonomous flight policy based on data collected in the replay buffer corresponding to the current status information of the flapping drone.
[0035] According to one aspect of the present invention described above, by providing a flapping drone, a method for controlling autonomous flight of a flapping drone, and an autonomous flight controller for performing the same, autonomous flight can be enabled using only wing deformation information of the flapping drone without using separate external devices such as sensors such as vision sensors, accelerometers, and gyro sensors, and motion capture cameras.
[0036] Figure 1 is a block diagram for explaining the configuration of a flapping drone according to one embodiment of the present invention.
[0037] FIG. 2 is a drawing schematically illustrating a learning process for controlling autonomous flight of a flapping drone according to one embodiment of the present invention;
[0038] Figures 3 and 4 are drawings for explaining a strain sensor according to one embodiment of the present invention.
[0039] FIG. 5 is a drawing for explaining a strain sensor and a driving device according to one embodiment of the present invention;
[0040] FIG. 6 is a drawing for comparing and explaining the autonomous flight process of a flapping drone according to one embodiment of the present invention with the autonomous flight process of a conventional flapping drone.
[0041] FIG. 7 is a drawing for explaining the configuration of an autonomous flight controller according to one embodiment of the present invention;
[0042] Figure 8 is a flowchart for explaining an autonomous flight method according to one embodiment of the present invention, and
[0043] FIGS. 9 to 13 are drawings for explaining the experimental process and experimental results for verifying the performance of a flapping drone according to one embodiment of the present invention.
[0044] The following detailed description of the present invention refers to the accompanying drawings, which illustrate specific embodiments in which the present invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present invention. It should be understood that the various embodiments of the present invention, while different from each other, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the positions or arrangements of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the present invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the present invention is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled, if properly described. Like reference numerals in the drawings designate the same or similar functionality throughout the several aspects.
[0045] The components according to the present invention are defined by functional distinctions rather than physical distinctions, and can be defined by the functions each component performs. Each component may be implemented as hardware or program code and processing units that perform each function, and the functions of two or more components may be implemented by including them in a single component. Therefore, the names given to the components in the following embodiments are not intended to physically distinguish each component, but rather to suggest the representative functions performed by each component, and it should be noted that the technical spirit of the present invention is not limited by the names of the components.
[0046] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings.
[0047] FIG. 1 is a block diagram for explaining the configuration of a flapping drone (10) according to one embodiment of the present invention, and FIG. 2 is a diagram for schematically explaining a learning process for controlling autonomous flight of a flapping drone (10) according to one embodiment of the present invention.
[0048] The flapping drone (10, hereinafter referred to as the drone) according to the present embodiment can quickly respond to external flow changes and fly stably even in environments where various types of wind occur.
[0049] To this end, the drone (10) according to the present embodiment can perform repetitive training under complex air flow conditions to control the rotation angle and degree of freedom as illustrated in FIG. 2 and control autonomous flight based on reinforcement learning.
[0050] And the drone (10) according to the present embodiment includes a strain sensor (100), an autonomous flight controller (200), and a driving device (300).
[0051] In addition, the drone (10) can be installed and run with software (application) for performing an autonomous flight control method, and the strain sensor (100), autonomous flight controller (200), and driving device (300) can be controlled by the software (application) for performing an autonomous flight control method.
[0052] The strain sensor (100) according to the present embodiment can measure the aeroelastic load generated on the wing during the flight of the drone (10) to generate wing deformation data and transmit it to the autonomous flight controller (200).
[0053] And a strain sensor (100) can be attached to each wing provided on the drone (10) to measure the aeroelastic load.
[0054] FIGS. 3 and 4 are drawings for explaining a strain sensor (100) according to one embodiment of the present invention, and FIG. 5 is a drawing for explaining a strain sensor (100) and a driving device (300) according to one embodiment of the present invention.
[0055] The strain sensor (100) according to the present embodiment may be based on the function of campaniform sensilla, which can quickly adapt to complex airflow conditions and immediately respond to sudden changes by mimicking the agile flight ability of insects.
[0056] As shown in Figure 3, the saccadic sensory organ is a sensory organ distributed on the bottom of the wing and can detect deformation of insect wings. The spike rate of the neuron varies depending on the flapping motion of the wing, and the saccadic sensory organ can detect rotational speed during flight.
[0057] The strain sensor (100) according to the present embodiment imitates the sensory system of an insect, such as a saccular sensory organ, and can be provided as an ultrasensitive crack-based strain sensor.
[0058] More specifically, the strain sensor (100) of the present invention can be produced through the following process.
[0059] First, a shadow mask was covered on a 7.5㎛ polyimide (PI) film substrate, and then a 50nm chromium and 20nm gold thin film was deposited using a thermal evaporator at a thickness of 0.5 It can be deposited at a rate of / s.
[0060] After a 2% tensile test using a tensile tester, nanoscale cracks are generated on the thin film surface. These cracks can be regularly aligned at 3 μm intervals. Therefore, if the wing is physically deformed, the gaps between the cracks widen, resulting in a significant change in electrical resistance.
[0061] At this time, the strain sensor (100) has a high sensitivity of about a gauge factor (GF) of 30,000 at a strain of 2%. The crack-based strain sensor (100) thus generated can be mounted on the flapping drone wing.
[0062] When attaching a strain sensor (100) to the wing of a drone (10), a PDMS dry transfer method can be used on the thin wing root. Additionally, conductive epoxy can be used to electrically connect the sensor (100) and the wire, and a commercial epoxy adhesive can be applied to reduce deformation of the portion connected to the wire.
[0063] Accordingly, the strain sensor (100) according to the present embodiment can precisely and stably measure wing deformation data, which is the trajectory change occurring in the wing and the aerodynamic force acting on the wing due to the long-term flight of the drone (10), and transmit the data to the autonomous flight controller (200).
[0064] Meanwhile, the driving device (300) is provided to provide thrust to the drone (10) and control the flight direction, and can receive a signal for the flight of the drone (10) from the autonomous flight controller (200).
[0065] The driving device (300) according to the present embodiment may be provided as a motor, for example, and may be provided as a first motor provided on the body of the drone (10) and a second motor provided on the tail of the drone (10) as shown in FIG. 5.
[0066] Accordingly, the driving device (300) can enable the drone (10) to fly autonomously by controlling the first motor and the second motor according to the signal received from the autonomous flight controller (200).
[0067] More specifically, the first motor provided on the body of the drone (10) can be combined with a Scotch Yoke mechanism that converts rotational motion into flapping motion to generate thrust.
[0068] The second motor equipped on the tail of the drone (10) can control the direction of the drone (10) by controlling the tail wing. When a positive voltage is applied to the second motor, the tail wing rotates clockwise, allowing the drone (10) to rotate counterclockwise.
[0069] Meanwhile, the autonomous flight controller (200) according to the present embodiment can collect wing deformation data according to flight from the strain sensor (100) and extract external and internal features of the drone (10) from the wing deformation data.
[0070] And the autonomous flight controller (200) can perform reinforcement learning to find an autonomous flight policy that allows the drone (10) to reach a target location using external and internal features, and to find an autonomous flight policy that provides the highest reward.
[0071] Additionally, the autonomous flight controller (200) can control the driving device (300) so that the drone (10) can fly according to the autonomous flight behavior determined by the autonomous flight policy.
[0072] FIG. 6 is a drawing for comparing and explaining the autonomous flight process of a flapping drone (10) according to one embodiment of the present invention with the autonomous flight process of a conventional flapping drone.
[0073] Unlike conventional flight control systems that use sensors such as vision systems, gyroscopes, and accelerometers in autonomous flight of conventional flapping drones, the drone (10) according to the present embodiment can reduce computational costs and control complexity by using only a strain sensor (100).
[0074] Below, the autonomous flight controller (200) according to the present embodiment will be described in more detail.
[0075] FIG. 7 is a drawing for explaining the configuration of an autonomous flight controller (200) according to one embodiment of the present invention, wherein the autonomous flight controller (200, hereinafter referred to as the controller) controls a driving device (300) to enable a drone (10) to fly autonomously and stably.
[0076] For this purpose, the autonomous flight controller (200) may include a collection unit (210), a feature extraction unit (230), an update unit (250), a flight control unit (270), and a replay buffer (290).
[0077] At this time, the controller (200) may be a separate terminal or a module of the terminal. In addition, the configuration of the collection unit (210), feature extraction unit (230), update unit (250), flight control unit (270), and replay buffer (290) may be formed as an integrated module or may be composed of one or more modules. However, conversely, each configuration may be composed of a separate module.
[0078] In addition, the controller (200) may be mobile or fixed. The controller (200) may be in the form of a server or an engine, and may be called by other terms such as a device, an apparatus, a terminal, a UE (user equipment), an MS (mobile station), a wireless device, a handheld device, etc. In addition, the device (100) may execute or produce various software based on an operating system (OS), that is, a system. Here, the operating system is a system program for enabling software to use the hardware of the device, and may include all mobile computer operating systems such as Android OS, iOS, Windows Mobile OS, Bada OS, Symbian OS, and Blackberry OS, as well as computer operating systems such as Windows series, Linux series, Unix series, MAC, AIX, and HP-UX.
[0079] First, a collection unit (210) can be provided to collect wing deformation data according to the flight of a drone (10) from a strain sensor (100).
[0080] And the collection unit (210) can transmit the collected wing deformation data to at least one of the feature extraction unit (230) and the replay buffer (290).
[0081] Additionally, the collection unit (210) can collect wing deformation data from the strain sensor (100) using DAQ. In addition, the collection unit (210) can perform preprocessing by using a band-pass filter in the range of 0.1 to 55 Hz and normalizing the data to 0 to 1 to improve the efficiency of reinforcement learning.
[0082] The feature extraction unit (230) may be provided to extract external and internal features of the drone (10) from wing deformation data received from the collection unit (210).
[0083] Here, the external features extracted by the feature extraction unit (230) may include at least one of the wind direction and wind speed when the drone (10) is flying.
[0084] And the internal features extracted by the feature extraction unit (230) may include at least one of the attitude of the drone (10) and the wing flapping speed.
[0085] In addition, the feature extraction unit (230) according to the present embodiment can extract external features using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to wind direction and wind speed.
[0086] Below, the process of learning a one-dimensional convolutional neural network model so that the feature extraction unit (230) can identify the deformation relationship of the wing according to the wind direction and wind speed will be described in more detail.
[0087] First, to understand the relationship between various flow environments and wing deformation, 62 wind directions and wind speeds of 3 m / s, 5 m / s, and 7 m / s are combined to generate 186 different wind types.
[0088] And the drone (10) is fixed to a three degree of freedom system with three motors along the x, y and z axes and rotated in various combinations of yaw, pitch and roll angles in 186 different wind types.
[0089] And to classify 186 winds, we use a neural network consisting of two 1D convolution layers and two fully connected layers.
[0090] Each convolutional layer contains five filters of size 50, for a total of 64 filters. The ReLU activation function and a max pooling layer are applied to the output of each convolutional layer.
[0091] For fully connected layers, dropout can be applied to prevent overfitting. The final output can be transformed into a probability distribution using the softmax function, and the loss function can be cross-entropy.
[0092] The feature extraction unit (230) can input wing deformation data collected from a strain sensor (100) of 0.2 seconds into a one-dimensional convolutional neural network model including such a configuration.
[0093] Afterwards, when the feature extraction unit (230) receives wing deformation data of a drone (10) performing free flight rather than a fixed state drone (10), it can classify the current wind direction and wind speed using the learned one-dimensional convolutional neural network model to extract external features.
[0094] Meanwhile, the update unit (250) may perform reinforcement learning to find an autonomous flight policy that allows the drone (10) to reach the target location using external and internal features, and update the autonomous flight policy that provides the highest reward. At this time, the target location of the drone (10) may be specified as x,y coordinates.
[0095] The drone (10) flies according to changes in the output of the driving device (300) including the first motor and the second motor. The movement of the drone (10) can be divided into four cases: clockwise rotation, counterclockwise rotation, acceleration, and deceleration.
[0096] Accordingly, the controller (200) can be said to control two degrees of freedom (DOF) with the first motor and the second motor, so that the drone (10) flies with six degrees of freedom.
[0097] And the update unit (250) can receive location information of the drone (10) from the motion capture camera to provide a reward for reinforcement learning, and the reward function can use a two-dimensional multivariate probability distribution function and input the target location as an average.
[0098] Accordingly, the update unit (250) can perform reinforcement learning to determine the optimal flight path with the highest reward by adjusting the output of the driving device (300) to finally reach the target position.
[0099] Additionally, the update unit (250) can update the autonomous flight policy based on data collected in the replay buffer (290) corresponding to the current status information of the drone (10).
[0100] Below, the reinforcement learning performed by the update unit (250) will be described in detail.
[0101] The update unit (250) can perform reinforcement learning according to a problem definition set in advance.
[0102] The problem definition can be defined as a partially observable Markov decision process (POMDP) (the system is a Partially Observable Markov Decision Process).
[0103] A POMDP is a state s, a set of observations , a set of conditional observation probabilities O and a discount factor A set of states S when action a is selected, a set of possible actions A, a reward function r(s,a), and a probability distribution of the next state. may include.
[0104] In this POMDP, the agent that decides the work, i.e. the controller (200), is the policy You can choose the task accordingly.
[0105] The goal of reinforcement learning performed by the update unit (250) is given A policy that maximizes the expected sum of discounted rewards is to find the parameter Through a neural network with , so the policy is parameterized. . Therefore, the goal of reinforcement learning performed by the update unit (250) according to this embodiment is to set parameters that satisfy the following mathematical expression 1. It may be looking for .
[0106] [Mathematical Formula 1]
[0107]
[0108] Here is a policy It can mean the total sum of rewards that can be obtained when progressing through the episode.
[0109] And the update unit (250) according to the present embodiment can perform reinforcement learning according to a problem definition set in advance to generate observation information including extracted external features and configure status information of the drone (10) according to the generated observation information.
[0110] Specifically, the update unit (250) can first define POMDP by discretizing time t stepwise at 0.05 second intervals.
[0111] And at each stage, the update unit (250) collects wing deformation data from the strain sensor (100). Observe and then The state can be constructed based on the last 32 observations, which correspond to 1.6 seconds of observation.
[0112] After that, the update section (250) The action can be determined according to the motor of the driving device (300) which controls the wing. The action continues as follows. Action at time t If it is 1, the motor output is maximum, and if it is 0, the motor output is minimum. At this time, the discount factor can be set to 0.98.
[0113] The replay buffer (290) described later includes observation information, which is wing deformation data generated according to the flight of the drone (10), and transition information including the behavior of the drone (10). This is collected, and the update unit (250) can sample a mini-batch consisting of 64 transitions among the transitions collected from the replay buffer (290) and use it for learning. Through this, new parameters are generated and sent directly to the neural network to replace the previous parameters, and this process can be repeated by the update unit (250) until the algorithm converges.
[0114] In addition, the update unit (250) according to the present embodiment can use the SAC (Soft Actor-Critic) algorithm to perform reinforcement learning, and the update unit (250) can use the gradient descent method to update the state value. , action value and policies The ultimate goal can be to learn the three function parameters.
[0115] First, the update unit (250) is the status value Learning can be performed to minimize the squared residual error as expressed in the following mathematical equation 2.
[0116] [Equation 2]
[0117]
[0118] Here is a parameter is the objective function of the value function according to , is the expected value according to the state sampled from the database at time t, can mean the expected sum of reward and entropy at the state at time t.
[0119] also is the parameter at time t Policy according to Expected value according to the behavior sampled from , is the expected sum of reward and entropy when a specific action is taken at the state at time t, is the parameter at the state at time t It may mean an action determined by a policy according to . And D may mean the state and action distribution of the replay buffer (290).
[0120] And the update part (250) is the action value As shown in the following mathematical expression 3, it is a soft Bellman residue. Learning can be performed to minimize .
[0121] [Equation 3]
[0122]
[0123]
[0124] Here Is It stands for exponential moving average and was introduced to stabilize the learning process.
[0125] also is a parameter The objective function of the action value function Q according to , and the parameter to reduce the soft Bellman residual can be updated.
[0126] and is the expected value when an action is taken in a state sampled from the dataset at time t, is the reward for taking action at point t, can mean a soft action value function when an action is taken at time t.
[0127] also is a variable in the state at time t+1 It can mean a soft state value function according to the moving average of , can mean a rate of decrease with a discount factor.
[0128] And the update section (250) uses the Kullback-Leibler (KL) divergence The purpose of can be defined as in the following mathematical formula 4.
[0129] [Equation 4]
[0130]
[0131] Here is distributed It can mean a partitioning function that normalizes .
[0132] In particular, the update unit (250) according to this embodiment is a policy and action value We can use the SAC algorithm that only uses , and learn the action a by explicitly setting the target entropy.
[0133] Specifically, the update unit (250) uses a one-dimensional convolutional neural network because the width of the time series data is fixed. A neural network whose policy is a set of states constructed based on 32 observations, such as When input, two consecutive 1D CNN layers and a max-pooling layer can be applied in the time dimension.
[0134] Here, the 1D CNN layer can use 32 filters with a width of 5, and the pooling layer can use a window with a width of 2. The output tensor can then be flattened into a 1D vector and inserted into a 256-dimensional vector through a linear layer. Finally, the action and log probability can be returned through separate linear layers.
[0135] Action value has essentially the same structure, but an additional connection layer may be added immediately before the final output layer. The update unit (250) according to the present embodiment may insert an action into a 64-dimensional vector and connect it to the flattened output of a convolutional neural network (CNN) layer.
[0136] and and has different network parameters, and the nonlinear ReLU function can be applied to the output of all convolutional and linear layers. The relevant hyperparameters can be as shown in Table 1 below.
[0137] HyperparameterValueLearning rate of π0.0005Learning rate of q0.001Learning rate of α0.001Initial α0.01 0.98 (for EMA update)0.01Batch size32Target entropy-1.0Maximum beffer size3000Initial training buffer size1000Maximum episode length (number of steps)300Decision period(s)0.05
[0138] Meanwhile, the flight control unit (270) can control the driving device (300) so that the drone (10) can fly according to the autonomous flight behavior determined by the autonomous flight policy updated by the update unit (250).
[0139] Meanwhile, the replay buffer (290) is provided to collect observation information, which is wing deformation data at a certain time after the flight control unit (270) controls the driving device (300), and transitions including the behavior of the drone (10).
[0140] The replay buffer (290) according to this embodiment includes observation information, which is wing deformation data generated according to the flight of the drone (10), and transition information including the behavior of the drone (10). This can be collected.
[0141]
[0142] Meanwhile, FIG. 8 is a flowchart for explaining an autonomous flight method according to one embodiment of the present invention. Since the autonomous flight method according to one embodiment of the present invention is performed on a configuration substantially identical to that of the controller (200) illustrated in FIG. 7, the same components as those of the controller (200) illustrated in FIG. 7 are given the same drawing reference numerals, and repeated explanations are omitted.
[0143] The autonomous flight control method according to the present embodiment includes a step of collecting wing deformation data (S110), a step of extracting external features and internal features (S130), a step of performing reinforcement learning (S150), and a step of controlling a flapping drone (S170).
[0144] In the step of collecting wing deformation data (S110), the collection unit (210) can collect wing deformation data according to the flight of the flapping drone (10) from a strain sensor (100) attached to the wing of the flapping drone (10).
[0145] In the step of extracting external features and internal features (S130), the feature extraction unit (230) can extract external features and internal features of the flapping drone (10) from wing deformation data.
[0146] At this time, the external characteristic may include at least one of the wind direction and wind speed when the flapping drone (10) is in flight, and the internal characteristic may include at least one of the attitude of the flapping drone (10) and the wing flapping speed.
[0147] And in the step (S130) of extracting external features and internal features, the feature extraction unit (230) can extract external features using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to the wind direction and wind speed.
[0148] In the step of performing reinforcement learning (S150), the update unit (250) can perform reinforcement learning to update the autonomous flight policy.
[0149] Specifically, in the step (S150) of performing reinforcement learning, the update unit (250) may perform reinforcement learning to find an autonomous flight policy that allows the flapping drone (10) to reach the target location using external and internal features, and update the autonomous flight policy that provides the highest reward.
[0150] In the step (S150) of performing reinforcement learning, the update unit (250) can perform reinforcement learning according to a problem definition set in advance so as to generate observation information including the extracted external features and configure the status information of the flapping drone (10) according to the generated observation information.
[0151] And in the step (S150) of performing reinforcement learning, the update unit (250) can learn the behavior of the flapping drone (10) according to the following mathematical expression 5.
[0152] [Equation 5]
[0153]
[0154]
[0155] Meanwhile, in the step of controlling the flapping drone (S170), the flight control unit (270) can control the flapping drone (10) according to the autonomous flight behavior determined by the autonomous flight policy.
[0156] In other words, in the step of controlling the flapping drone (S170), the flapping drone (10) can be controlled by outputting a control signal to the driving device (300) provided in the flapping drone (10) so that the drone (10) flies according to the determined autonomous flight behavior.
[0157] And the autonomous flight control method of the present invention may further include a step of collecting observation information, which is wing deformation data at a certain time after the step (S170) in which the replay buffer (290) controls the flapping drone, and a transition including the behavior of the flapping drone (10).
[0158] Therefore, in the step (S150) of performing reinforcement learning, the update unit (250) can update the autonomous flight policy based on the data collected in the replay buffer (290) corresponding to the current state information of the flapping drone (10).
[0159] The autonomous flight method of the present invention may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, and the like, either singly or in combination.
[0160] The program commands recorded on the above computer-readable recording medium may be specially designed and configured for the present invention or may be known and available to those skilled in the art of computer software.
[0161] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.
[0162] Examples of program instructions include not only machine language codes, such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0163] Although various embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
[0164] Meanwhile, FIGS. 9 to 13 are drawings for explaining the experimental process and experimental results for verifying the performance of a flapping drone (10) according to one embodiment of the present invention.
[0165] Figure 9a is a diagram illustrating an experimental design for changing wind direction and speed. To more accurately verify whether the wing deformation data collected by the strain sensor in this embodiment includes information on changes in wind direction and speed, an experiment was conducted to restrict the drone's movements, as shown in Figure 9a.
[0166] First, to determine the relationship between various flow environments and wing deformation, the drone was fixed to a three-degree-of-freedom system with three motors along the x, y, and z axes, as illustrated in Fig. 9a, and the experiment was conducted. Then, wind was generated using a fan and the drone was rotated with various combinations of yaw, pitch, and roll angles.
[0167] Fig. 9b illustrates a strain sensor that changes according to the wind direction and wind speed generated in an experimental environment. Using the system illustrated in Fig. 9a, 62 wind directions were generated as illustrated in Fig. 9b, and three wind speeds of 3 m / s, 5 m / s, and 7 m / s were generated, generating a total of 186 different wind types. Accordingly, as illustrated in Fig. 9b, the strain sensor attached to the wing changes according to the direction and speed of the blowing wind.
[0168] Meanwhile, Figures 9c and 9d are diagrams showing the results of verifying a one-dimensional convolutional neural network model learned to extract external features.
[0169] In order to extract external features from the wing deformation data measured by the strain sensor, the wing deformation data, which is raw data measured by the strain sensor, is input into a one-dimensional convolutional neural network, and the one-dimensional convolutional neural network is trained to predict and output wind direction and wind speed.
[0170] Accordingly, as shown in Fig. 9c, the validity of state recognition was verified using only strain information based on the ROC-AUC (average area under the curve) and confusion matrix of the one-dimensional convolutional neural network model for which learning was completed, and the verification result confirmed that the average AUC was approximately 0.99.
[0171] In addition, as shown in Fig. 9d, it was confirmed that the learned one-dimensional convolutional neural network model classified wind direction and wind speed with an average accuracy of 80%.
[0172] And Fig. 9e shows the accuracy of the trained one-dimensional convolutional neural network model as a color map in spherical coordinates. As shown in Fig. 9e, the prediction accuracy was highest in the front, lower left, and lower right areas, and the accuracy was lowest in the upper rear area. This is because the wind blowing in the front, left, and right areas caused the largest mechanical deformation of the wing base. These results confirm that the drone can accurately determine the wind direction and wind speed in various airflows, especially when the wind blows from the front, left, and right sides, using only the wing deformation data collected from the two strain sensors on the wing base.
[0173] Meanwhile, Fig. 10 shows wing deformation data, i.e. sensing values, measured from a drone in a free-flight state rather than a fixed state as in Fig. 9a described above.
[0174] Specifically, FIG. 10a shows the change in amplitude measured from a high-speed camera image and a strain sensor during free flight, and FIG. 10b shows the change in amplitude measured from a strain sensor according to the trajectory and stroke cycle of a drone wing viewed from the front.
[0175] Figure 10c shows the results of a fast Fourier transform analysis of the sensor response during free flight.
[0176] Figure 10d is a comparison result between the 6-axis load cell data and the calculated lift and thrust, and Figures 10e and 10f are changes in the wing trajectory and sensor response when the wind blows from left to right.
[0177] Figure 10g shows the results of a marathon test where the drone flapped its wings for 30 minutes at a flapping frequency of 14 Hz.
[0178] Figure 11a is a diagram illustrating an experimental environment for performing reinforcement learning in a free flight situation.
[0179] The drone is tethered with a nylon rope to prevent it from colliding with the experimental environment and being damaged, and to return to its initial position after each learning episode.
[0180] The drone is controlled by changing the output of motor 1, which controls thrust, and motor 2, which controls direction (yawing).
[0181] A drone's movements can be divided into four categories: clockwise rotation, counterclockwise rotation, acceleration, and deceleration. The drone can be said to fly with six degrees of freedom, controlled by two motors and two degrees of freedom. The drone's target arrival location is specified by x and y coordinates.
[0182] The drone's position information measured by the motion capture camera is not used for drone control, but rather as a reward function during reinforcement learning. This reward function uses a two-dimensional multivariate probability distribution function, with the target position input as the mean.
[0183] Reinforcement learning uses the SAC algorithm. Ultimately, the reinforcement learning-based controller trains to determine the optimal path with the highest reward by adjusting the output of the motors to reach the target position.
[0184] Figure 11b is a representative image of a flapping drone flying to a target location using only information obtained from strain sensors attached to the wings.
[0185] Figure 11c shows the drone's trajectory in xyz space, and the black line in Figure 11c is the flight trajectory 7.5 seconds ago.
[0186] Figure 11d is a diagram showing changes in wing deformation and motor output during flight toward the target position (7.5 to 9.5).
[0187] Meanwhile, Fig. 12 is a diagram showing the results of reinforcement learning on a drone, and Figs. 12a and 12b are diagrams showing wing deformation and motor output of the drone when heading to a target location on the left side of the drone.
[0188] Figures 12c and 12d are diagrams showing wing deformation and motor output of the drone when heading toward a target location on the right.
[0189] Figures 12e and 12f are diagrams showing three-dimensional trajectories of a trained drone toward target locations on the left and right, respectively.
[0190] Meanwhile, Figure 13 is a diagram comparing the difference between a flapping drone that performed reinforcement learning and a flapping drone that did not perform reinforcement learning.
[0191] Figures 13a and 13b illustrate the trajectories of a non-reinforcement-learned flapping drone toward a target position in the middle, and Figures 13c and 13d illustrate the trajectories of a reinforcement-learned flapping drone toward a target position in the middle.
[0192] And Figures 13e and 13f are diagrams showing changes in entropy and score of an artificial neural network according to the number of training epochs.
[0193] As can be seen from the above drawing, it has been confirmed that the drone control method according to the present embodiment can fly stably using only wing deformation information even when experiencing various types of wind.
[0194] [Explanation of symbols]
[0195] 10: Flapping drone 100: Strain sensor
[0196] 200: Autonomous flight controller 210: Collection unit
[0197] 230: Feature extraction unit 250: Update unit
[0198] 270: Flight Control Unit 290: Replay Buffer
[0199] 300: Drive Unit
Claims
1. A method for controlling autonomous flight of a flapping drone in an autonomous flight controller that controls autonomous flight of a flapping drone, A step of collecting wing deformation data according to the flight of the flapping drone from a strain sensor attached to the wing of the flapping drone; A step of extracting external features and internal features of the flapping drone from the above wing deformation data; A step of performing reinforcement learning to find an autonomous flight policy that allows the flapping drone to reach the target location using the external and internal features, and to update the autonomous flight policy that provides the highest reward; and A method for controlling autonomous flight of a flapping drone, comprising a step of controlling the flapping drone according to an autonomous flight behavior determined by the autonomous flight policy.
2. In paragraph 1, The above external features are, Including at least one of the wind direction and wind speed when the above flapping drone is in flight, The above internal features are: A method for controlling autonomous flight of a flapping drone, comprising at least one of the attitude of the flapping drone and the wing flapping speed.
3. In paragraph 2, In the step of extracting the above features, A method for controlling autonomous flight of a flapping drone, wherein the external features are extracted using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to the wind direction and wind speed.
4. In paragraph 1, In the step of performing the above reinforcement learning, A method for controlling autonomous flight of a flapping drone, wherein reinforcement learning is performed according to a problem definition that is set in advance to generate observation information including the extracted external features and configure state information of the flapping drone according to the generated observation information.
5. In paragraph 4, In the step of performing the above reinforcement learning, A method for controlling autonomous flight of a flapping drone, which learns the behavior of the flapping drone according to the following mathematical formula. [Mathematical formula] 6. In paragraph 4, The autonomous flight control method of the above flapping drone is, A method for controlling autonomous flight of a flapping drone, further comprising a step of collecting, in a replay buffer, observation information, which is wing deformation data, and a transition including an action of the flapping drone at a predetermined time after the step of controlling the flapping drone.
7. In paragraph 6, In the step of performing the above reinforcement learning, A method for controlling autonomous flight of a flapping drone, wherein the autonomous flight policy is updated based on data collected in the replay buffer corresponding to current state information of the flapping drone.
8. For flapping drones capable of autonomous flight, A driving device that provides thrust to the above flapping drone and controls its flight direction; A strain sensor attached to the wing of the above flapping drone; and A flapping drone comprising an autonomous flight controller which collects wing deformation data according to flight from the strain sensor, extracts external features and internal features of the flapping drone from the wing deformation data, finds an autonomous flight policy that allows the flapping drone to reach a target position using the external features and internal features, performs reinforcement learning to find the autonomous flight policy that provides the highest reward, and controls the actuator so that the flapping drone can fly according to an autonomous flight behavior determined by the autonomous flight policy.
9. In an autonomous flight controller that controls the driving device equipped in a flapping drone, A collection unit that collects wing deformation data according to the flight of the flapping drone from a strain sensor attached to the wing of the flapping drone; A feature extraction unit for extracting external features and internal features of the flapping drone from the above wing deformation data; An update unit that performs reinforcement learning to find an autonomous flight policy that allows the flapping drone to reach a target location using the external and internal features, and updates the autonomous flight policy that provides the highest reward; An autonomous flight controller including a flight control unit that controls the actuator so that the flapping drone can fly according to an autonomous flight behavior determined by the autonomous flight policy.
10. In paragraph 9, The above external features are, Including at least one of the wind direction and wind speed when the above flapping drone is in flight, The above internal features are: An autonomous flight controller comprising at least one of the attitude and wing flapping speed of the flapping drone.
11. In paragraph 10, The above feature extraction unit, An autonomous flight controller that extracts the external features using a one-dimensional convolutional neural network model trained to identify the deformation relationship of the wing according to the wind direction and wind speed.
12. In paragraph 9, The above update section, An autonomous flight controller that performs reinforcement learning according to a problem definition that is set in advance to generate observation information including the extracted external features and configure state information of the flapping drone according to the generated observation information.
13. In paragraph 12, The above update section, An autonomous flight controller that learns the behavior of the flapping drone according to the following mathematical formula. [Mathematical formula] 14. In paragraph 12, The above autonomous flight controller, An autonomous flight controller further comprising a replay buffer for collecting transitions including observation information, which is wing deformation data at a predetermined time after controlling the driving device, and actions of the flapping drone.
15. In paragraph 14, The above update section, An autonomous flight controller that updates the autonomous flight policy based on data collected in the replay buffer corresponding to the current state information of the flapping drone.
Citation Information
Patent Citations
Bionic flapping-wing robot flapping-wing mechanism capable of being folded in order and control method
CN114735212A
The apparatus and method of wireless flapping flight with auto control flight and auto navigation flight
KR101204720B1
Apparatus and Method for Diagnosing Safety of Light Aircraft
KR102016124B1
Flapping-wing aerial robot formation control method
US20230083210A1
KR20200114698A
Cited By
Adaptive dynamic programming control method of flexible flapping wing system under time-varying constraint
CN120972538A