Drone and control method therefor
The drone system uses a combination of policy-based reinforcement learning and Bayesian networks to estimate and offset wind-induced disturbances, addressing the issue of drone instability in windy environments and ensuring stable control.
Patent Information
- Application Number
- PCT/KR2024/019561
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-26
AI Technical Summary
Drones used in windy environments, such as lithium salt lakes, are prone to drifting or tilting due to gusts of wind, leading to potential crashes or loss of control, as traditional PID controllers are vulnerable to external disturbances.
A drone system that incorporates a policy-based reinforcement learning network and a Bayesian network to estimate disturbances, allowing the drone to adjust its motor control signals and maintain attitude stability by offsetting actual disturbances.
The system enables stable drone control with low delay in windy conditions, preventing crashes and maintaining attitude stability by accurately estimating and offsetting wind-induced disturbances.
Smart Images

Figure KR2024019561_26062025_PF_FP_ABST
Abstract
Description
Drone and method of controlling the drone
[0001] The present embodiments relate to a drone for controlling an aircraft in a specific situation and a method for controlling the drone.
[0002] For example, to produce brine lithium in Argentina, target concentrations must be maintained in each pond. To maintain these concentrations, management tasks include sampling and measuring the depth of each brine pond. Drones can be used to improve the efficiency of this process, but lithium brine ponds are located in windy areas to maximize evaporation.
[0003] Drones are typical under-actuated systems and are vulnerable to external disturbances such as wind. In addition, the most commonly used PID controllers are vulnerable to gusts because their hyperparameters are tuned for path following.
[0004] Most drones carry a risk of drifting or tilting under certain conditions, such as gusts of wind, which can lead to a crash or loss of control.
[0005] The present embodiments can provide a drone and a method for controlling the drone that can stably control the drone with low latency in specific situations such as gusts of wind.
[0006] Embodiments of the present invention provide a drone and a control method thereof, which control the drone to fly along a specific flight trajectory by converting a motor signal into a motor control signal in a general situation, and measure an attitude based on sensor data to recognize a specific situation, and estimate the difference between an input value output from a policy-based reinforcement learning network that is trained to receive a desired current state as an input and output an input value and an estimated input value output from a Bayesian network that is trained to receive a current state and a next state as an input and output an estimated input value, and control the drone to fly while maintaining an attitude by adjusting a motor control signal in a specific situation with a control signal that offsets an actual disturbance by subtracting the estimated disturbance from the input value.
[0007] In one aspect, the present embodiments provide a drone including a drone sensor unit that provides sensor data, a drone flight unit that provides a motor signal, and a drone control unit that controls the drone to fly along a specific flight trajectory by converting the motor signal into a motor control signal in a general situation, measures an attitude based on the sensor data to recognize a specific situation, and estimates the difference between an input value output from a policy-based reinforcement learning network that receives a desired current state as an input and outputs an input value and an estimated input value output from a Bayesian network that receives a current state and a next state as an input and outputs an estimated input value, and controls the drone to fly while maintaining an attitude by adjusting the motor control signal in a specific situation with a control signal that offsets a disturbance that may actually occur by subtracting the estimated disturbance from the input value.
[0008] In another aspect, the present embodiments can provide a method for controlling a drone, including a first step of controlling the drone to fly along a specific flight trajectory by converting a motor signal into a motor control signal in a general situation, and a second step of controlling the drone to fly while maintaining the attitude by adjusting the motor control signal in a specific situation by measuring the attitude based on sensor data, recognizing a specific situation, estimating the difference between an input value output from a policy-based reinforcement learning network trained to receive a desired current state as an input and output an input value, and an estimated input value output from a Bayesian network trained to receive a next state as an input and output an estimated input value, and subtracting the estimated disturbance from the input value to offset a disturbance that may actually occur.
[0009] According to the drone and the control method thereof according to the present embodiments, the drone can be stably controlled with low delay in specific situations such as gusts of wind.
[0010] Figure 1 is a perspective view of a drone to which embodiments are applied.
[0011] Figure 2 is a configuration diagram of a drone according to one embodiment.
[0012] Figure 3 is a configuration diagram of an example of the drone control unit of Figure 2.
[0013] Fig. 4 is a relationship diagram between the first control unit and the drone drive unit of Figs. 2 and 3.
[0014] Fig. 5 is a relationship diagram between the second control unit and the drone drive unit of Figs. 2 and 3.
[0015] Figure 6 is a conceptual diagram of a policy-based reinforcement learning network used in the second control unit (154) of Figure 5.
[0016] Figure 7 illustrates an example of the relationship between the policy-based reinforcement learning network of the second control unit of Figure 5 and the Bayesian network.
[0017] Figure 8 illustrates another example of the relationship between the policy-based reinforcement learning network of the second control unit of Figure 5 and the Bayesian network.
[0018] Fig. 9 is an example of application of the second control unit and drone drive unit of Figs. 2 and 3 in the learning stage.
[0019] Fig. 10 is an example of application of the second control unit and drone drive unit of Figs. 2 and 3 in the use stage.
[0020] Fig. 11 is a conceptual diagram of the drone driving unit of Figs. 2 and 3.
[0021] Figure 12 illustrates an example of a simulator of the drone of Figures 2 and 3.
[0022] Figure 13 is a flowchart of a drone control method according to another embodiment.
[0023] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, when describing the present disclosure, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present disclosure.
[0024] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.
[0025] In embodiments of the present invention, the form of energy may be electrical energy, thermal energy, light energy, etc. Hereinafter, embodiments of the present invention will be mainly described with the case where the form of energy is electrical energy, but the form of energy is not limited thereto.
[0026] Hereinafter, an energy management device, a multi-agent based battery cooling plate design device, and a multi-agent based battery cooling plate design device according to embodiments of the present invention will be described with reference to related drawings.
[0027] The embodiments are described in detail with reference to the drawings below.
[0028] Figure 1 is a front view of a drone according to one embodiment.
[0029] Referring to FIG. 1, a drone (100) according to one embodiment can fly using a propeller (110) and an electric motor (120). The electric motor (120) is used to convert electric power into electrical energy to rotate the propeller (110). This rotational motion pushes air, creating a force that propels the drone (100) upward. The propeller (110) uses this rotational motion to help move the air, and the drone (100) can fly using this principle.
[0030] Drones (100) are equipped with various sensors, cameras, GPS, communication systems, etc. and are used for flight control and data collection.
[0031] The drone (100) has multiple (e.g., four in FIG. 1) propellers (110), and each propeller (110) can be controlled to perform ascending, descending, forward, backward, left and right movement, rotation, etc. To this end, the direction and height of the drone (100) are controlled by controlling the rotation speed of each propeller (110).
[0032] The drone (100) has a built-in flight control system. This system controls the speed of each electric motor (120) and adjusts the attitude to move in the desired direction according to commands entered by the user.
[0033] According to one embodiment, a drone (100) stably controls its attitude and speed without drifting or tilting, falling, or losing control in a specific situation such as a gust of wind.
[0034] Figure 2 is a configuration diagram of the drone of Figure 1.
[0035] Referring to FIG. 2, a drone (100) according to one embodiment includes a drone sensor unit (130) that provides sensor data, a drone flight unit (140) that provides a motor signal, and a drone control unit that converts the motor signal into a motor control signal and controls the drone to fly in a specific flight trajectory.
[0036] The drone control unit (150) controls the drone to fly along a specific flight trajectory by converting a motor signal into a motor control signal in a general situation, measures the attitude based on sensor data to recognize a specific situation, and estimates the difference between the input value output from a policy-based reinforcement learning network that is trained to receive a desired current state as input and output an input value and the estimated input value output from a Bayesian network that is trained to receive the current state and the next state as input and output an estimated input value as an estimated disturbance, and controls the drone to fly while maintaining the attitude by adjusting the motor control signal in a specific situation with a control signal that offsets the disturbance that may actually occur by subtracting the estimated disturbance from the input value.
[0037] Hereinafter, specific situations are described as examples of gusts of wind as described above, but are not limited thereto, and generally include all cases other than general situations in which the drone (100) may drift or tilt, causing it to crash or lose control. That is, this specification defines two situations in which the drone (100) may perform general control operations in a first situation such as a general situation, and perform special control operations in a second situation such as a specific situation. As described below, the drone (100) may perform software control operations or hardware control operations in both general and special situations.
[0038] The drone sensor unit (130) includes a gyro sensor (132) and an acceleration sensor (134), and the sensor data may be, but is not limited to, angular velocity sensed by the gyro sensor (132) and acceleration sensed by the acceleration sensor (134).
[0039] The gyro sensor (132) can measure angular velocity. An example of the gyro sensor (132) may be a gyroscope that measures angular velocity. The acceleration sensor (134) is a sensor that measures acceleration. The acceleration sensor (132) can measure acceleration in the x-axis, y-axis, and z-axis directions, for example.
[0040] The drone flight unit (140) can convert the motor control signal adjusted by the drone control unit (150) into a specific PWM signal that adjusts the motor speed and direction.
[0041] The drone flight unit (140) converts the motor control signal adjusted by the drone control unit (150) into a corresponding PWM signal that adjusts the motor speed and direction. The control logic of the drone flight unit (140) uses aerodynamic and flight dynamics models to compensate for wind disturbances and maintain the desired flight trajectory.
[0042] The drone flight unit (140) provides the converted PWM signal to the electric motor (120) and controls the electric motor (120), thereby inducing the flight of the drone (200).
[0043] The drone control unit (150) includes a microprocessor, and the drone (100) may additionally include memory (not shown). The memory stores various sensor data and programs. The memory may be volatile memory (e.g., SRAM, DRAM) or non-volatile memory (e.g., NAND Flash).
[0044] The drone (100) according to the above-described embodiment learns the model dynamics of the drone through a Bayesian network by utilizing a simulator of the commercial drone to be used.
[0045] A Bayesian network is a probabilistic graphical model used to represent probabilistic dependencies between variables. The parameters of a Bayesian network are the conditional probabilities of each variable constituting the network, and these parameters are used to model the relationships among variables. For example, the parameter estimation method used in the Bayesian network learning process utilizes a method called Maximum Likelihood Estimation (MLE). This method identifies the most likely parameters from data.
[0046] A trained Bayesian network uses learned parameters to calculate probabilities when making predictions on new data.
[0047] In a broad sense, Bayesian networks exist in various types, including Bayesian networks, which represent the conditional dependencies between general variables as a graph, Markov networks, which are represented as graph structures consisting of nodes and edges, where edges represent conditional independence, or Markov random fields, dynamic Bayesian networks, which are used to model the dependencies between variables in situations that change over time, and structural Bayesian networks, which automatically learn structures from data to create networks.
[0048] Additionally, the Bayesian network used may be a lightweight sparse Bayesian network that calculates the signal-to-noise ratio for the parameter distribution at each step during learning, and if the signal-to-noise ratio value is small, it is considered a parameter with little influence on the result value, so it may be pruned.
[0049] For example, the drone (100) according to the aforementioned embodiment considers the difference between the system input estimate output from the trained sparse Bayesian network and the intended system input as a disturbance and compensates for it in the next input. The sparse Bayesian network compresses the network during the learning process, thereby reducing capacity and inference speed. The drone (100) according to the aforementioned embodiment can prevent crashes due to gusts when using the drone in the relatively windy airspace over Argentina.
[0050] In certain situations, such as gusts of wind, drones can be controlled to hover (suspend in mid-air) when they crash, using reinforcement learning. However, due to the nature of reinforcement learning, it is vulnerable to disturbances not encountered in the training environment. Therefore, for real-world applications, methods to improve the robustness of policy-based reinforcement learning networks are needed. To improve the robustness of reinforcement learning, model dynamics can be estimated using a conventional artificial neural network, followed by the application of a disturbance observer. However, the inherent large capacity and relatively slow inference speed of conventional artificial neural networks limit their use in performance-constrained environments such as embedded systems.
[0051] The drone (100) according to the above-described embodiment can increase the robustness of a policy-based reinforcement learning network and reduce network capacity and inference speed by utilizing a Bayesian network, for example, a sparse Bayesian network and a disturbance observer.
[0052] Figure 3 is a configuration diagram of an example of the drone control unit of Figure 2.
[0053] Referring to FIG. 3, the drone control unit (150) may include a first control unit (152) that controls the drone to fly along a specific flight trajectory by converting a motor signal into a motor control signal in a general situation, and a second control unit (154) that controls the drone to fly while maintaining an attitude by adjusting the motor control signal in a specific situation with a control signal that offsets a disturbance that may actually occur by subtracting an estimated disturbance, which is the difference between an input value output from a policy-based reinforcement learning network and an estimated input value output from a Bayesian network, from the input value.
[0054] The first control unit (152) and the second control unit (154) may be implemented as separate hardware, may be implemented as software, or may be implemented as hardware and the other as software.
[0055] Fig. 4 is a relationship diagram between the first control unit and the drone drive unit of Figs. 2 and 3. Fig. 5 is a relationship diagram between the second control unit and the drone drive unit of Figs. 2 and 3.
[0056] Referring to FIG. 4, the first control unit (152) converts a motor signal into a motor control signal in a general situation and controls the drone to fly along a specific flight trajectory. The first control unit (152) receives a desired current state (Sr) in a general situation and outputs a motor control signal (u) to the drone drive unit (140). The drone drive unit (140) drives the electric motor with the received motor control signal (u).
[0057] Referring to Fig. 5, the second control unit (154) is configured to control the desired current state (s r ) and the output value (s) of the drone driving unit (140), i.e., the error value (e s ) input and output a control signal value (u) from a policy-based reinforcement learning network (policy network, 156) that has been trained to output a control signal value (u) and the current state (s). r ) and the next state (s) are input to estimate the input value ( ) output from the Bayesian network (158) trained to output the estimated input value ( ) difference( -u) estimated disturbance ( ) and the estimated disturbance ( ) to offset the actual control signal (u) that may actually occur by subtracting the disturbance (d). real ) can be controlled to maintain attitude and fly by adjusting the motor control signals in certain situations.
[0058] For a simple example, if the control signal of the drone drive unit (140) is an RPM value, let us assume that the control signal (u) output from the policy-based reinforcement learning network (policy network, 156) is 4000 RPM. At this time, the estimated disturbance ( ) is 200 RPM, the actual control signal value input to the drone drive unit (140) can be 3800 RPM.
[0059] Unlike the general method of calculating the inverse function of a dynamic model to estimate input values, the Bayesian network (158) is trained to estimate input values through supervised learning using training data pairs in the reinforcement learning buffer, as described below. The second control unit (154) can implement a disturbance observer in the policy-based reinforcement learning network (154) without physical modeling by replacing the artificial neural network inverse model dynamics described above with the trained Bayesian network (158).
[0060] Hereinafter, the general reinforcement learning and the policy-based reinforcement learning network and Bayesian network (158) used in the second control unit (154) of the drone (100) according to the aforementioned embodiment will be described in detail with reference to FIGS. 6 to 10.
[0061] Figure 6 is a conceptual diagram of a policy-based reinforcement learning network used in the second control unit (154) of Figure 5.
[0062] Referring to FIG. 6, the policy-based reinforcement learning network (158) used in the second control unit (154) is a system in which an environment (156a) and an agent (156b) exchange states / rewards and actions. The environment (156a) may include various types of simulation environments or simulators, as described below. In this specification, the terms environment (156a) and simulator are used interchangeably, but are not limited thereto.
[0063] The environment (simulator) (156a) is in the current state (S t ) and action (A t ) is input and the next state (Next State, S) is t+1 ) and reward (R) t+1) can be output. In this process, the agent (156b) learns a policy (policy, π(a / s): S==>R∈{0,1}).
[0064] Figure 7 illustrates an example of the relationship between the policy-based reinforcement learning network of the second control unit of Figure 5 and the Bayesian network.
[0065] Referring to Fig. 7, in the policy-based reinforcement learning network (156) used in the second control unit (154), the agent (156b) repeatedly performs the next action based on the state and reward generated in the environment (156a), and updates the action decision so that the sum of the accumulated rewards is maximized.
[0066] The Bayesian network (158) used in the second control unit (154) is a learning data pair (s) generated from policy-based reinforcement learning of the policy-based reinforcement learning network (156). t , , s t+1 ) can be stored in a buffer (159) and then learned by sampling as much as the mini batch size (M).
[0067] Mathematical expression 1 is an example of the objective formula of a Bayesian network (158).
[0068]
[0069] In mathematical equation 1, , θ and M represent the loss function of the Bayesian network, the training parameters of the Bayesian network, and the mini-batch size, respectively. r sbl is the scaling value, D KL represents the KL divergence term as described later. Data for learning is sampled from the buffer (159) of reinforcement learning in amounts equal to the mini-batch size (M).
[0070] For example, a mini-batch with a mini-batch size (M) can be expressed as
[0071] The first item in Equation 1 is the behavioral data used in the simulation (156a). ) and the estimated action output from the policy-based reinforcement learning network (156). ) is designed to minimize the mean square error (MSE) between the two networks. The second term of Equation 1, the KL divergence term, encourages the posterior distribution to remain close to the prior distribution, and the signal-to-noise ratio (Q(W)) is calculated at each step to remove less influential network parameters.
[0072] Figure 8 illustrates another example of the relationship between the policy-based reinforcement learning network of the second control unit of Figure 5 and the Bayesian network.
[0073] Referring to FIG. 8, the Bayesian network (158) used in the second control unit (154) calculates the signal-to-noise ratio for the parameter distribution at each step during learning, and if the corresponding signal-to-noise ratio value (SNR(Q(W))) is small, it is considered a parameter that has little influence on the result value, so it can be pruned to make it lightweight. A network that has been pruned in the Bayesian network (158) to make it lightweight may be the sparse Bayesian network (158b) illustrated in FIG. 8. Hereinafter, the Bayesian network (158) may be a Bayesian network (158) that is not made lightweight, or may be a sparse Bayesian network (158b). Hereinafter, the sparse Bayesian network (158b) will be described as an example of the Bayesian network (158).
[0074] The learned policy-based reinforcement learning network (156) and the sparse Bayesian network (158) are designed as a structure of a disturbance observer consisting of a second control unit (154) that observes disturbances and a drone driving unit (140).
[0075] Fig. 9 is an example of application of the second control unit and drone drive unit of Figs. 2 and 3 in the learning stage.
[0076] Referring to Figures 8 and 9, in the policy-based reinforcement learning network (156) used in the second control unit (154), the agent (156b) implemented as an artificial neural network obtains the desired current state (s t ) and input the input value ( ) outputs the simulator environment or simulator (156a) to this input value ( ) is input and the next state (S t+1 ) is generated. The sparse Bayesian network (158) receives the current state and the next state as input and outputs an estimated input value, and the difference between the input value and the estimated input value is calculated in the comparator, and the sparse Bayesian network (158) is trained so that the accumulated difference value is minimized according to mathematical formula 1 in which this process is performed repeatedly.
[0077] As described above, the Bayesian network (158) used in the second control unit (154) is a learning data pair (s) generated from policy-based reinforcement learning of the policy-based reinforcement learning network (156). t , , s t+1 ) can be learned by sampling as much as the mini batch size (M).
[0078] Fig. 10 is an example of application of the second control unit and drone drive unit of Figs. 2 and 3 in the use stage.
[0079] Referring to Fig. 10, the learned sparse Bayesian network (158) measures the current state (s) measured from the drone drive unit (140) of the actual drone (100). t ) and the next state (s t+1 ) and estimate the system input paired with it.
[0080] The output of the sparse Bayesian network (158) ( ) and the output obtained from the policy-based reinforcement learning network (156) ( ) is considered as an effect of unknown disturbance (d) and model uncertainty. Consequently, the unintended disturbance (d) can be offset by incorporating its estimate into the next control input.
[0081] In other words, the second control unit (154) is configured to control the desired current state (s t ) and input the input value ( ) from a policy-based reinforcement learning network that has been trained to output the input value ( ) and current status (s t ) and the next state (S t+1 ) is input and the estimated input value ( ) from the Bayesian network trained to output the estimated input values ( ) difference( - ) is estimated as a disturbance ( ) and the input value ( ) estimated disturbance( ) can be controlled to maintain attitude and fly by adjusting the motor control signal in a specific situation with a control signal that offsets the disturbance (d) that may actually occur.
[0082] Fig. 11 is a conceptual diagram of the drone driving unit of Figs. 2 and 3.
[0083] Referring to Fig. 11, the policy-based reinforcement learning network (156) of the second control unit, when applied to a reinforcement learning problem, state S is a three-dimensional position (Px, Py, Pz) and a three-dimensional rotation (roll, pitch, yaw), action A is the thrust of each motor (F1, F2, F3, F4), and the sparse Bayesian network (158) is the current state (s t ) and the next state (s t+1 ) as input and the current input (which is the thrust of each motor) ) can be learned to output.
[0084] Figure 12 illustrates an example of a simulator of the drone of Figures 2 and 3.
[0085] Referring to Fig. 12, a simulator for a commercial drone is provided, and reinforcement learning is performed to output the input of the drone (100) based on the state data generated from the simulator during learning. The data is stored in a buffer (159) and is later used to train the model using a sparse Bayesian network (158).
[0086] According to the drone (100) according to the other embodiment described above, it is possible to control stably with low delay in specific situations such as gusts of wind.
[0087] Figure 13 is a flowchart of a drone control method according to another embodiment.
[0088] Referring to FIG. 13, a drone control method (200) according to another embodiment includes a first step (S210) of controlling the drone to fly along a specific flight trajectory by converting a motor signal into a motor control signal in a general situation, and a second step (S210) of controlling the drone to fly while maintaining the attitude by adjusting the motor control signal in a specific situation with a control signal obtained by subtracting the estimated disturbance from the input value by estimating the difference between the input value output from a policy-based reinforcement learning network that is trained to input a desired current state and output an input value and the estimated input value output from a Bayesian network that is trained to input a next state and output an estimated input value by estimating the difference as an estimated disturbance.
[0089] In the second step (S220), the adjusted motor control signal can be converted into a specific PWM signal that adjusts the motor speed and direction.
[0090] As described above with reference to FIGS. 4 and 5, the second step (S220) can control flight while maintaining attitude by adjusting the motor control signal in a specific situation by subtracting the estimated disturbance, which is the difference between the input value output from the policy-based reinforcement learning network and the estimated input value output from the Bayesian network, from the input value and then offsetting the disturbance that may actually occur.
[0091] Step 2 (S220) is to obtain the desired t state (s t ) and input the input value ( ) from a policy-based reinforcement learning network that has been trained to output the input value ( ) and the t state (s t ) and the t-state (S t ) is the next state, the t+1 state (S t+1 ) is input and the estimated input value ( ) from the Bayesian network trained to output the estimated input values ( ) difference( - ) is estimated as a disturbance ( ) and the input value ( ) estimated disturbance( ) can be controlled to maintain attitude and fly by adjusting the motor control signal in a specific situation with a control signal that offsets the disturbance (d) that may actually occur.
[0092] As described above with reference to FIGS. 6 and 7, the Bayesian network used in the second step (S220) is a learning data pair (s) generated from policy-based reinforcement learning of a reinforcement learning network. t , , s t+1 ) can be stored in a buffer and then trained by sampling in the size of the mini-batch.
[0093] As described above with reference to FIGS. 8 and 9, the Bayesian network used in the second step (S220) calculates the signal-to-noise ratio for the parameter distribution at each step during learning, and if the signal-to-noise ratio value is small, it is considered a parameter with little influence on the result value, so it can be pruned to make it lighter.
[0094] As described above with reference to FIG. 10, the second step (S220) is to obtain a desired t-state (s t ) and input the input value ( ) from a policy-based reinforcement learning network that has been trained to output the input value ( ) and the t state (s t ) and the t-state (S t ) is the next state, the t+1 state (S t+1 ) is input and the estimated input value ( ) from the Bayesian network trained to output the estimated input values ( ) difference( - ) is estimated as a disturbance ( ) and the input value ( ) estimated disturbance( ) can be controlled to maintain attitude and fly by adjusting the motor control signal in a specific situation with a control signal that offsets the disturbance (d) that may actually occur.
[0095] As described above with reference to FIG. 11, the policy-based reinforcement learning network of the second stage (S220), when applied to a reinforcement learning problem, the state S is a three-dimensional position (Px, Py, Pz) and a three-dimensional rotation (roll, pitch, yaw), the action A is the thrust of each motor (F1, F2, F3, F4), and the sparse Bayesian network is the current state (s t ) and the next state (s t+1 ) as input and the current input (which is the thrust of each motor) ) can be learned to output.
[0096] According to the drone control method (200) according to the other embodiment described above, the drone can be stably controlled with low delay in specific situations such as gusts of wind.
[0097] Although the drone and the control method thereof according to the embodiments have been described with reference to the above drawings, the present invention is not limited thereto.
[0098] The drone (100) described above may be implemented by a computing device including at least some of a processor, a memory, a user input device, and a presentation device. The memory is a medium that stores computer-readable software, applications, program modules, routines, instructions, and / or data, etc., which are coded to perform a specific task when executed by the processor. The processor can read and execute the computer-readable software, applications, program modules, routines, instructions, and / or data stored in the memory. The user input device may be a means for allowing a user to input a command to cause the processor to perform a specific task or to input data necessary for the execution of a specific task. The user input device may include a physical or virtual keyboard or keypad, key buttons, a mouse, a joystick, a trackball, a touch-sensitive input device, or a microphone. The presentation device may include a display, a printer, a speaker, or a vibration device.
[0099] Computing devices can include a variety of devices, including smartphones, tablets, laptops, desktops, servers, and clients. A computing device may be a single, standalone device, or it may include multiple computing devices operating in a distributed environment, each of which collaborates with another through a communications network.
[0100] In addition, the drone (100) described above can be executed by a computing device having a processor and a memory storing computer-readable software, applications, program modules, routines, instructions, and / or data structures coded to perform the drone control method (200) described above when executed by the processor.
[0101] The aforementioned drone control method (200) can be implemented through various means. For example, the aforementioned drone control method (200) can be implemented through hardware, firmware, software, or a combination thereof.
[0102] In the case of hardware implementation, the above-described drone control method (200) can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, or microprocessors.
[0103] For example, the above-described drone control method (200) according to the embodiments may be implemented using an artificial intelligence semiconductor device in which neurons and synapses of a deep neural network are implemented using semiconductor elements. In this case, the semiconductor elements may be currently used semiconductor elements, such as SRAM, DRAM, NAND, etc., or may be next-generation semiconductor elements, such as RRAM, STT MRAM, PRAM, etc., or may be a combination thereof.
[0104] When implementing the above-described drone control method (200) using an artificial intelligence semiconductor device, the results (weights) of learning a deep learning model using software can be transferred to synapse-mimicking elements arranged in an array, or learning can be performed in the artificial intelligence semiconductor device.
[0105] In the case of implementation via firmware or software, the aforementioned drone control method (200) may be implemented in the form of a device, procedure, or function that performs the functions or operations described above. The software code may be stored in a memory unit and driven by a processor. The memory unit may be located within or outside the processor and may exchange data with the processor via various known means.
[0106] Additionally, terms such as "system," "processor," "controller," "component," "module," "interface," "model," or "unit" as described above may generally refer to a computer-related entity, such as hardware, a combination of hardware and software, software, or software in execution. For example, the aforementioned components may be, but are not limited to, a process driven by a processor, a processor, a controller, a control processor, an object, a thread of execution, a program, and / or a computer. For example, both an application running on a controller or a processor and the controller or the processor may be components. One or more components may be within a process and / or thread of execution, and the components may be located on a single device (e.g., a system, a computing device, etc.) or distributed across two or more devices.
[0107] Meanwhile, a computer program stored on a computer storage medium is provided for performing the aforementioned drone control method (200). In addition, another embodiment provides a computer-readable storage medium storing a program for realizing the aforementioned multi-agent-based battery cooling plate design device.
[0108] The program recorded on the recording medium can be read, installed and executed by a computer, thereby executing the steps described above.
[0109] In this way, in order for a computer to read a program recorded on a recording medium and execute functions implemented as a program, the above-mentioned program may include code coded in a computer language such as C, C++, JAVA, or machine language that can be read by the computer's processor (CPU) through the computer's device interface.
[0110] Such code may include functional code related to functions defining the aforementioned functions, and may also include control code related to execution procedures required for the computer's processor to execute the aforementioned functions according to a predetermined procedure.
[0111] Additionally, such code may further include memory reference related code regarding where in the internal or external memory of the computer the additional information or media required for the computer's processor to execute the aforementioned functions should be referenced.
[0112] Additionally, if the computer's processor needs to communicate with another computer or server located remotely in order to execute the functions described above, the code may further include communication-related code regarding how the computer's processor should communicate with another computer or server located remotely using the computer's communication module, and what information or media should be sent and received during the communication.
[0113] The computer-readable recording medium that records the program as described above includes, for example, ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical media storage device, etc., and may also include one implemented in the form of a carrier wave (e.g., transmission via the Internet).
[0114] Additionally, computer-readable recording media can be distributed across network-connected computer systems, allowing computer-readable code to be stored and executed in a distributed manner.
[0115] In addition, the functional program for implementing the present invention and the code and code segments related thereto may be easily inferred or changed by programmers in the technical field to which the present invention belongs, taking into consideration the system environment of the computer that reads the recording medium and executes the program.
[0116] The drone control method (200) described above may also be implemented in the form of a recording medium containing computer-executable commands, such as an application or program module executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. In addition, computer-readable media may include all computer storage media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented with any method or technology for storing information, such as computer-readable commands, data structures, program modules, or other data.
[0117] The drone control method (200) described above can be executed by an application that is installed by default on the terminal (which may include a program included in a platform or operating system installed by default on the terminal), or by an application (i.e., a program) that the user directly installs on the master terminal through an application providing server such as an application store server, an application, or a web server related to the service. In this sense, the multi-agent-based battery cooling plate design device described above can be implemented as an application (i.e., a program) that is installed by default on the terminal or directly installed by the user, and can be recorded on a computer-readable recording medium such as the terminal.
[0118] Although all components constituting the embodiments of the present disclosure have been described above as being combined or operating in combination, the present disclosure is not necessarily limited to such embodiments. That is, within the scope of the purpose of the present disclosure, all components may be selectively combined and operated one or more times. In addition, although all components may be implemented as individual independent hardware, some or all of the components may be selectively combined and implemented as a computer program having program modules that perform some or all of the functions combined in one or more hardware pieces. The codes and code segments constituting the computer program may be easily inferred by those skilled in the art of the present disclosure. Such a computer program may be stored in a computer-readable storage medium and read and executed by a computer, thereby implementing the embodiments of the present disclosure. Storage media for the computer program may include magnetic recording media, optical recording media, etc.
[0119] In addition, terms such as "include," "comprise," or "have" described above, unless specifically stated otherwise, mean that the corresponding component may be included, and therefore should be interpreted to include other components rather than excluding other components. All terms, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which this disclosure pertains, unless otherwise defined. Commonly used terms, such as terms defined in dictionaries, should be interpreted to be consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an ideal or overly formal sense, unless explicitly defined in this disclosure.
[0120] The above description is merely an example of the technical idea of the present disclosure, and those skilled in the art to which the present disclosure pertains will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present disclosure. Therefore, the embodiments disclosed in the present disclosure are not intended to limit the technical idea of the present disclosure, but rather to explain it, and the scope of the technical idea of the present disclosure is not limited by these embodiments. The scope of protection of the present disclosure should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present disclosure.
[0121]
[0122] CROSS-REFERENCE TO RELATED APPLICATION
[0123] This patent application claims priority under 35 USC § 119(a) to Korean Patent Application No. 10-2023-0184113, filed December 18, 2023, the entire contents of which are incorporated herein by reference. Furthermore, this patent application claims priority in countries other than the United States for the same reasons, the entire contents of which are incorporated herein by reference.
Claims
1. Drone sensor section providing sensor data; A drone flight section providing motor signals; and A drone comprising a drone control unit that controls the drone to fly along a specific flight trajectory by converting the motor signal into the motor control signal in a general situation, and measures the attitude based on the sensor data to recognize a specific situation, and estimates the difference between the input value output from a policy-based reinforcement learning network trained to receive a desired current state as an input and output an input value and the estimated input value output from a Bayesian network trained to receive the current state and the next state as inputs, and controls the drone to fly while maintaining the attitude by adjusting the motor control signal in the specific situation as a control signal that offsets the disturbance that may actually occur by subtracting the estimated disturbance from the input value.
2. In paragraph 1, The above drone flight section is a drone that converts the adjusted motor control signal into a specific PWM signal that adjusts the motor speed and direction.
3. In paragraph 1, The above drone control unit In the above general situation, a first control unit that converts the motor signal into the motor control signal and controls the flight to a specific flight trajectory; and A drone including a second control unit that controls the drone to fly while maintaining the attitude by adjusting the motor control signal in the specific situation with the control signal that offsets the disturbance that may actually occur by subtracting the estimated disturbance, which is the difference between the input value output from the policy-based reinforcement learning network and the estimated input value output from the Bayesian network, from the input value.
4. In paragraph 3, The Bayesian network used in the second control unit above is a learning data pair (s) generated from policy-based reinforcement learning of the reinforcement learning network. t , , s t+1 ) is stored in the buffer, and the trained drone is sampled by the mini batch size.
5. In paragraph 4, The Bayesian network used in the second control unit calculates the signal-to-noise ratio for the parameter distribution at each step during learning, and if the signal-to-noise ratio value is small, it is considered a parameter with little influence on the result value, so pruning is performed to make the drone lightweight.
6. In paragraph 3, The above second control unit is configured to control the desired current state (s t ) and input the input value ( ) is output from the above policy-based reinforcement learning network (policy network) that has been reinforced to output the input value ( ) and the current state (s) above t ) and the following state (s) t+1 ) is input and the estimated input value ( ) from the Bayesian network trained to output the estimated input value ( ) difference( - ) to estimate the disturbance ( ) is estimated as the input value ( ) in the above estimated disturbance ( ) is controlled to fly while maintaining the attitude by adjusting the motor control signal in the specific situation with a control signal that offsets the disturbance (d) that may actually occur.
7. In paragraph 3, The policy-based reinforcement learning network of the second control unit, when applied to a reinforcement learning problem, the state S is a three-dimensional position (Px, Py, Pz) and a three-dimensional rotation (roll, pitch, yaw), and the action A is the thrust of each motor (F1, F2, F3, F4). The above sparse Bayesian network is the current state (s t ) and the following state (s) t+1 ) as input and the current input (which is the thrust of each motor) ) trained to output 8. A first step for controlling the flight along a specific flight trajectory by converting the motor signal into the motor control signal in a general situation; and A method for controlling a drone, comprising a second step of controlling the drone to fly while maintaining the attitude by adjusting the motor control signal in the specific situation by estimating the difference between the input value output from a policy-based reinforcement learning network trained to receive a desired current state as input and output an input value and the estimated input value output from a Bayesian network trained to receive a next state as input and output an estimated input value as an estimated disturbance, and controlling the drone to fly while maintaining the attitude by subtracting the estimated disturbance from the input value and using a control signal that offsets disturbance that may actually occur.
9. In paragraph 8, A method for controlling a drone by converting, in the second step, the adjusted motor control signal into a specific PWM signal for adjusting the motor speed and direction.
10. In paragraph 8, The second step is a method for controlling a drone to fly while maintaining the attitude by adjusting the motor control signal in the specific situation with the control signal that offsets the disturbance that may actually occur by subtracting the estimated disturbance, which is the difference between the input value output from the policy-based reinforcement learning network and the estimated input value output from the Bayesian network, from the input value.
11. In paragraph 10, The Bayesian network used in the second step above is a learning data pair (s) generated from policy-based reinforcement learning of the reinforcement learning network. t , , s t+1 ) is stored in a buffer, and then a method for controlling a learned drone is sampled by the size of the mini batch.
12. In paragraph 11, The Bayesian network used in the second step calculates the signal-to-noise ratio for the parameter distribution at each step during learning, and if the signal-to-noise ratio value is small, it is considered a parameter with little influence on the result value, so pruning is performed to control a lightweight drone.
13. In paragraph 10, The second step above is to obtain the desired t state (s t ) and input the input value ( ) is output from the above policy-based reinforcement learning network (policy network) that has been reinforced to output the input value ( ) and the above t state (s t ) and the above t state (s t ) is the next state, the t+1 state (s). t+1 ) is input and the estimated input value ( ) from the Bayesian network trained to output the estimated input value ( ) difference( - ) to estimate the disturbance ( ) is estimated as the input value ( ) in the above estimated disturbance ( ) is subtracted from the drone to control the drone so that it can fly while maintaining the attitude by adjusting the motor control signal in the specific situation with a control signal that offsets the disturbance (d) that may actually occur.
14. In paragraph 10, The above second stage policy-based reinforcement learning network, when applied to a reinforcement learning problem, the state S is a three-dimensional position (Px, Py, Pz) and a three-dimensional rotation (roll, pitch, yaw), and the action A is the thrust of each motor (F1, F2, F3, F4). The above sparse Bayesian network is the current state (s t ) and the next state (s t+1 ) as input and the current input (which is the thrust of each motor) ) A method for controlling a drone learned to output.
Citation Information
Patent Citations
Unmanned flying body, information processing method and program
JP2020013340A
Spiking Neural Network with STDP apllied to the threshold of neuron
KR102204107B1
Autonomous flying drone using artificial intelligence neural network
KR102313115B1
Cleaner
KR102448090B1
Drone Terrain Surveillance with Camera and Radar Sensor Fusion for Collision Avoidance
US20180253980A1
Cited By
Unmanned aerial vehicle safety control method based on fault risk learning
CN121680477A