Multi-rotor unmanned aerial vehicle interception control system based on reinforcement learning

By combining image error and drone state information, the multi-rotor drone interception strategy is optimized, and the problems of low interception efficiency and poor stability in the existing technology are solved, real-time and efficient interception in complex environments are achieved.

CN120406551AActive Publication Date: 2025-08-01SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Patent Information

Application Number
CN202510897023.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

The existing image-based visual servo control algorithm lacks environmental adaptability when facing variable maneuvering targets, resulting in low interception efficiency and poor stability, especially in complex dynamic environments that are prone to error accumulation and response lag.

Method used

Combining image error and drone status information, the interception strategy is optimized using the reinforcement learning model, and the control instructions for intercepting the target drone are generated through the IBVS controller and the flight controller, and the reward function of the reinforcement learning model is used to adjust the drone acceleration and attitude in real time to avoid unstable flight state.

Benefits of technology

Real-time and efficient interception of target drones is achieved, the success rate and security of intercepting tasks are improved, and the system's adaptability to complex environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406551A_ABST
    Figure CN120406551A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles, in particular to a multi-rotor unmanned aerial vehicle interception control system based on reinforcement learning. The system comprises an image acquisition module used for acquiring a target image of a target unmanned aerial vehicle; the IBVS controller is used for extracting a target position in the target image, calculating an image error, constructing and training a reinforcement learning model, constructing a current state as input of the reinforcement learning model based on the image error and the state information of the multi-rotor unmanned aerial vehicle, obtaining an expected attitude output by the reinforcement learning model, and outputting the expected attitude to the multi-rotor unmanned aerial vehicle. Calculating an error between the current attitude and the expected attitude to generate a control instruction for intercepting the trajectory of the target unmanned aerial vehicle; and the flight controller is used for receiving the control instruction, converting the control instruction into a pulse width modulation signal and outputting the pulse width modulation signal to motors to adjust the rotating speed of each motor and execute an interception task. The method can dynamically adapt to target movement in a complex environment, and realizes real-time and efficient interception of the target unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unmanned aerial vehicles, and particularly to a multi-rotor unmanned aerial vehicle interception control system based on reinforcement learning. Background Art

[0002] With the opening of low-altitude airspace and the development of small unmanned aerial vehicle technology, non-cooperative targets (such as unauthorized unmanned aerial vehicles) increasingly frequently appear in important airspaces, posing a serious threat to flight safety and public order.

[0003] To address such threats, existing countermeasures mainly include technologies such as radio frequency interference, high-energy weapon strikes, and physical capture. Radio frequency interference technology makes the target unmanned aerial vehicle out of control or return by interfering with its communication link or satellite navigation signal, and has a strong interference effect. High-energy weapons, such as lasers and microwaves, can directly destroy the target, with rapid response and high hit rate. However, these methods generally have risks of environmental pollution, strong electromagnetic interference, and collateral damage to surrounding personnel or equipment. In contrast, the physical capture method intercepts the target by firing a net gun, deploying a capture device, etc., which is pollution-free and more environmentally friendly. However, when facing high-speed or highly maneuverable targets, the interception efficiency of this type of method is low, the success rate is limited, and its effect is limited especially in a dynamically complex environment.

[0004] In recent years, the technology of using small unmanned aerial vehicles equipped with visual sensors as interception platforms has gradually become a research hotspot. Such systems have advantages such as rapid response, low cost, and high safety. In particular, camera-based vision systems have characteristics such as lightweight and multi-functionality, and show broad application prospects in non-cooperative target recognition, tracking, and interception.

[0005] In the prior art, there are papers proposing a method of combining an Image-Based Visual Servoing (IBVS) algorithm with a Proportional Navigation Guidance Law (PNG). In its experiment, the Circular Error Probable (CEP) can reach 0.089 meters, showing good interception accuracy. However, such control algorithms are often designed based on fixed parameters and lack the ability to adapt to target maneuverability and environmental changes, resulting in problems such as error accumulation and response lag in a complex dynamic environment, seriously affecting the interception efficiency and system stability.

[0006] Therefore, how to improve the real-time performance and robustness of such systems in dealing with variable and maneuverable targets remains a key technical problem to be solved urgently. Summary of the Invention

[0007] An embodiment of the present application provides a multi-rotor UAV interception control system based on reinforcement learning, including: An image acquisition module for acquiring a target image of a target UAV; An IBVS controller for extracting the target position in the target image, calculating the image error, constructing and training a reinforcement learning model, constructing the current state as the input of the reinforcement learning model based on the image error and the state information of the multi-rotor UAV, obtaining the expected attitude output by the reinforcement learning model, and calculating the error between the current attitude and the expected attitude to generate a control command for intercepting the target UAV's trajectory; A flight controller for receiving the control command, converting it into a pulse width modulation signal and outputting it to the motor to adjust the rotation speed of each motor, and executing the interception task to ensure that the UAV can execute the interception task according to the control strategy given by the reinforcement learning model; Wherein, the IBVS controller is configured as: The reinforcement learning model includes an actor network and a critic network. The actor network is used to generate actions, and the critic network is used to evaluate the actions generated by the actor network. The reinforcement learning model uses a set reward function to learn and optimize the strategy. The reward function is used to encourage the multi-rotor UAV to achieve the expected acceleration, track the expected attitude and keep the target UAV within the field of view, and punish the maximum acceleration and lift actions close to the physical limit of the multi-rotor UAV, so as to avoid unstable or dangerous states during flight.

[0008] In the above embodiment, by combining the image error and the UAV state information, and using the reinforcement learning model to optimize the interception strategy, it can dynamically adapt to the target movement in a complex environment and achieve real-time and efficient interception of the target UAV. The flexible selection of the image acquisition module (monocular camera or binocular camera) further improves the control accuracy. During flight, the reward function of the reinforcement learning model can adjust the acceleration and attitude of the UAV in real time to avoid unstable or dangerous flight states and improve the success rate and safety of the interception task.

[0009] In some of the embodiments, the reward function is expressed as the following calculation model:

[0010] Wherein, 、 、 、 、 are weight coefficients respectively, is the actual acceleration of the multi-rotor UAV in the earth coordinate system and the square of the Euclidean distance between the expected acceleration , reflecting the motion deviation to make the actual acceleration close to the expected acceleration, is the trace of the matrix, is the transformation matrix from the body coordinate system to the earth coordinate system, which is used to represent the actual attitude of the multi-rotor UAV, is used to represent the desired attitude of the multi-rotor UAV. The desired attitude is used to ensure that the lift direction of the multi-rotor UAV is aligned with the desired acceleration. Through the adjustment of the pitch angle and yaw angle, target tracking is achieved, are the horizontal pixel error and vertical pixel error of the target UAV at the center point of the target image, which are used to keep the target within the field of view of the multi-rotor UAV by reducing the image error, is the normalized desired angular velocity, consists of The three are the components of the desired angular velocity on each axis of the body coordinate system, is the normalized desired lift, which is used to represent the action magnitude and encourage safe control.

[0011] Based on this reward function, by quantifying the acceleration error and attitude error, the system is guided to achieve more accurate acceleration tracking and attitude control. The introduction of the image error term improves the real-time performance and reliability of target tracking. The introduction of the angular velocity and lift limit terms effectively reduces the risk of the aircraft losing control and crashing, and enhances the system's adaptability to complex environments.

[0012] In some embodiments, the following constraint conditions are used to prevent the target UAV from exceeding the field of view angle due to pitch actions or altitude changes: , where is the moving range of the target UAV in the vertical direction of the target image during the process of the multi-rotor UAV intercepting the target UAV, that is, the difference between the maximum value and the minimum value of the vertical pixel error, is a pre-configured vertical movement threshold, and its value is a small constant.

[0013] In some embodiments, the desired attitude is represented by a rotation matrix, which only involves the pitch angle and yaw angle, and the roll angle is ignored to simplify the control, is the tilt rotation matrix used to represent the pitch angle and yaw angle required for tracking the target UAV.

[0014] In some embodiments, constructing and training a reinforcement learning model further includes: Initializing the actor network, critic network, target actor network, and target critic network, initializing the experience replay buffer. The target actor network and target critic network are initialized as copies of the actor network and critic network for stable training; The actor network generates an action based on the state through a deterministic policy and adds noise , and outputs the action , , achieve action exploration by adding noise; Define the environment, which receives actions and updates to output the next state 、 Reward , recycle the experience data to the experience replay buffer for updating the actor network and the critic network; Randomly sample a mini-batch of data from the experience replay buffer, calculate the target Q-value based on the mini-batch of data and the critic network for estimating the expected return of the state-action pair; Update the critic network by minimizing the loss function; Update the parameters of the actor network based on the policy gradient of the critic network; Update the parameters of the target actor network and the target critic network according to the soft update formula; Repeat the above steps until the learning goal is reached.

[0015] By introducing the actor-critic structure, target network, soft update mechanism, experience replay, and policy noise design, the training process of this application has good convergence and stability. Experience replay effectively breaks the temporal correlation between samples and improves the training efficiency; the target network and the soft update mechanism significantly reduce the policy oscillation problem and enhance the training robustness; the overall training architecture can achieve precise policy learning in complex and continuous state spaces and action spaces, enabling the multi-rotor UAV to have real-time and intelligent interception capabilities in dynamic environments, significantly improving the autonomy and execution efficiency of the system.

[0016] In some of the embodiments, constructing and training a reinforcement learning model further includes: Define the state of the multi-rotor UAV and action ; wherein, the state is defined as: , wherein, are the pitch angle, roll angle, and yaw angle of the multi-rotor UAV in the body coordinate system; is the angular velocity vector of the multi-rotor UAV in the body coordinate system; is the expected acceleration component of the multi-rotor UAV in the body coordinate system, is the velocity component of the UAV in the earth coordinate system; the action is defined as: are respectively the multi-rotor UAV in the body coordinate system along axis direction, axial direction, desired angular velocity in the axial direction, is the desired lift force.

[0017] In some of these embodiments, the environment is configured as a dynamic model of a multi-rotor unmanned aerial vehicle: , wherein, used to represent the velocity of the unmanned aerial vehicle in the Earth coordinate system is equal to its velocity vector in the Earth coordinate system; is the actual acceleration of the multi-rotor unmanned aerial vehicle in the Earth coordinate system, is the position vector of the unmanned aerial vehicle in the Earth coordinate system, and the velocity can be calculated based on its position change ; is the attitude update equation, also known as the attitude change rate, used to reflect the attitude of the unmanned aerial vehicle changing over time, which is determined by its angular velocity; is the angular velocity vector of the unmanned aerial vehicle in the body coordinate system, expressed as , used to describe the rotation rate of the unmanned aerial vehicle, is the skew-symmetric matrix of; is the inertia matrix of the multi-rotor unmanned aerial vehicle, describing the inertial characteristics of the unmanned aerial vehicle rotating about each axis, usually a diagonal matrix; is the angular acceleration vector of the unmanned aerial vehicle in the body coordinate system; is the Coriolis moment caused by the angular velocity, reflecting the inertial effect inside the rotating object; is the gyroscopic moment, that is, the additional moment generated by the rotation of the rotor; is the aerodynamic moment, caused by the rotor rotation direction and air resistance, usually part of the control input, in the above formula: .

[0018] In some of these embodiments, the actor network uses a fully connected neural network, including 3 hidden layers, each hidden layer includes 256 neurons, and the activation function uses ReLU.

[0019] In some of these embodiments, the output layer of the actor network uses the tanh activation function for output and action scaling. Specifically: ; .

[0020] In some of these embodiments, after action scaling, the clipping function is used to clip the action, so that: , for limiting the control quantity to avoid overload wherein, the clipping function is expressed as: , then , , is the maximum angular velocity of the multi-rotor UAV is the maximum lift force of the UAV

[0021] In some embodiments, the critic network adopts a fully connected neural network, including 3 hidden layers, each hidden layer includes 256 neurons, the activation function adopts ReLU, wherein, the input state and action are merged after the first hidden layer, and the output layer is set without an activation function

[0022] Compared with the related art, the multi-rotor UAV interception control system provided by the embodiments of the present application combines the image error and the UAV state information, and uses the reinforcement learning model to optimize the interception strategy, which can dynamically adapt to the target movement in a complex environment and achieve real-time and efficient interception of the target UAV; during the flight, the reward function of the reinforcement learning model can adjust the acceleration and attitude of the UAV in real time to avoid unstable or dangerous flight states and improve the success rate and safety of the interception task

[0023] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 is a schematic diagram of the coordinate system of a quadrotor aircraft according to the related art Figure 2 is a structural block diagram of a multi-rotor UAV interception control system according to an embodiment of the present application DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be described and explained below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present application without creative efforts fall within the scope of protection of the present application

[0026] Obviously, the accompanying drawings in the following description are merely some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.

[0027] Reference to "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0028] A quadrotor unmanned helicopter is an underactuated dynamic rotor helicopter with four input forces and six coordinate outputs, and thus it can be known that the system is an autonomous aircraft capable of quasi-static flight (hovering flight and close-range hovering flight). The motion equation of any system is established for a specific reference coordinate system. An unmanned aerial vehicle essentially belongs to a multi-body dynamics system. The motion of the unmanned aerial vehicle body can be regarded as a rigid body motion with six degrees of freedom, including rotation around three axes and linear motion of the center of gravity along three axes.

[0029] The earth coordinate system is an inertial coordinate system, with the earth as a reference and fixed on the ground. It is usually used to describe the position, velocity, and acceleration of a multi-rotor unmanned aerial vehicle and a target in the global environment. The earth coordinate system is represented as: . The z-axis usually points in the vertically upward direction, opposite to the direction of gravity; and The x-axis and y-axis are located in the horizontal plane, and the directions are defined according to specific applications. An example is shown in Figure 1 .

[0030] The body coordinate system is fixed on the body of the multi-rotor unmanned aerial vehicle, and the coordinate origin is located at the center of mass of the unmanned aerial vehicle. It is used to describe the attitude of the unmanned aerial vehicle ( ) and the position of sensors (such as cameras) on the body relative to the body. The body coordinate system is represented as: 。 The x-axis usually points along the forward direction of the body, The y-axis points along the lateral direction of the body, The z-axis points along the vertical direction of the body (usually upward). An example is shown inFigure 1 as shown

[0031] The camera coordinate system is fixed on the monocular camera carried by the multi-rotor UAV. The coordinate origin is located at the optical center of the camera and is used to describe the three-dimensional position of the target in the camera view and the relative attitude of the camera with respect to the airframe. The camera coordinate system is expressed as: 。 The z-axis generally points vertically upward, opposite to the direction of gravity; and The x and y axes are in the horizontal plane, and their directions are defined according to specific applications.

[0032] This application controls the multi-rotor UAV based on the principle of proportional navigation. Its basic idea is to make the rate of change of the velocity direction of the UAV proportional to the rate of change of the line-of-sight direction :

[0033] : The velocity direction angle of the UAV represents the direction of the velocity vector in space.

[0034] : The line-of-sight direction angle represents the direction of the line of sight from the UAV to the target.

[0035] : The proportional navigation constant controls the sensitivity of navigation.

[0036] When the three-dimensional space is decomposed into a two-dimensional horizontal plane and a vertical plane, the changes in the line-of-sight direction and the velocity direction can be processed separately in the horizontal and vertical planes. Let: The line-of-sight angle in the vertical plane is: , , The line-of-sight angle in the horizontal plane: , , where is the line-of-sight direction vector, The velocity angle in the vertical plane: ,

[0037] The velocity angle in the horizontal plane: ,

[0038] where is the velocity direction vector of the UAV.

[0039] An embodiment of the present application provides a multi-rotor UAV interception control system based on reinforcement learning. Referring to Figure 2 as shown, the system includes an image acquisition module, an IBVS controller, and a flight controller that are electrically connected, and are specifically configured as follows: The image acquisition module is used to acquire the target image of the target UAV. Optionally, the image acquisition module uses a monocular camera, such as a Jetson CSI camera. In another embodiment, the image acquisition module uses a binocular camera to provide richer depth information, thereby enhancing the accuracy of the interception control. The IBVS controller is used to extract the target position in the target image, calculate the image error, construct and train a reinforcement learning model, construct the current state based on the image error and the state information of the multi-rotor UAV as the input of the reinforcement learning model, obtain the desired attitude output by the reinforcement learning model, and calculate the error between the current attitude and the desired attitude to generate a control command for intercepting the target UAV's trajectory. The reinforcement learning model generates the desired angular velocity and the desired lift as the goals of the UAV flight control according to the state and the image error of the UAV. The IBVS controller calculates the errors between these values and the current attitude of the UAV, such as the pitch angle error, the yaw angle error, the roll angle error, etc., and generates control commands for adjusting the angular velocity (pitch angle, yaw angle, roll angle) to adjust the attitude and trajectory of the UAV to achieve the interception task.

[0040] The flight controller is used to receive the control command, convert it into a pulse width modulation signal, and output it to the motor to adjust the speed of each motor, and execute the interception task to ensure that the UAV can execute the interception task according to the control strategy given by the reinforcement learning model. Optionally, the flight controller uses Pixhawk, which can accurately control the speed of the motor and ensure that the UAV can stably execute the desired flight path. Among them, the IBVS controller is configured as follows: The reinforcement learning model includes an actor network and a critic network. The actor network is used to generate actions, and the critic network is used to evaluate the actions generated by the actor network. The reinforcement learning model uses a set reward function to learn and optimize the strategy. The reward function is used to encourage the multi-rotor UAV to achieve the desired acceleration, track the desired attitude, and keep the target UAV within the field of view, and punish the maximum acceleration and lift actions close to the physical limit of the multi-rotor UAV, so as to avoid unstable or dangerous states during the flight process.

[0041] In the above embodiments, by combining the image error and the UAV state information and using the reinforcement learning model to optimize the interception strategy, it is possible to dynamically adapt to the target movement in a complex environment and achieve real-time and efficient interception of the target UAV. The flexible selection of the image acquisition module (monocular camera or binocular camera) further improves the control accuracy. During the flight, the reward function of the reinforcement learning model can adjust the acceleration and attitude of the UAV in real time to avoid unstable or dangerous flight states and improve the success rate and safety of the interception mission.

[0042] In some of these embodiments, the reward function is expressed as the following calculation model:

[0043] Where , , , , are weight coefficients respectively, is the actual acceleration of the multi-rotor UAV in the earth coordinate system and the desired acceleration The square of the Euclidean distance, which reflects the motion deviation to make the actual acceleration close to the desired acceleration, is the trace of the matrix, is the transformation matrix from the body coordinate system to the earth coordinate system, which is used to represent the actual attitude of the multi-rotor UAV, is used to represent the desired attitude of the multi-rotor UAV. The desired attitude is used to ensure that the lift direction of the multi-rotor UAV is aligned with the desired acceleration, and target tracking is achieved through pitch and yaw adjustments, is the horizontal pixel error and vertical pixel error of the target UAV at the center point of the target image, which is used to keep the target within the field of view of the multi-rotor UAV by reducing the image error, is the normalized desired angular velocity, is composed of the desired angular velocity The three are the components of the desired angular velocity on each axis of the body coordinate system, is the normalized desired lift, which is used to represent the action size and encourage safe control.

[0044] Based on this reward function, by quantifying the acceleration error and attitude error, the system is guided to achieve more precise acceleration tracking and attitude control. The introduction of the image error term improves the real-time performance and reliability of target tracking. The introduction of the angular velocity and lift limit terms effectively reduces the risk of the aircraft losing control and crashing, and enhances the system's adaptability to complex environments.

[0045] In the above embodiments, the actual acceleration of the multi-rotor UAV is jointly determined by the gravitational acceleration and the lift generated by the rotors. The lift is controllable and is used to adjust the movement trajectory of the UAV. , is the gravitational acceleration vector, usually expressed as , where , the direction is vertically downward, is the mass of the multi-rotor UAV, is the lift vector generated by the rotors of the multi-rotor UAV in the Earth coordinate system, and the direction is usually opposite to the axis of the body coordinate system, , is the desired velocity, indicating the acceleration applied by the multi-rotor UAV from the current velocity to the desired velocity .

[0046] In the above embodiment, the desired velocity vector is composed of the velocity magnitude and the normalized direction vector , and is used to drive the multi-rotor UAV to fly along the proportional navigation trajectory. The velocity magnitude is dynamically adjusted according to the target distance and the UAV performance, defines the flight direction of the UAV: , , are the desired vertical plane velocity angle and horizontal plane velocity angle, which are calculated through the line-of-sight angle change between the current moment and the previous moment : , to ensure that the UAV gradually approaches the target.

[0047] In the above embodiment, the camera uses the pinhole model to project the target in the three-dimensional space onto the two-dimensional image plane: , , , ([[]] is the pixel coordinate of the target UAV in the target image, and its origin is set at the upper left corner of the target image, ([[]] is the central pixel coordinate of the target image, then the image error can be expressed as: .

[0048] In some of these embodiments, to facilitate the multi-rotor UAV to track the target through yaw angle adjustment and prevent large movements of the target UAV in the vertical direction of the target image, the following constraint conditions are used to avoid the target UAV exceeding the field of view angle FOV due to pitch movement or altitude change: , where, is the movement range of the target UAV in the vertical direction of the target image during the process of the multi-rotor UAV intercepting the target UAV, that is, the difference between the maximum and minimum values of the pixel error in the vertical direction, is a pre-configured vertical movement threshold, and its value is a small constant.

[0049] In the Earth coordinate system, using the line-of-sight direction vector is defined as the normalized relative position vector of the multi-rotor UAV for interception pointing to the target aircraft, which is calculated from the perspective of the relative position and can also be specifically expressed as: The line-of-sight direction can be calculated from the position of the target UAV in the target image and coordinate transformation, and is expressed as the following formula: where and is the transformation matrix from the body coordinate system to the camera coordinate system.

[0050] The above-mentioned rate of change of the image error is related to the instantaneous motion vector of the camera in the camera coordinate system and is described by the Jacobian matrix as: where the instantaneous motion vector is: including the linear velocity and the angular velocity ; The Jacobian matrix maps the six-degree-of-freedom motion of the camera to the rate of change of the image error, thus facilitating the connection from image feedback to motion control. The Jacobian matrix is expressed as: ; where is the depth of the target UAV along the optical axis in the camera coordinate system; is the normalized image error, is the camera focal length, .

[0051] In some of the embodiments, the desired attitude is represented by a rotation matrix, only involving the pitch angle and the yaw angle, and ignoring the roll angle to simplify the control. is the tilt rotation matrix used to represent the pitch angle and yaw angle required for tracking the target UAV, where: is the identity matrix; is the cross product of the current lift direction and the desired lift direction representing the axis of rotation; : the current lift direction and the desired lift direction The included angle between them represents the rotation angle; the desired lift direction is the desired orientation of the axis of the multi-rotor UAV body coordinate system, which is used to constrain the current lift direction to be aligned with the desired lift direction The desired lift direction is the net acceleration after considering gravity compensation, which is expressed as: ; : the rotation axis skew-symmetric matrix of, which is used to calculate the rotation matrix; , , , is the rotation matrix about the axis of the body coordinate system.

[0052] In the above embodiment, the present application adjusts the pitch angle and yaw angle by configuring an angular velocity controller, so that the actual attitude quickly approaches the desired attitude , and at the same time, in combination with the requirement of field of view maintenance, the angular velocity controllers of the pitch angle and yaw angle are set as: , where the vex( ) function is used to extract the corresponding three-dimensional vector from a three-dimensional skew-symmetric matrix.

[0053] Further, an input of a field of view maintenance controller is configured in combination with the above angular velocity controller: , where is the yaw angular velocity generated by the field of view maintenance controller, is the saturation function, which is defined as: .

[0054] To ensure the stability of attitude control, the present application is based on the above Lyapunov candidate function: , where is the Frobenius norm.

[0055] In some of the embodiments, constructing and training a reinforcement learning model further includes: Step S1: Define the state of the multi-rotor UAV, the action , and define the environment; Step S2: Initialize the actor network, the critic network, the target actor network, and the target critic network. Specifically: Initialize the parameters of the actor network and the parameters of the critic network as well as the parameters of the target actor network and the target critic network 、 , initialize the experience replay buffer, and initialize the target actor network and the target critic network as copies of the actor network and the critic network for stable training; Step S3: The current state of the actor network generates an action from the state through a deterministic policy and adds noise , and outputs the action . , Action exploration is achieved by adding noise. Optionally, the noise is Ornstein-Uhlenbeck noise, and the noise satisfies: ; is the regression rate, is the mean (usually 0), is the volatility, is the Brownian motion.

[0056] Step S4: The environment receives the action and updates to output the next state 、 and the reward , and recovers the experience data to the experience replay buffer for updating the actor network and the critic network; Step S5: Randomly sample a mini-batch of data from the experience replay buffer, Step S6: Calculate the target Q-value based on the mini-batch data and the critic network to estimate the expected return of the state-action pair. The target Q-value is expressed as: , i where is the sample index; Step S7: Update the critic network by minimizing the loss function, and the loss function is expressed as: ; Step S8: Update the parameters of the actor network based on the policy gradient of the critic network. The policy gradient uses the gradient ascent method, which is expressed as: ; Step S9: Update the parameters of the target actor network and the target critic network according to the soft update formula, and the soft update formula is expressed as:

[0057]

[0058] in, is the soft update rate, which is 0.001; Step S10: Repeat steps S3 to S9 until the learning goal is achieved; Wherein, the state is defined as:

[0059] ,in, The pitch angle, roll angle and yaw angle of the multi-rotor UAV in the body coordinate system; is the angular velocity vector of the multi-rotor UAV in the body coordinate system; is the expected acceleration component of the multi-rotor UAV in the body coordinate system, is the velocity component of the UAV in the Earth coordinate system; The action is defined as: They are the coordinates of the multi-rotor UAV along the body coordinate system. Axis direction, Axis direction, The desired angular velocity in the axis direction, namely the roll angle, pitch angle, and yaw angle, The expected lift is used to ensure that the multirotor drone generates sufficient thrust to achieve the desired acceleration while being limited within the hardware capabilities of the multirotor drone. It can be calculated based on the following formula: , is the expected lift force considering gravity compensation.

[0060] In the above embodiment, the environment is configured as a dynamic model of a multi-rotor drone: , in, Used to represent the speed of the drone in the earth coordinate system is equal to its velocity vector in the Earth coordinate system; is the actual acceleration of the multi-rotor UAV in the earth coordinate system, is the position vector of the UAV in the Earth coordinate system. The speed can be calculated based on its position change. ; is the attitude update equation, also known as the attitude change rate, which is used to reflect that the attitude change of the drone over time is determined by its angular velocity; is the angular velocity vector of the UAV in the body coordinate system, expressed as , used to describe the rotation rate of the drone, is the skew-symmetric matrix of the angular velocity vector, is the inertia matrix of the multi-rotor UAV, which describes the inertia characteristics of the UAV rotating around each axis and is usually a diagonal matrix; is the angular acceleration vector of the UAV in the body coordinate system; is the Coriolis torque caused by the angular velocity, reflecting the inertial effect inside the rotating object; is the gyroscopic torque, that is, the additional torque generated by the rotation of the rotor; is the aerodynamic torque, caused by the rotor rotation direction and air resistance, and is usually part of the control input; 。

[0061] In the above embodiments, is the angular velocity update equation, which is constructed based on the Euler dynamics equation Considering that the external torque (such as the control torque generated by the rotor) and the internal inertial effect jointly determine the rotation behavior of the UAV, therefore, the angular velocity update equation of this embodiment is determined by the combined action of multiple torques.

[0062] Based on the above configuration, the state space of the embodiments of the present application comprehensively characterizes the core information such as the flight attitude, desired acceleration, visual error, and geographical velocity of the multi-rotor UAV at any moment. By capturing the pitch angle, roll angle, yaw angle, and the corresponding angular velocity, the system can reflect the current attitude and rotation trend. The desired acceleration reflects the requirements of the mission objective for maneuverability, and the image error ensures that the target remains at the center of the field of view continuously, while the velocity vector can be used for state dynamic modeling and navigation estimation. The desired angular velocity in the action space controls the attitude adjustment of three rotational degrees of freedom, and the desired lift determines the thrust output in the vertical direction. The calculation process of the desired lift is clipped by upper and lower limits to avoid the lift command exceeding the actual flight control capabilities, which not only ensures the feasibility of the flight action but also guarantees the safety and stability of the flight process. The above configuration provides a clear input and output interface for the reinforcement learning strategy, which helps to construct a stable and convergent training process.

[0063] The environment of the present application constructs a dynamic model of the multi-rotor UAV based on the position update (i.e., velocity), velocity update (i.e., acceleration), attitude update, and angular velocity update equation of the UAV, which can accurately simulate the complex dynamic response of the multi-rotor UAV during flight. It is the key basis for the reinforcement learning algorithm training to interact with the environment, generate state transitions, calculate the reward function, and predict the effect.

[0064] The actor network is used to generate flight actions, and the critic network is responsible for evaluating the value of the actions in a given state (i.e., the Q-value). The two work together to build a policy optimization mechanism. The experience replay mechanism records the state transitions and reward information during the environmental interaction process, reducing the correlation between training samples and facilitating the improvement of learning stability and efficiency. The introduced Ornstein-Uhlenbeck noise has temporal correlation, which can enhance the exploration of the policy in the continuous action space and prevent getting stuck in local optima. In each round of training, the model updates the critic network by batch sampling the experience data to fit the target Q-value, and then improves the action quality of the actor network through policy gradients. The soft update mechanism of the target network provides a smooth transition for policy updates, thus avoiding the instability of the training process caused by frequent changes in the target network. This mechanism can achieve a complete training closed-loop from state-action mapping to action value evaluation and then to policy adjustment.

[0065] By introducing the actor-critic structure, target network, soft update mechanism, experience replay, and policy noise design, the training process of this application has good convergence and stability. Experience replay effectively breaks the temporal correlation between samples and improves training efficiency; the target network and soft update mechanism significantly reduce the policy oscillation problem and enhance training robustness; the Ornstein-Uhlenbeck noise enables effective exploration of the action space, improves the policy search ability, and helps the model obtain a better flight control strategy. The overall training architecture can achieve precise policy learning in complex and continuous state spaces and action spaces, enabling the multi-rotor UAV to have real-time and intelligent interception capabilities in dynamic environments, and significantly improving the autonomy and execution efficiency of the system.

[0066] In some embodiments, the actor network adopts a fully connected neural network, including 3 hidden layers, each hidden layer including 256 neurons, and the activation function adopts ReLU.

[0067] In some embodiments, the output layer of the actor network uses the tanh activation function for output and action scaling. Specifically: ; .

[0068] In some embodiments, after action scaling, the clipping function is used to clip the actions so that: , , which is used to limit the control amount to avoid overload, where the clipping function is expressed as: , then , , is the maximum angular velocity of the multi-rotor UAV, is the maximum lift of the UAV.

[0069] In some of these embodiments, the critic network employs a fully-connected neural network, including 3 hidden layers, each hidden layer including 256 neurons, with the activation function being ReLU. Among them, the input state and action are merged after the first hidden layer, and the output layer is set without an activation function.

[0070] It should be noted that the above-mentioned various modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned various modules can be located in the same processor; or the above-mentioned various modules can also be located in different processors in any combined form.

[0071] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A multi-rotor UAV interception control system based on reinforcement learning, characterized in that, Including: An image acquisition module, configured to acquire a target image of a target UAV; An IBVS controller, configured to extract a target position in the target image, calculate an image error, construct and train a reinforcement learning model, construct a current state as an input to the reinforcement learning model based on the image error and the state information of the multi-rotor UAV, obtain an expected attitude output by the reinforcement learning model, and calculate an error between the current attitude and the expected attitude to generate a control command for intercepting the trajectory of the target UAV; A flight controller, configured to receive the control command, convert it into a pulse width modulation signal and output it to the motor to adjust the rotational speed of each motor, and execute an interception task; Wherein, the IBVS controller is configured as: The reinforcement learning model includes an actor network and a critic network. The actor network is used to generate actions, and the critic network is used to evaluate the actions generated by the actor network. The reinforcement learning model uses a set reward function to learn and optimize the policy. The reward function is used to encourage the multi-rotor UAV to achieve an expected acceleration, track an expected attitude and keep the target UAV within the field of view, and punish the maximum acceleration and lift actions close to the physical limit of the multi-rotor UAV.

2. The multi-rotor UAV interception control system based on reinforcement learning according to claim 1, characterized in that, The reward function is expressed as the following calculation model: Among them, , , , , are weight coefficients respectively, is the actual acceleration of the multi-rotor UAV in the Earth coordinate system and the expected acceleration is the square of the Euclidean distance between them, is the trace of the matrix, is the transformation matrix from the body coordinate system to the Earth coordinate system, which is used to represent the actual attitude of the multi-rotor UAV, is used to represent the expected attitude of the multi-rotor UAV, are the horizontal pixel error and vertical pixel error of the target UAV at the center point of the target image, is the normalized expected angular velocity, is the normalized expected lift.

3. The multi-rotor UAV interception control system based on reinforcement learning according to claim 2, wherein The following constraint conditions are used to prevent the target UAV from exceeding the field of view angle due to pitch movement or altitude change: , Among them, is the moving range of the target UAV in the vertical direction of the target image during the process of the multi-rotor UAV intercepting the target UAV, that is, the difference between the maximum value and the minimum value of the pixel error in the vertical direction. is a pre-configured vertical movement threshold.

4. The multi-rotor UAV interception control system based on reinforcement learning according to claim 3, characterized in that, Desired attitude Represented by a rotation matrix, is an inclined rotation matrix constructed to represent the pitch angle and yaw angle required for the tracked target UAV.

5. The multi-rotor UAV interception control system based on reinforcement learning according to any one of claims 2 to 4, characterized in that Constructing and training a reinforcement learning model further includes: Initializing the actor network, the critic network, the target actor network, and the target critic network, and initializing an experience replay buffer; The actor network is based on the state generates actions through a deterministic policy and adds noise to output the actions , ; Define the environment, which receives actions and updates the output of the next state 、 Reward , and recycle the experience data to the experience replay buffer; Randomly sampling a small batch of data from the experience replay buffer; Calculating a target Q value based on the small batch of data and the critic network; Updating the critic network by minimizing a loss function; Updating the parameters of the actor network based on the policy gradient of the critic network; Updating the parameters of the target actor network and the target critic network according to a soft update formula; Repeating the above steps until a learning target is reached.

6. The multi-rotor UAV interception control system based on reinforcement learning according to claim 5, wherein, Constructing and training a reinforcement learning model further includes: Define the state of the multi-rotor UAV , actions ; Wherein, the state is defined as: , Wherein, are the pitch angle, roll angle and yaw angle of the multi-rotor UAV in the body coordinate system; is the angular velocity vector of the multi-rotor UAV in the body coordinate system; are the expected acceleration components of the multi-rotor UAV in the body coordinate system, are the velocity components of the UAV in the earth coordinate system; The said actions are defined as: They are respectively the desired angular velocities of the multi-rotor UAV in the body coordinate system along the axis direction, direction, direction, and is the desired lift force.

7. The multi-rotor UAV interception control system based on reinforcement learning according to claim 6, wherein, The environment is configured as a dynamic model of a multi-rotor UAV: , Among them, used to represent the velocity of the UAV in the Earth coordinate system is equal to its velocity vector in the Earth coordinate system; is the actual acceleration of the multi-rotor UAV in the Earth coordinate system; is the attitude change rate, is the angular velocity vector of the UAV in the body coordinate system, expressed as , is the skew-symmetric matrix of; is the inertia matrix of the multi-rotor UAV; is the angular acceleration vector of the UAV in the body coordinate system; is the Coriolis moment caused by the angular velocity; is the gyroscopic moment; is the aerodynamic moment.

8. The multi-rotor UAV interception control system based on reinforcement learning according to claim 5, characterized in that, The output layer of the actor network uses a tanh activation function for output and action scaling.

9. The multi-rotor UAV interception control system based on reinforcement learning according to claim 8, characterized in that, After the action is scaled, use the clipping function to clip the action.

10. The multi-rotor UAV interception control system based on reinforcement learning according to claim 5, wherein, The actor network and the critic network adopt fully connected neural networks, each including three hidden layers.

Citation Information

Patent Citations

  • Four-rotor unmanned aerial vehicle route following control method based on deep reinforcement learning

    CN110673620A

  • Quad-rotor unmanned aerial vehicle trajectory control method based on reinforcement learning

    CN112650058A

  • Unmanned ship missile interception and avoidance algorithm and system based on reinforcement learning and readable storage medium

    CN117666589A

  • Hypersonic aircraft attitude control method based on robust adversarial reinforcement learning

    CN118567386A

  • Reinforcement learning guidance control integration method for intercepting three-dimensional maneuvering target

    CN118938676A

Cited By

  • Control method and system of grain scraping conveyor

    CN120817452A

  • A control method and system for a paddy conveyer

    CN120817452B

  • Unmanned aerial vehicle autonomous tracking method and system based on monocular vision

    CN121596906A