Unmanned aerial vehicle take-off and landing control device and method based on action recognition

Through the drone take-off and landing control device and method based on action recognition, the operator's actions are recognized by onboard video images and the control instructions are sent, the problem that the drone cannot be effectively controlled in complex environments or inclement weather is solved, and the environmental perception ability and flight safety of the drone are improved.

CN119960469APending Publication Date: 2025-05-09CHINESE AERONAUTICAL RADIO ELECTRONICS RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983983.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-09

Smart Images

  • Figure CN119960469A_ABST
    Figure CN119960469A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle take-off and landing control device and method based on action recognition, and the device comprises an action recognition assembly which is used for recognizing the action of an operator; the unmanned aerial vehicle take-off and landing control assembly is used for completing instruction conversion on the operator actions recognized by the action recognition assembly according to a corresponding instruction set and sending the instruction to a flight control system, and the flight control system controls take-off and landing of the aircraft based on the instruction; according to the method, action recognition is completed through the video image based on the airborne image, medium-free control instruction sending is achieved, takeoff, landing and sliding control of the unmanned aerial vehicle under the complex environment, severe weather or link-free communication conditions is completed, and the environment perception ability of the unmanned aerial vehicle and the safety of sliding, takeoff, landing and flying of the unmanned aerial vehicle can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of unmanned aerial vehicle control, and in particular, relates to a device and method for controlling the take-off and landing of an unmanned aerial vehicle based on motion recognition. Background Art

[0002] The operation of drones during the taxiing, take-off and landing phases relies on established plans and the judgment and control made by the aircraft operator through the camera. In complex environments or severe weather conditions, due to the limitations of environmental changes and environmental perception capabilities, it is impossible to make reliable judgments in a timely and effective manner. Summary of the invention

[0003] The purpose of the present invention is to: when the link communication is lost, the pilot cannot issue effective control instructions, and at this time it is necessary to rely on the autonomous environmental perception and decision-making capabilities of the UAV. At this time, a method for sending control instructions to the UAV without a transmission medium is required.

[0004] In a first aspect, the present application provides a drone take-off and landing control device based on motion recognition, the device comprising:

[0005] A motion recognition component for recognizing the operator's motion;

[0006] The UAV take-off and landing control component is used to convert the operator's actions recognized by the action recognition component into instructions according to the corresponding instruction set, and send the instructions to the flight control system, and the flight control system controls the aircraft take-off and landing based on the instructions.

[0007] Preferably, the action recognition component includes:

[0008] The target tracking module is used to receive the airborne video image, detect the target image frame by frame, obtain the pixel position of the target in the image, associate the target of the current frame image with the target of the previous frame image through the Kalman filter, realize target detection and association, and output the target image;

[0009] The limb skeleton extraction module is used to detect and infer key points of the detected target image and output the operator skeleton information;

[0010] A feature extraction module, used for denoising the operator skeleton information and extracting features of the skeleton information;

[0011] The action recognition module is used to perform continuous action recognition using an LSTM classifier based on the operator's skeleton information and skeleton feature information, and output the recognized operator actions.

[0012] Preferably, the target tracking module is further used to use a YOLOX module to perform target detection, input a video image onboard the drone, and output pixel coordinates of the operator image in the video image. The SORT algorithm is used to complete the tracking of the detection frame.

[0013] Preferably, the limb skeleton extraction module is also used to adopt the RTMpose key point detection method, input the operator image obtained by the target tracking module, and obtain the operator's skeleton information through corresponding reasoning.

[0014] Preferably, the feature extraction module is also used to denoise the data to eliminate interference caused by special data. Then we extract the features of the skeleton. For the feature extraction of skeleton information, commonly used features include the skeleton's limb angle, node relative distance, joint kinetic energy and acceleration of joint angle.

[0015] Preferably, the instruction set includes a take-off and landing control instruction set and a taxiing control instruction set.

[0016] In a second aspect, the present application also provides a method for controlling the take-off and landing of a UAV based on motion recognition, the method comprising:

[0017] The motion recognition component recognizes the operator's motion;

[0018] The drone take-off and landing control component converts the operator action recognized by the action recognition component into an instruction according to a corresponding instruction set, and sends the instruction to the flight control system, and the flight control system controls the aircraft take-off and landing based on the instruction; wherein the instruction set includes a take-off and landing control instruction set and a taxiing control instruction set.

[0019] Preferably, the action recognition component recognizes the action of the operator, including:

[0020] The target tracking module receives the airborne video image, detects the target image frame by frame, obtains the pixel position of the target in the image, associates the target in the current frame image with the target in the previous frame image through the Kalman filter, realizes target detection and association, and outputs the target image;

[0021] The limb skeleton extraction module performs key point detection and reasoning on the detected target image and outputs the operator skeleton information;

[0022] The feature extraction module denoises the operator skeleton information and extracts the features of the skeleton information;

[0023] The action recognition module uses an LSTM classifier to perform continuous action recognition based on the operator's skeleton information and skeleton feature information, and outputs the recognized operator actions.

[0024] Beneficial technical effects of the present invention:

[0025] The method provided in this application completes action recognition through video images based on airborne images, realizes the sending of control instructions without a medium, and completes the take-off, landing and taxiing control of drones in complex environments, bad weather or without link communication. It can improve the environmental perception ability of drones and the safety of taxiing and take-off and landing flights of drones. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A skeleton extraction flowchart provided for an embodiment of the present application. DETAILED DESCRIPTION

[0027] See also Figure 1 The content of the present invention includes an action recognition component, a UAV take-off and landing control component, a take-off and landing control instruction set and a taxiing control instruction set.

[0028] Among them, the action recognition component includes a target tracking module, a limb skeleton extraction module, a feature extraction module and an action recognition module.

[0029] Among them, the target tracking module:

[0030] The present invention uses the YOLOX module for target detection, inputs the video image on the drone, and outputs the pixel coordinates of the operator image in the video image. The SORT (Simple Online and Realtime Tracking) algorithm is used to complete the tracking of the detection frame. The main steps of the SORT algorithm include:

[0031] 1) Perform object detection on each input frame, using the YOLOX module result as input. The result is a four-tuple (x, y, w, h), where (x, y) is the coordinate of the center point of the bounding box, and w and h are the width and height of the bounding box;

[0032] 2) Kalman filter prediction: An associated Kalman filter is set for each tracked target, and the Kalman filter is used to predict the position of the target in the new frame;

[0033] 3) Data association: associate the new detection results with the old tracking targets, calculate the distance between each old tracking target and the new detection result, and then use the Hungarian algorithm to find the match with the minimum distance, using the IoU (Intersection over Union) distance;

[0034] 4) Kalman filter update: Once a match is found, the state of the Kalman filter is updated using the new detection results.

[0035] Among them, the limb skeleton extraction module:

[0036] The present invention adopts the RTMpose key point detection method, inputs the operator image obtained by the target tracking module, and obtains the operator's skeleton information through corresponding reasoning.

[0037] Feature extraction module:

[0038] First, we denoise the data to eliminate the interference caused by special data, and then we extract the features of the skeleton. For the feature extraction of skeleton information, the commonly used features include the skeleton's limb angle, node relative distance, joint kinetic energy, and acceleration of the joint angle.

[0039] a) Data Denoising

[0040] The mean filtering algorithm is used for the skeleton data. The selected window is a time window. For the data at a certain moment, the filtered value is the average value of the sum of the data in the previous period and the data in the period after the moment. For each limb skeleton key point, each joint point can be represented as a two-dimensional vector (x, y). For the gesture key point, it is represented as a three-dimensional vector (x, y, z). Taking the x dimension of the skeleton data as an example, the calculation formula of the filtering process is as follows:

[0041]

[0042] For the mth frame, the filtering process is the mean of the sum of the N / 2 frames before and after the current frame.

[0043] Pm,x represents the mth frame of the x dimension of a joint point. After the original skeleton data is processed by mean filtering, the influence of mutation points can be basically eliminated. It should be noted that the limb skeleton extracted in this project is 2D, so only the x and y axes need to be denoised. The extracted gesture skeleton is 3D, so the z axis needs to be denoised additionally. After the original data is processed by mean filtering, smooth human skeleton data can be obtained, and then the limb angle and relative distance of each frame are calculated.

[0044] b) Limb angle

[0045] For the angle feature, calculate the angle between each joint point and its adjacent joint point. For every three joint points, the angle calculation needs to be performed. Arrange all the angles in order, so that effective angle information describing the posture is obtained, and the angle feature is represented as A.

[0046] c) Relative distance

[0047] Taking into account the height differences among individuals, the normalization idea is adopted when calculating the relative distance, and the eight groups of distances obtained are uniformly divided by the distance D between the clavicle joint and the waist joint.

[0048] Arrange all the distances in order, so that effective spatial information describing the posture is obtained, and the angle feature is represented as B.

[0049] d) Feature Fusion

[0050] When fusion of features is performed, the limb angle feature A and the relative distance feature B are directly spliced, so that the skeleton sequence can be described as:

[0051] x = [A, B]

[0052] Where x is the skeleton feature obtained from each frame of the image, and all the skeleton features captured by a window are represented as the serialized features X[x1, x2, ..., xt] of the input recognition module.

[0053] Among them, the action recognition module:

[0054] The present invention selects LSTM as a classifier. The input is represented as a sequence X[x1, x2, ..., xt], where each xi is the skeleton sequence feature of each sampling frame. The length t of each window is the time span for each detection and recognition, and t=80 is selected as the window length to ensure that the model can take into account the time dependency in a continuous action. The model first sends the input sequence to the LSTM layer for processing, and then aggregates the output of the LSTM to obtain a total output result. Afterwards, this output result is processed by the softmax function to obtain the predicted probability of each instruction. The category with the highest predicted probability will be used as the action instruction detected by this window, and this action instruction corresponds to the behavior of the operator during the time period.

[0055] In order to achieve continuous action recognition, we use a sliding window method to process the input sequence. Each window is a subsequence in the time series, and these windows overlap. 25 is selected as the step size to balance the computational efficiency and the accuracy of action recognition. When the classification model determines that the input sequence is a valid instruction, it will be compared with the instruction label recognized by the previous sequence. If it is a new instruction, the drone will be controlled according to the new instruction.

[0056] The drone take-off and landing control component mainly converts and sends the operator's actions identified by the action recognition component according to the corresponding instruction set.

[0057] Table 1 UAV take-off and landing / taxi control instruction set

[0058]

Claims

1. A UAV take-off and landing control device based on motion recognition, characterized in that: The device comprises: A motion recognition component for recognizing the operator's motion; The UAV take-off and landing control component is used to convert the operator's actions recognized by the action recognition component into instructions according to the corresponding instruction set, and send the instructions to the flight control system, and the flight control system controls the aircraft take-off and landing based on the instructions.

2. The device according to claim 1, characterized in that The action recognition component includes: The target tracking module is used to receive the airborne video image, detect the target image frame by frame, obtain the pixel position of the target in the image, associate the target of the current frame image with the target of the previous frame image through the Kalman filter, realize target detection and association, and output the target image; The limb skeleton extraction module is used to detect and infer key points of the detected target image and output the operator skeleton information; A feature extraction module, used for denoising the operator skeleton information and extracting features of the skeleton information; The action recognition module is used to perform continuous action recognition using an LSTM classifier based on the operator's skeleton information and skeleton feature information, and output the recognized operator actions.

3. The device according to claim 2, characterized in that The target tracking module is also used to perform target detection using the YOLOX module, input the video image on the drone, and output the pixel coordinates of the operator image in the video image. The SORT algorithm is used to complete the tracking of the detection frame.

4. The device according to claim 2, characterized in that The limb skeleton extraction module is also used to adopt the RTMpose key point detection method, input the operator image obtained by the target tracking module, and obtain the operator's skeleton information through corresponding reasoning.

5. The device according to claim 2, characterized in that The feature extraction module is also used to denoise the data to eliminate the interference caused by special data. Then we extract the features of the skeleton. For the feature extraction of skeleton information, commonly used features include the skeleton's limb angle, node relative distance, joint kinetic energy and acceleration of joint angle.

6. The device according to claim 1, characterized in that The instruction set includes a take-off and landing control instruction set and a taxiing control instruction set.

7. A method for controlling the take-off and landing of an unmanned aerial vehicle based on motion recognition, characterized in that: The method comprises: The motion recognition component recognizes the operator's motion; The drone take-off and landing control component converts the operator action recognized by the action recognition component into an instruction according to a corresponding instruction set, and sends the instruction to the flight control system, and the flight control system controls the aircraft take-off and landing based on the instruction; wherein the instruction set includes a take-off and landing control instruction set and a taxiing control instruction set.

8. The method according to claim 7, characterized in that The action recognition component recognizes the action of the operator, including: The target tracking module receives the airborne video image, detects the target image frame by frame, obtains the pixel position of the target in the image, associates the target in the current frame image with the target in the previous frame image through the Kalman filter, realizes target detection and association, and outputs the target image; The limb skeleton extraction module performs key point detection and reasoning on the detected target image and outputs the operator skeleton information; The feature extraction module denoises the operator skeleton information and extracts the features of the skeleton information; The action recognition module uses an LSTM classifier to perform continuous action recognition based on the operator's skeleton information and skeleton feature information, and outputs the recognized operator actions.