Distributed UAV Cooperative Motion Control Method and Device Based on Pure Vision
By carrying on-board sensors and perceived motor neural networks on the drone, independent drone cluster motion and obstacle avoidance for wireless communication are achieved, solving the reliability and cost problems of drone motion control in complex environments in the prior art.
Patent Information
- Application Number
- CN202211690249.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-12-27
AI Technical Summary
Existing multi-UAV systems are difficult to achieve independent cluster motion and obstacle avoidance for wireless communication in complex and unknown environments, and have problems with low accuracy, high cost and reliability.
The distributed drone collaborative motion control method based on pure vision is adopted to obtain environmental information through onboard sensors, and control instructions are generated using perceived motion neural networks to realize cluster motion and obstacle avoidance of drones in complex environments.
Without relying on wireless communication, the efficient cluster movement and obstacle avoidance of drones in complex environments is achieved, which improves the autonomy and reliability of the system and reduces costs.
Smart Images

Figure CN116009583B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of unmanned aerial vehicles, and particularly to a distributed unmanned aerial vehicle cooperative motion control method and device based on pure vision. Background Art
[0002] With the rapid development of technology, unmanned aerial vehicles have been widely used in various fields. However, with the increasingly complex application environment and diverse mission requirements, a single unmanned aerial vehicle has significant limitations in both hardware and software and cannot meet the relevant needs. In contrast, a multi-unmanned aerial vehicle system can effectively solve the limitations of a single unmanned aerial vehicle, expand the mission execution mode, and improve the system reliability.
[0003] Currently, multi-unmanned aerial vehicle systems generally adopt distributed control. The unmanned aerial vehicles obtain data from the outside world and share it with other unmanned aerial vehicles through a data link, and on this basis, they cooperate to achieve complex behaviors. However, this method has significant drawbacks. On the one hand, the information sources of unmanned aerial vehicles mainly include the Global Navigation Satellite System (GNSS), Real-Time Kinematic (RTK) system, or motion capture system. GNSS is suitable for outdoor environments, but in an environment with dense obstacles, the accuracy is low, the error is large, and any signal loss will have a fatal impact on the control of this high-dynamic system. The additional deployed RTK survey instrument or motion capture system not only increases the cost but also is not suitable for large-scale deployment. On the other hand, multi-unmanned aerial vehicle systems rely on communication networking to obtain information, and the cooperation of unmanned aerial vehicles requires precise and frequent information exchange between individuals. The amount of data transmitted will increase sharply with the increase in the formation scale, and there are limitations in communication distance and bandwidth. At the same time, there are also reliability problems with the communication link. In a complex environment, there are not only problems such as data loss and communication delay, but also it is vulnerable to communication interference and network attacks, and even communication interruption and hijacking may occur. The limitations in communication also greatly increase the complexity of multi-aircraft cooperation. In addition, in most cases, multi-unmanned aerial vehicle systems lack autonomy and can only fly in an obstacle-free and known environment. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a distributed unmanned aerial vehicle cooperative motion control method and device based on pure vision that can control the individual unmanned aerial vehicles to achieve cluster motion and obstacle avoidance without relying on wireless communication in a complex and unknown environment.
[0005] A distributed unmanned aerial vehicle cooperative motion control method based on pure vision, the method includes:
[0006] Construct a distributed UAV swarm motion model based on pure vision. In the distributed UAV swarm motion model, the UAVs are equipped with on-board sensors and on-board computers. Among them, the on-board sensors sense the external environment of the UAV flight and obtain on-board sensor information, which includes grayscale images, depth images, and UAV motion information. The on-board computer calculates the on-board sensor information according to the perceptual motion neural network therein to obtain control instructions, and controls the UAV motion according to the control instructions. The perceptual motion neural network includes an expert system and a student system.
[0007] Obtain the prior information of UAV flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, collision avoidance rule, and migration rule during the flight of the UAV swarm to obtain the first control instruction output by the expert system.
[0008] The student system includes an imitation learning network and a multi-layer perceptron. The grayscale image, depth image, and UAV motion information in the on-board sensor information are respectively obtained and processed according to the three branches of the imitation learning network to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector. The multi-layer perceptron connects and processes the grayscale feature vector, depth feature vector, and motion feature vector to obtain the second control instruction output by the student system.
[0009] Train the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and control the UAV motion according to the final control instruction output by the trained student system.
[0010] In one embodiment, the UAV motion information obtained by the inertial measurement unit includes the UAV's own speed, acceleration, attitude information, and reference flight direction. Among them, the reference flight direction refers to the flight direction from the current position of the UAV to the target position without considering conflicts and collisions.
[0011] In one embodiment, the prior information of UAV flight includes: the accurate state information of the UAV itself, the accurate state information of neighboring UAVs, target information, and obstacle information. Among them, the accurate state information of the UAV itself includes the UAV's own position information, UAV's own speed information, and UAV's own acceleration information, and the accurate state information of neighboring UAVs includes the neighboring UAV's position information and neighboring UAV's speed information.
[0012] In one embodiment, obtaining the prior information of UAV flight according to the expert system and performing calculations in combination with the separation rule, aggregation rule, alignment rule, collision avoidance rule, and migration rule during the flight of the UAV swarm to obtain the first control instruction output by the expert system includes:
[0013] Obtain the prior information of UAV flight according to the expert system, and calculate the prior information of UAV flight by combining the separation rule, aggregation rule, alignment rule, collision avoidance rule and migration rule, and respectively obtain the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term of UAV flight;
[0014] By summing the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term, the final speed of UAV flight is obtained, and the final speed is used as the first control instruction output by the expert system.
[0015] In one embodiment, by summing the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term, the final speed of UAV flight is obtained, including:
[0016] In each time step, sum the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term to obtain the desired speed of UAV flight, expressed as
[0017]
[0018] Wherein, represents the separation speed term of UAV i, represents the aggregation speed term of UAV i, represents the alignment speed term of UAV i, represents the collision avoidance speed term of UAV i approaching obstacle s, represents the migration speed term of UAV i;
[0019] According to the preset speed upper limit v max constrain the desired speed to obtain the final speed of UAV flight, expressed as
[0020]
[0021] In one embodiment, before constraining the desired speed according to the preset speed upper limit v max to obtain the final speed of UAV flight, it further includes:
[0022] Constrain the desired speed according to the preset acceleration upper limit, and the change of the desired speed does not exceed the preset acceleration upper limit.
[0023] In one embodiment, according to the three branches of the imitation learning network, respectively obtain and process the grayscale image, depth image and UAV motion information in the on-board sensor information to obtain the grayscale feature vector, depth feature vector and motion feature vector, including:
[0024] The first neural network branch of the imitation learning network includes an object detection layer, a two-dimensional convolutional neural network, and a one-dimensional time-domain convolutional neural network; the object detection layer performs object recognition on the input grayscale image to obtain a five-dimensional feature vector of the recognized object; the two-dimensional convolutional neural network processes the extended five-dimensional feature vector to obtain historical data of the feature vector; the one-dimensional time-domain convolutional neural network processes the historical data of the feature vector to obtain a grayscale feature vector;
[0025] The second neural network branch of the imitation learning network includes a depth image feature extraction network, a one-dimensional convolutional neural network, and a one-dimensional time-domain convolutional neural network; the depth image feature extraction network extracts features from the input depth image and outputs depth image features; the one-dimensional convolutional neural network processes the depth image features and outputs historical data of the depth image features; the one-dimensional time-domain convolutional neural network processes the historical data of the depth image features to obtain a depth feature vector;
[0026] The third neural network branch of the imitation learning network includes a state sampling module and a five-layer perceptron network; the state sampling module samples and connects the drone's own speed, acceleration, attitude information, and reference flight direction in the input drone motion information to obtain the connected sampling information; the five-layer perceptron network processes the connected sampling information to obtain a motion feature vector.
[0027] In one embodiment, the student system is trained through an off-policy and data aggregation strategy until a trained student system is obtained, including:
[0028] According to the off-policy, minimize the action difference between the first control instruction and the second control instruction during the training process, and update the action difference according to the onboard sensor information collected by the drone and the control instruction currently executed by the drone during each training obtained by the data set strategy until a trained student system is obtained.
[0029] In one embodiment, minimizing the action difference between the first control instruction and the second control instruction according to the off-policy includes:
[0030] During the training process, when the action difference between the first control instruction and the second control instruction is less than the preset action difference threshold and the drone will not collide when executing the second control instruction for movement, control the drone movement according to the second control instruction; otherwise, control the drone movement according to the first control instruction; wherein, after each training, update the preset action difference threshold ξ to ζ′ = min(ζ + 0.5, 10);
[0031] After the training is completed, control the drone movement according to the final control instruction output by the trained student system.
[0032] A distributed UAV cooperative motion control device based on pure vision, the device comprising:
[0033] A distributed UAV cluster motion construction module for constructing a distributed UAV cluster motion model based on pure vision. In the distributed UAV cluster motion model, the UAVs are equipped with on-board sensors and on-board computers. Among them, the on-board sensors sense the external environment of the UAV flight and obtain on-board sensor information, and the on-board sensor information includes grayscale images, depth images, and UAV motion information. The on-board computer calculates the on-board sensor information according to the perception motion neural network therein to obtain a control instruction, and controls the UAV motion according to the control instruction. The perception motion neural network includes an expert system and a student system;
[0034] A first control instruction output module for obtaining the prior information of UAV flight according to the expert system, and calculating by combining the separation rule, aggregation rule, alignment rule, anti-collision rule, and migration rule during the UAV cluster flight to obtain the first control instruction output by the expert system;
[0035] A second control instruction output module for respectively obtaining and processing the grayscale image, depth image, and UAV motion information in the on-board sensor information according to the three branches of the imitation learning network in the student system to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector; connecting and processing the grayscale feature vector, depth feature vector, and motion feature vector according to the multi-layer perceptron in the learning system to obtain the second control instruction output by the student system;
[0036] A training module for training the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and controlling the UAV motion according to the final control instruction output by the trained student system.
[0037] The above-mentioned distributed UAV cooperative motion control method and device based on pure vision constructs a distributed UAV cluster motion model based on pure vision. In the model, the UAVs sense the external environment of the UAV flight through on-board sensors and obtain on-board sensor information, and calculate the on-board sensor information through the perception motion neural network in the on-board computer to obtain a control instruction, and control the UAV motion according to the control instruction. By using this method, it can be ensured that the UAVs do not rely on communication, only rely on the on-board sensor vision to perceive environmental obstacles and neighboring UAVs purely, and directly map the sensor perception data into high-level control signals according to the perception motion neural network in the on-board computer, so as to realize the UAVs to perform cluster, obstacle avoidance, navigation and other motions in a complex environment. Description of the Drawings
[0038] Figure 1 Schematic flowchart of a purely vision-based distributed UAV cooperative motion control method in an embodiment;
[0039] Figure 2 Schematic flowchart of a purely vision-based distributed UAV swarm motion model in an embodiment;
[0040] Figure 3 Schematic diagram of the network architecture of a student system in an embodiment. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] In one embodiment, as Figure 1 shown, a purely vision-based distributed UAV cooperative motion control method is provided, including the following steps:
[0043] Step S1, constructing a purely vision-based distributed UAV swarm motion model. In the distributed UAV swarm motion model, the UAVs are equipped with on-board sensors and on-board computers; wherein, the on-board sensors sense the external environment of the UAV flight and obtain on-board sensor information, and the on-board sensor information includes grayscale images, depth images and UAV motion information; the on-board computer calculates the on-board sensor information according to the perception motion neural network therein to obtain control instructions, and controls the UAV motion according to the control instructions; the perception motion neural network includes an expert system and a student system.
[0044] Among them, the constructed purely vision-based distributed UAV swarm motion model is as Figure 2 shown. The demonstration, data collection and verification of the entire model are all carried out in a simulation environment built by Gazebo, and the UAV model is constructed based on ros (Robot Operating System). Combining Figure 2It can be known that the grayscale images in the airborne sensor information come from the grayscale monocular cameras on the left, right, and rear, and the front binocular depth camera carried by the UAV. The depth images come from the front binocular depth camera carried by the UAV, as well as the UAV motion information obtained by the Inertial Measurement Unit (IMU). The UAV motion information specifically includes the UAV's own speed, acceleration, attitude information, and the reference flight direction. Among them, the reference flight direction refers to the flight direction from the current position of the UAV to the target position without considering conflicts and collisions. The prior information of UAV flight includes: the accurate state information of the UAV itself, the accurate state information of neighboring UAVs, target information, and obstacle information; among them, the accurate state information of the UAV itself includes the position information, speed information, and acceleration information of the UAV itself, and the accurate state information of neighboring UAVs includes the position information and speed information of neighboring UAVs.
[0045] It can be understood that the visual perception obtained by the airborne sensors can not only provide unparalleled information density, but also is not dependent on communication and has real-time performance. Compared with other UAV sensing devices, the airborne sensor cameras have obvious advantages in terms of weight, cost, size, power consumption, and field of view. For UAVs, the visual perception information is rich enough and does not require networking, and there are no problems of network latency and network interference.
[0046] Step S2: Obtain the prior information of UAV flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, anti-collision rule, and migration rule during the collective flight of UAVs to obtain the first control instruction output by the expert system.
[0047] It can be understood that the expert system includes a motion model based on swarm intelligence (Reynolds-Boids). The prior information of UAV flight obtained through this model provides high-quality decision-making behavior data (i.e., the first control instruction) for the student system.
[0048] Step S3: The student system includes an imitation learning network and a multi-layer perceptron. Obtain and process the grayscale image, depth image, and UAV motion information in the airborne sensor information according to the three branches of the imitation learning network to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector; connect and process the grayscale feature vector, depth feature vector, and motion feature vector according to the multi-layer perceptron to obtain the second control instruction output by the student system.
[0049] It can be understood that the student system is an end-to-end perception and motion controller that cannot obtain any prior information and is only equipped with the recorded sensor information from the airborne sensors, and generates the second control instruction through the imitation learning network and the multi-layer perceptron.
[0050] Step S4: Train the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and control the movement of the drone according to the final control instruction output by the trained student system.
[0051] It can be understood that the student system adopts an imitation learning mechanism and realizes the mapping training from visual input to control instructions by imitating the demonstrations provided by the expert system. Finally, a trained end-to-end perception-motion controller is obtained. During the training process, the onboard sensor information collected by the drone and the control instructions currently executed by the drone flight are mainly collected according to the data aggregation strategy (Dagger). According to the off-policy, the action difference between the first control instruction and the second control instruction during the training process is minimized until a trained student system is obtained, and the movement of the drone is controlled according to the final control instruction output by the trained student system.
[0052] It can be understood that in the face of uncertain, diverse, and dynamic environments and tasks, compared with traditional controllers that decouple the control of drones into multiple subtasks, the end-to-end perception-motion controller directly predicts control commands from sensor data, reducing the delay between perception and action, and at the same time being robust to perception artifacts (such as motion blur, data loss, and sensor noise). In addition, the imitation learning mechanism has dynamic adaptability, can well solve the generalization problem, thus having the so-called intelligence, and due to having relatively high-quality decision-making behavior data, imitation learning can reduce the sample complexity.
[0053] In the above-mentioned pure vision-based distributed drone cooperative motion control method, a pure vision-based distributed drone cluster motion model is constructed. In the model, the drone perceives the external environment of the drone flight through the onboard sensor and obtains the onboard sensor information, calculates the onboard sensor information through the perception-motion neural network in the onboard computer to obtain a control instruction, and controls the movement of the drone according to the control instruction; and respectively outputs the first control instruction and the second control instruction according to the expert system and the student system in the perception-motion neural network; finally, the student system is trained through an off-policy and data aggregation strategy until a trained student system is obtained, and the movement of the drone is controlled according to the final control instruction output by the trained student system. Using this method can ensure that drones do not rely on communication, only purely rely on the onboard sensor vision to perceive environmental obstacles and neighboring drones, and directly map the sensor perception data into high-level control signals according to the perception-motion neural network in the onboard computer, realizing the movement of drones such as clustering, obstacle avoidance, and navigation in a complex environment.
[0054] In one of the embodiments, the prior information of the UAV flight is obtained according to the expert system, and calculations are performed in combination with the separation rules, aggregation rules, alignment rules, anti-collision rules, and migration rules during the UAV swarm flight to obtain the first control instruction output by the expert system, including:
[0055] First, the kinematic equations of the UAVs in the UAV swarm are described according to the swarm intelligence-based motion model in the expert system. Among them, the UAV swarm includes N quadrotor UAVs, each quadrotor UAV has the same motion characteristics, and each quadrotor UAV has four propellers and one controller. The controller can issue control commands to each propeller respectively. To simply simulate the quadrotor UAV, it is assumed that the flight speed of the quadrotor UAV is slow enough to ignore the external aerodynamic forces acting on the quadrotor, such as air resistance and blade vortices; secondly, it is assumed that the response speed of the propeller to the thrust command is fast enough to ignore the time delay from the controller issuing the thrust command to the propeller actually generating thrust. Therefore, under the above assumptions of ignoring air resistance and motor dynamics, the kinematic equation of the quadrotor UAV can be expressed as
[0056]
[0057]
[0058]
[0059]
[0060] where p WB , v WB , q WB respectively represent the position, linear velocity, and attitude of the UAV in the world coordinate system, respectively represent the first-order derivatives of p WB , v WB , q WB with respect to time, g w represents the gravitational acceleration in the world coordinate system, q WB ⊙c B represents the transformation of the mass-normalized thrust vector c B =(0, 0, c) T under q WB , c represents the thrust magnitude, and the quaternion form of q WB is expressed as q WB =(q w , q x , q y , q z ) T , Λ(ω B ) represents the vector The skew-symmetric matrix, J = diag(J xx , J yy , J zz ) represents the moment of inertia of the UAV, represents the torque exerted by the motor thrust on the UAV, represents the set of 3D real vectors.
[0061] Then, the prior information of UAV flight is obtained according to the expert system, and the prior information of UAV flight is calculated by combining the separation rule, aggregation rule, alignment rule, collision avoidance rule and migration rule, and the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term of UAV flight are obtained respectively.
[0062] Specifically, the separation rule means that in the movement of the UAV swarm, it is necessary to avoid the UAVs in the swarm being too close, ensure that the UAVs maintain an appropriate distance, and prevent collisions between UAVs. For the separation speed term its magnitude is related to the distance r ij between UAVs and the approach speed . The repulsion range between UAVs is related to the approach speed between UAVs. The repulsion range is defined as When the distance between UAVs is less than this value, a local repulsive force is generated between UAVs, generating a separation speed term. The closer the UAVs are, the stronger the repulsive force, and the greater the approach speed of the UAVs, the stronger the repulsive force. The repulsion range and approach speed are specifically expressed as
[0063]
[0064]
[0065] where, is the minimum allowable separation distance between UAVs. When the distance between UAVs is less than this distance, repulsion will definitely occur between UAVs. r ij = |p i - p j | is the spatial distance between UAV i and neighboring UAV j, T is the prediction interval, generally set to 2s; according to the above repulsion range and approach speed, the separation speed between UAV i and neighboring UAV j is calculated as
[0066]
[0067] Furthermore, the total separation speed term generated by UAV i is calculated as
[0068]
[0069] Specifically, the aggregation rule refers to the rule that ensures the aggregation of the UAV swarm and prevents it from dispersing during the swarm movement. For the aggregation velocity term, its magnitude is related to the distance r between UAVs ij and the velocity away is related. The aggregation distance is defined as When the distance between UAVs is greater than this value, the local attraction between UAVs comes into play, generating the aggregation velocity term. The greater the distance between UAVs, the greater the attraction, and the greater the velocity away between UAVs, the stronger the attraction. The aggregation distance and the velocity away are specifically expressed as
[0070]
[0071]
[0072] where is the aggregation threshold between UAVs. When the distance between UAVs is greater than this distance, there will definitely be an attraction between UAVs; according to the above-mentioned aggregation distance and velocity away, the aggregation velocity between UAV i and neighboring UAV j is calculated as
[0073]
[0074] Furthermore, the total aggregation velocity term generated by UAV i is calculated as
[0075]
[0076] Specifically, the alignment rule refers to the rule in swarm movement that makes UAVs try to be consistent with the average direction of neighboring individuals, enables the swarm to move in the same direction, and ensures the orderliness of the UAV swarm. The alignment velocity between UAV i and neighboring UAV j calculated according to the alignment rule is Furthermore, the total alignment velocity term generated by UAV i is calculated as where C frict represents the alignment parameter and is a constant.
[0077] Specifically, for the collision avoidance rule, the obstacle is decomposed into multiple points, and the UAV generates a repulsive force with the obstacle within the radius range r. The closer the UAV is to the obstacle, the greater the approaching velocity of the UAV to the obstacle, and the greater the collision avoidance velocity term. Assuming the position of the obstacle s is The repulsive range between UAV i and obstacle s is defined as where r is is the spatial distance between the UAV and the obstacle at the current moment, is the velocity of the UAV approaching the obstacle. The collision avoidance velocity between UAV i and obstacle s calculated according to the collision avoidance rule is expressed as
[0078]
[0079]
[0080]
[0081] Similarly, the migration speed term is calculated according to the migration rule. Guide the UAV to move towards the target, and the direction of the migration speed term is the direction of the target point, and the magnitude is a constant term.
[0082] Finally, within each time step, sum up the above-mentioned separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term, and migration speed term to obtain the expected speed of the UAV flight, denoted as
[0083]
[0084] After obtaining the expected speed by superposition, in order to prevent the UAV speed from being too large, according to the preset speed upper limit v max constrain the expected speed to obtain the final speed of the UAV flight, denoted as
[0085]
[0086] At the same time, considering the maneuverability of the UAV, constrain the expected speed according to the preset acceleration upper limit, and the change of the expected speed does not exceed the preset acceleration upper limit.
[0087] In one embodiment, the network architecture of the student system is as Figure 3 , including an imitation learning network with three branches and a multi-layer perceptron.
[0088] Among them, the first neural network branch of the imitation learning network includes an object detection layer, a two-dimensional convolutional neural network, and a one-dimensional time-domain convolutional neural network. First, according to the object detection layer, the input grayscale image Perform target recognition. The target detection layer adopts the yolo3v3-tiny architecture pre-trained on an automatically labeled image dataset. The network architecture consists of a total of 13 convolutional layers, interspersed with max pooling layers and leaky rectified linear units (ReLUs). The target detection layer outputs a five-dimensional feature vector [direction, x, y, size_x, size_y] of the recognized target. [direction] is the image direction marker, that is, the direction of the camera from which the target is detected (front, left, right, back). [x, y] are the coordinates of the recognized target on the image, and [size_x, size_y] are the length and width of the recognition box. Then, expand the five-dimensional feature vector of the recognized target to the dimension of a tensor and input it into a two-dimensional convolutional neural network for processing. The two-dimensional convolutional neural network contains 4 hidden layers, with filters of (32, 64, 128, 128) respectively, interspersed with LeakyRelu activation connection layers. Finally, after processing by the global average pooling layer (globalAveragePoling2D), historical data with a time length of T = 5 of the feature vector is output. The historical data is sufficient to infer the motion information of the nearby drones. Finally, input the historical data into a 1D time-domain convolutional network for processing. This network contains 4 hidden layers, with filters of (128, 64, 64, 64) respectively. Finally, map the signal to 128 dimensions through a fully connected layer to obtain a grayscale feature vector.
[0089] The second neural network branch of the imitation learning network includes a depth image feature extraction network, a one-dimensional convolutional neural network, and a one-dimensional time-domain convolutional neural network; first, perform feature extraction on the input depth image according to the depth image feature extraction network. The depth image feature extraction network uses the pre-trained MobileNet structure to extract depth image features from the depth image. Then, process the depth image features according to the one-dimensional convolutional neural network. This network includes 4 hidden layers, with filters of (128, 64, 64, 64) respectively, interspersed with LeakyRelu activation connection layers, and outputs historical data with a time length of T = 5 of the depth image features. Finally, process the historical data of the depth image features according to the one-dimensional time-domain convolutional neural network, and output a 120-dimensional depth feature vector.
[0090] The third neural network branch of the imitation learning network includes a state sampling module and a five-layer perceptron network; first, according to the state sampling module, sample the drone's own speed acceleration attitude information represented by a rotation matrix and the reference flight direction Sampling and connection are performed to obtain the connected sampling information. Then, the connected sampling information is processed according to a five-layer perceptron network, which is a filtering layer of [128, 64, 64, 64, 32], with LeakyReLU activation layers interspersed in the middle. Finally, the signal is mapped through a fully connected layer to obtain a 128-dimensional motion feature vector.
[0091] After obtaining the grayscale feature vector, depth feature vector, and motion feature vector output by the three branches of the imitation learning network, the outputs of each branch are connected and processed by a multi-layer perceptron, which contains 4 hidden layers, namely (128, 64, 64, 64) filters, and then mapped to a 3-dimensional feature vector [v x ,v y ,v z through a fully connected layer to obtain the second control instruction output by the student system. Similar to the expert system, in order to prevent sudden changes in the speed instruction generated by the neural network, which may cause strong pitching motion of the drone, for the speed control instruction generated by the student system, a speed upper limit v max is set. At the same time, an acceleration upper limit a max is set, and the maximum change in speed cannot exceed a max .
[0092] In one embodiment, the student system is trained through an off-policy and data aggregation strategy until a trained student system is obtained, including:
[0093] First, off-policy training is carried out. During the k-th flight of the drone, at each moment t, the expert system generates a first control instruction k based on the prior information s of the drone flight. The student system generates a second control instruction k on the policy π based on the on-board sensor information o of the perceived real world. Supervised learning is used to train the neural network until the optimal policy is found, and finding a student system with the same performance as the expert system is reduced to minimizing the action difference between the two policies during the motion process, expressed as
[0094]
[0095] Among them, represents the difference, represents the student policy, ρ(π) represents the trajectory of the drone under the control of the student system, represents the student policy and the control instruction generated under the on-board sensor information o k , that is, the second control instruction
[0096] Then, according to the data set strategy, in the (k + 1)-th iteration, the control instructions obtained from the k-th learning training are used to control the flight of the drone, and the on-board sensor information o collected during this flight is collected. k+1 (t) and the corresponding control instructions And a data set is constructed When the drone completes a round-trip flight, the data sets collected by all drones are added to the data pool, and then all the data sets are used to train a new This process is continuously repeated until the training is completed.
[0097] Furthermore, in order to prevent collisions during training, during the training process, when the action difference between the first control instruction and the second control instruction is less than the preset action difference threshold, and the drone will not collide when executing the second control instruction for movement, the drone is controlled to move according to the second control instruction; otherwise, the drone is controlled to move according to the first control instruction; where, after each training, the preset action difference threshold ξ is updated to ξ′ = min(ξ + 0.5, 10); after the training is completed, the drone is completely controlled to move according to the final control instruction output by the trained student system.
[0098] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0099] In one embodiment, a distributed drone cooperative motion control device based on pure vision is provided, including: a distributed drone cluster motion construction module, a first control instruction output module, a second control instruction output module, and a training module, where:
[0100] Distributed UAV Cluster Motion Construction Module, which is used to construct a distributed UAV cluster motion model based on pure vision. In the distributed UAV cluster motion model, the UAVs are equipped with on-board sensors and on-board computers. Among them, the on-board sensors sense the external environment of the UAV flight and obtain on-board sensor information, and the on-board sensor information includes grayscale images, depth images, and UAV motion information. The on-board computer calculates the on-board sensor information according to the perception motion neural network therein to obtain control instructions, and controls the UAV motion according to the control instructions. The perception motion neural network includes an expert system and a student system.
[0101] The First Control Instruction Output Module is used to obtain the prior information of UAV flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, collision avoidance rule, and migration rule during the flight of the UAV cluster to obtain the first control instruction output by the expert system.
[0102] The Second Control Instruction Output Module is used to respectively obtain and process the grayscale image, depth image, and UAV motion information in the on-board sensor information according to the three branches of the imitation learning network in the student system to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector. The multi-layer perceptron in the learning system is used to connect and process the grayscale feature vector, depth feature vector, and motion feature vector to obtain the second control instruction output by the student system.
[0103] The Training Module is used to train the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and control the UAV motion according to the final control instruction output by the trained student system.
[0104] For the specific limitations of the distributed UAV cooperative motion control device based on pure vision, reference can be made to the limitations of the distributed UAV cooperative motion control method in the above text, which will not be elaborated here. Each module in the above-mentioned distributed UAV cooperative motion control device based on pure vision can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered as the scope recorded in this specification.
[0106] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A distributed UAV cooperative motion control method based on pure vision, characterized in that, the method includes: Construct a distributed UAV cluster motion model based on pure vision. In the distributed UAV cluster motion model, the UAV is equipped with an on-board sensor and an on-board computer; wherein, the on-board sensor perceives the external environment of the UAV flight and obtains on-board sensor information, and the on-board sensor information includes grayscale images, depth images, and UAV motion information; the on-board computer calculates the on-board sensor information according to the perception motion neural network therein to obtain a control instruction, and controls the UAV motion according to the control instruction; the perception motion neural network includes an expert system and a student system; Obtain the prior information of UAV flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, anti-collision rule, and migration rule during the flight of the UAV cluster to obtain the first control instruction output by the expert system; The student system includes an imitation learning network and a multi-layer perceptron. The grayscale image, depth image, and UAV motion information in the on-board sensor information are respectively obtained and processed according to the three branches of the imitation learning network to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector; the multi-layer perceptron connects and processes the grayscale feature vector, depth feature vector, and motion feature vector to obtain the second control instruction output by the student system; Train the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and control the UAV motion according to the final control instruction output by the trained student system.
2. The method according to claim 1, characterized in that, The UAV motion information obtained by the inertial measurement unit includes the UAV's own speed, acceleration, attitude information, and reference flight direction, where the reference flight direction refers to the flight direction from the UAV's current position to the target position without considering conflicts and collisions.
3. The method according to claim 1, characterized in that, The prior information of UAV flight includes: the accurate state information of the UAV itself, the accurate state information of neighboring UAVs, target information, and obstacle information; wherein, the accurate state information of the UAV itself includes the UAV's own position information, UAV's own speed information, and UAV's own acceleration information, and the accurate state information of neighboring UAVs includes the position information of neighboring UAVs and the speed information of neighboring UAVs.
4. The method according to claim 1, characterized in that, Obtain the prior information of UAV flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, anti-collision rule, and migration rule during the flight of the UAV cluster to obtain the first control instruction output by the expert system, including: Obtain the prior information of UAV flight according to the expert system, and calculate the prior information of UAV flight by combining the separation rule, aggregation rule, alignment rule, anti-collision rule and migration rule, respectively obtaining the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term of UAV flight; By summing up the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term, obtain the final speed of UAV flight, and use the final speed as the first control instruction output by the expert system.
5. The method according to claim 4, wherein, obtaining the final speed of UAV flight by summing up the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term, includes: Within each time step, sum up the separation speed term, aggregation speed term, alignment speed term, collision avoidance speed term and migration speed term to obtain the desired speed of UAV flight, expressed as Among them, represents the separation speed term of UAV i, represents the aggregation speed term of UAV i, represents the alignment speed term of UAV i, represents the collision avoidance speed term when UAV i approaches obstacle s, represents the migration speed term of UAV i; According to the preset speed limit v max constrain the desired speed to obtain the final speed of the UAV flight, expressed as 6. The method according to claim 5, wherein, According to a preset speed upper limit v max Before constraining the desired speed to obtain the final speed of the UAV flight, it further includes: constrain the desired speed according to the preset acceleration upper limit, and the change of the desired speed does not exceed the preset acceleration upper limit.
7. The method according to claim 1, wherein, respectively obtain and process the grayscale image, depth image and UAV motion information in the on-board sensor information according to the three branches of the imitation learning network to obtain a grayscale feature vector, a depth feature vector and a motion feature vector, including: The first neural network branch of the imitation learning network includes an object detection layer, a two-dimensional convolutional neural network and a one-dimensional time-domain convolutional neural network; identify the object in the input grayscale image according to the object detection layer to obtain a five-dimensional feature vector of the identified object; process the extended five-dimensional feature vector according to the two-dimensional convolutional neural network to obtain the historical data of the feature vector; process the historical data of the feature vector according to the one-dimensional time-domain convolutional neural network to obtain the grayscale feature vector; The second neural network branch of the imitation learning network includes a depth image feature extraction network, a one-dimensional convolutional neural network and a one-dimensional time-domain convolutional neural network; extract features from the input depth image according to the depth image feature extraction network and output the depth image features; process the depth image features according to the one-dimensional convolutional neural network and output the historical data of the depth image features; process the historical data of the depth image features according to the one-dimensional time-domain convolutional neural network to obtain the depth feature vector; The third neural network branch of the imitation learning network includes a state sampling module and a five-layer perceptron network; sample and connect the UAV's own speed, acceleration, attitude information and reference flight direction in the input UAV motion information according to the state sampling module to obtain the connected sampling information; process the connected sampling information according to the five-layer perceptron network to obtain the motion feature vector.
8. The method according to claim 1, wherein, Training the student system through an off-policy and data aggregation strategy until a trained student system is obtained, including: Minimizing the action difference between the first control instruction and the second control instruction during the training process according to the off-policy, and updating the action difference based on the onboard sensor information collected by the drone and the control instruction currently executed by the drone during each training obtained according to the data aggregation strategy until a trained student system is obtained.
9. The method according to claim 8, wherein, Minimizing the action difference between the first control instruction and the second control instruction during the training process according to the off-policy includes: During the training process, when the action difference between the first control instruction and the second control instruction is less than the preset action difference threshold, and the UAV will not collide when executing the second control instruction for movement, control the movement of the UAV according to the second control instruction; otherwise, control the movement of the UAV according to the first control instruction; where, every time after a training, update the preset action difference threshold ξ to ξ ′ = min(ξ + 0.5, 10); After the training is completed, controlling the movement of the drone according to the final control instruction output by the trained student system.
10. A distributed drone cooperative motion control device based on pure vision, wherein, The device includes: A distributed drone cluster motion construction module, configured to construct a distributed drone cluster motion model based on pure vision. In the distributed drone cluster motion model, the drone is equipped with an onboard sensor and an onboard computer; wherein, the onboard sensor perceives the external environment of the drone flight and obtains onboard sensor information, and the onboard sensor information includes grayscale images, depth images, and drone motion information; the onboard computer calculates the onboard sensor information according to the perception motion neural network therein to obtain a control instruction, and controls the movement of the drone according to the control instruction; the perception motion neural network includes an expert system and a student system; A first control instruction output module, configured to obtain the prior information of the drone flight according to the expert system, and perform calculations in combination with the separation rule, aggregation rule, alignment rule, anti-collision rule, and migration rule during the flight of the drone cluster to obtain the first control instruction output by the expert system; A second control instruction output module, configured to respectively obtain and process the grayscale image, depth image, and drone motion information in the onboard sensor information according to the three branches of the imitation learning network in the student system to obtain a grayscale feature vector, a depth feature vector, and a motion feature vector; connecting and processing the grayscale feature vector, depth feature vector, and motion feature vector according to the multi-layer perceptron in the student system to obtain the second control instruction output by the student system; A training module, configured to train the student system through an off-policy and data aggregation strategy until a trained student system is obtained, and control the movement of the drone according to the final control instruction output by the trained student system.
Citation Information
Patent Citations
Self-learning autonomous navigation systems and methods for unmanned underwater vehicle
CA3067575A1
Unmanned aerial vehicle flight control method based on reinforcement learning and network model distillation
CN113110550A