Obstacle Avoidance Method, Device, Equipment and Medium Based on Imitation Learning

Through the obstacle-walking method based on imitation learning, the obstacle-walking model trained by multimodal loss function is used to output the multimodal velocity direction index, locate the target expert trajectory fragment and control the drone navigation, the robustness and flexibility of the existing drone control algorithm in complex environments is solved, and the efficient trajectory obstacle-avoidance effect is achieved.

CN119781503BActive Publication Date: 2025-05-27TIANJIN YUNSHENG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510279373.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Existing UAV control algorithms show low robustness and flexibility in complex dynamic environments. Deep learning methods have problems such as instability in training and insufficient generalization capabilities, resulting in poor trajectory obstacle avoidance.

Method used

By using the imitation learning-based obstacle course, the obstacle course model is trained by collecting the navigation data of the drone, using the multimodal loss function, the velocity direction index and probability prediction value of the multimodal, the target expert trajectory fragment is located and the drone navigation is controlled.

Benefits of technology

It significantly improves the robustness, flexibility and obstacle avoidance effect of trajectory planning, and realizes the decision-making process that imitates humans or algorithm experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781503B_ABST
    Figure CN119781503B_ABST
Patent Text Reader

Abstract

The present invention provides an obstacle avoidance method, device, equipment and medium based on imitation learning, including: collecting first navigation data of an unmanned aerial vehicle (UAV); through an obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data and target direction data in the first navigation data, outputting a multi-mode first speed direction index, and the multi-mode first speed direction index is used to locate a target expert trajectory segment to be executed in an expert trajectory library, and the target expert trajectory segment includes a plurality of navigation points and their corresponding expert speed directions; controlling the UAV to navigate according to the navigation points and their corresponding expert speed directions in the target expert trajectory segment to be executed, and continuously collecting new first navigation data during the UAV navigation until the UAV reaches the end position. The present invention can significantly improve the robustness, flexibility of trajectory planning and the effect of trajectory obstacle avoidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of trajectory planning, and in particular to an obstacle avoidance method, device, equipment and medium based on imitation learning. Background Art

[0002] With the rapid development of unmanned aerial vehicle (UAV) technology, the application of UAVs in fields such as logistics transportation, environmental monitoring, and disaster rescue has become increasingly popular. However, achieving precise control of UAVs in complex dynamic environments remains an urgent problem to be solved. Currently, traditional control algorithms usually rely on accurate dynamic modeling and environmental perception, but they often exhibit low robustness and flexibility when faced with variable obstacles and complex kinematic constraints. In addition, the development of deep learning has provided new solutions for UAV control, but the end-to-end method based entirely on data-driven has problems such as unstable training and insufficient generalization ability, resulting in poor obstacle avoidance effects of the planned trajectories. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an obstacle avoidance method, device, equipment and medium based on imitation learning, which can significantly improve the robustness, flexibility of trajectory planning and the obstacle avoidance effect of the trajectory.

[0004] In a first aspect, the present invention provides an obstacle avoidance method based on imitation learning, including:

[0005] Collect the first navigation data of the UAV, where the first navigation data includes scene image data, UAV pose data, and target direction data, and the target direction data is a vector pointing from the current position of the UAV to the end position;

[0006] Through an obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data, and target direction data in the first navigation data, output a multi-mode first speed direction index and its corresponding probability prediction value; wherein, the multi-mode first speed direction index is used to locate the target expert trajectory segment to be executed in the expert trajectory library, and the probability prediction value is used to characterize the possibility that the target expert trajectory segment corresponding to the first speed direction index is called. The target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the obstacle avoidance model is trained using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function. The classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-mode second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the UAV, and the speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment;

[0007] Control the UAV to navigate according to the navigation points and their corresponding expert velocity directions within the target expert trajectory segment to be executed, and continue to collect new first navigation data during the UAV navigation until the UAV reaches the end position.

[0008] In one implementation, before outputting the multi-mode first velocity direction index by the obstacle avoidance model based on imitation learning according to the scene image data, UAV pose data, and target direction data, the method further includes:

[0009] Obtain multiple pre-planned expert trajectories;

[0010] Control the UAV to navigate according to the expert trajectories in scenarios of different scales, and collect second navigation data during the UAV navigation. The second navigation data includes time stamps, scene image data, UAV pose data, and target direction data;

[0011] Intercept segments of the expert trajectories based on the time stamps to construct an expert trajectory library;

[0012] Use the expert trajectory library, as well as the scene image data, UAV pose data, and target direction data in the second navigation data, and combine with a multi-modal loss function to train the obstacle avoidance model based on imitation learning.

[0013] In one implementation, the expert trajectory library includes multiple target expert trajectory segments; intercepting segments of the expert trajectories based on the time stamps to construct the expert trajectory library includes:

[0014] Determine multiple groups of start time stamps and end time stamps according to the preset planning time to obtain multiple time stamp intervals to be intercepted;

[0015] Intercept the initial expert trajectory segments within the time stamp intervals from the expert trajectories;

[0016] Reconstruct the trajectory based on the navigation points included in the initial expert trajectory segments to obtain smooth expert trajectory segments;

[0017] Sample multiple navigation points from the smooth expert trajectory segments. The sampled navigation points and their corresponding expert velocity directions constitute the target expert trajectory segments.

[0018] In one implementation, using the expert trajectory library, as well as the scene image data, UAV pose data, and target direction data in the second navigation data, and combining with a multi-modal loss function to train the obstacle avoidance model based on imitation learning includes:

[0019] Output a multi-mode second velocity direction index through the obstacle avoidance model based on imitation learning according to the scene image data, UAV pose data, and target direction data in the second navigation data;

[0020] Extract the first target expert trajectory segment that matches the second navigation data from the expert trajectory library based on the timestamps in the second navigation data;

[0021] Based on the multimodal loss function, determine the classification loss value and the velocity constraint loss value based on the multimodal second velocity direction index and the first target expert trajectory segment, and fuse the classification loss value and the velocity constraint loss value to obtain the multimodal loss value;

[0022] Train the obstacle avoidance model based on imitation learning using the multimodal loss value.

[0023] In one implementation, determining the classification loss value based on the multimodal second velocity direction index and the first target expert trajectory segment according to the multimodal loss function includes:

[0024] Locate the second target expert trajectory segment corresponding to the second velocity direction index of each mode in the expert trajectory library;

[0025] Determine the label corresponding to each second target expert trajectory segment based on the first target expert trajectory segment, where the label is used to represent whether the second target expert trajectory segment is the first target expert trajectory segment;

[0026] Determine the classification loss value according to the probability prediction value corresponding to the second velocity direction index of each mode and the label corresponding to each second target expert trajectory segment.

[0027] In one implementation, determining the velocity constraint loss value based on the multimodal second velocity direction index and the first target expert trajectory segment according to the multimodal loss function includes:

[0028] Interpret the second velocity direction index of each mode to obtain the velocity direction prediction value corresponding to the second velocity direction index of each mode;

[0029] Determine the matching degree between the velocity direction prediction value corresponding to the second velocity direction index of each mode and the expert velocity direction included in the first target expert trajectory segment;

[0030] Determine the velocity constraint loss value based on the matching degree.

[0031] In one implementation, determining the classification loss value and the velocity constraint loss value based on the multimodal second velocity direction index and the first target expert trajectory segment according to the multimodal loss function, and fusing the classification loss value and the velocity constraint loss value to obtain the multimodal loss value includes:

[0032] Determine the multimodal loss value according to the following formula:

[0033] ;

[0034] Among them, is the multimodal loss value, is the quantity of the second navigation data, is the number of modes, is for the th second navigation data output of the th mode of the second speed direction index, is the second speed direction index corresponding to the label of the second target expert trajectory segment, is the second speed direction index corresponding to the probability prediction value of the second target expert trajectory segment, is the second speed direction index corresponding to the speed direction prediction value, is the th expert speed direction included in the first target expert trajectory segment matched by the second navigation data, is the speed direction prediction value and the expert speed direction the degree of matching between them, is the weight coefficient, is a constant.

[0035] In one implementation, through an obstacle avoidance model based on imitation learning, according to the scene image data, the UAV pose data, and the target direction data, a multi-mode first speed direction index is output, including:

[0036] Extract the first feature information of the scene image data, and extract the second feature information of the UAV pose data and the target direction data;

[0037] Perform downsampling processing and fusion processing on the first feature information and the second feature information to obtain target feature information;

[0038] Based on the target feature information, output the multi-mode first speed direction index and its corresponding probability prediction value.

[0039] In a second aspect, the present invention also provides an obstacle avoidance device based on imitation learning, including:

[0040] A data acquisition module, configured to acquire the first navigation data of the UAV, where the first navigation data includes scene image data, UAV pose data, and target direction data, and the target direction data is a vector pointing from the current position of the UAV to the end position;

[0041] An index output module, configured to output a multi-mode first speed direction index and its corresponding probability prediction value according to the scene image data, UAV pose data, and target direction data in the first navigation data through an obstacle avoidance model based on imitation learning; wherein, the multi-mode first speed direction index is used to locate the target expert trajectory segment to be executed in the expert trajectory library, and the probability prediction value is used to characterize the possibility that the target expert trajectory segment corresponding to the first speed direction index is called. The target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the obstacle avoidance model is trained using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function. The classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-mode second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the UAV, and the speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment;

[0042] A UAV control module, configured to control the UAV to navigate according to the navigation points and their corresponding expert speed directions in the target expert trajectory segment to be executed, and continue to collect new first navigation data during the UAV navigation until the UAV reaches the end position.

[0043] In a third aspect, the present invention further provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of the first aspect.

[0044] In a fourth aspect, the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the method according to any one of the first aspect.

[0045] The obstacle avoidance method, device, equipment and medium based on imitation learning provided by the present invention first collect the first navigation data of the unmanned aerial vehicle (UAV). The first navigation data includes scene image data, UAV pose data and target direction data. The target direction data is a vector pointing from the current position of the UAV to the end position. Then, through an obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data and target direction data in the first navigation data, a multi-mode first speed direction index is output. The multi-mode first speed direction index is used to locate the target expert trajectory segment to be executed in the expert trajectory library. The target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions. Finally, the UAV is controlled to navigate according to the navigation points and their corresponding expert speed directions in the target expert trajectory segment to be executed, and new first navigation data is continuously collected during the navigation of the UAV until the UAV reaches the end position. The above method combines expert experience and deep learning technology. Through an obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data and target direction data in the first navigation data, a multi-mode first speed direction index is output, and then the UAV is controlled to navigate according to the target expert trajectory segment located by the first speed direction index. The present invention realizes the decision-making process of imitating human or algorithm experts, not only has high robustness and flexibility, but also has a better obstacle avoidance effect.

[0046] Other features and advantages of the present invention will be described in the following specification, and, in part, will become apparent from the specification or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0047] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a schematic flowchart of an obstacle avoidance method based on imitation learning provided by an embodiment of the present invention;

[0050] Figure 2 It is a schematic structural diagram of an obstacle avoidance model based on imitation learning provided by an embodiment of the present invention;

[0051] Figure 3 Schematic diagram of the architecture of an obstacle avoidance method based on imitation learning provided by an embodiment of the present invention;

[0052] Figure 4 Schematic diagram of the structure of an obstacle avoidance device based on imitation learning provided by an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Currently, when traditional control algorithms face variable obstacles and complex kinematic constraints, they often exhibit low robustness and flexibility. Control algorithms based on deep learning have problems such as unstable training and insufficient generalization ability. Based on this, the embodiments of the present invention provide an obstacle avoidance method, device, equipment, and medium based on imitation learning, which can significantly improve the robustness, flexibility of trajectory planning, and the effect of trajectory obstacle avoidance.

[0056] For ease of understanding of this embodiment, first, a detailed introduction is given to an obstacle avoidance method based on imitation learning disclosed in the embodiments of the present invention. Refer to Figure 1 The flowchart of an obstacle avoidance method based on imitation learning shown in the figure. The method mainly includes the following steps S102 to S106:

[0057] Step S102, collect the first navigation data of the unmanned aerial vehicle.

[0058] Among them, the first navigation data includes scene image data, unmanned aerial vehicle pose data, and target direction data. The unmanned aerial vehicle is equipped with an image acquisition device, such as a monocular fisheye camera. The image collected by the monocular fisheye camera for the scene in front of the unmanned aerial vehicle is the scene image data, and the content shown in the scene image data may include obstacles existing in the forward direction of the unmanned aerial vehicle. The unmanned aerial vehicle is equipped with a sensor system, such as a positioning device, a speed sensor, an accelerometer, a gyroscope, etc. The sensor system is used to collect corresponding pose data during the navigation of the unmanned aerial vehicle. The pose data includes position, quaternion, speed, acceleration, angular velocity, etc. The target direction data is a vector pointing from the current position of the unmanned aerial vehicle to the end position.

[0059] Step S104: Based on the obstacle avoidance model using imitation learning, according to the scene image data, UAV pose data, and target direction data in the first navigation data, output a multi-mode first speed direction index and its corresponding probability prediction value.

[0060] Among them, the input of the obstacle avoidance model using imitation learning includes scene image data, UAV pose data, and target direction data, and the output includes a multi-mode first speed direction index and its corresponding probability prediction value. Multi-mode can be understood as that the number of the output first speed direction indexes is multiple. The first speed direction index is the speed direction index output by the model in the model application stage, and the first speed direction index is used to locate the target expert trajectory segment to be executed in the expert trajectory library. The expert trajectory library includes multiple target expert trajectory segments, and the target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the probability prediction value is used to characterize the possibility that the target expert trajectory segment corresponding to this first speed direction index is called. That is, the higher the probability prediction value corresponding to the first speed direction index, the higher the possibility of controlling the UAV to navigate according to the target expert trajectory segment corresponding to this first speed direction index.

[0061] Among them, the obstacle avoidance model is trained using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function. The classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-mode second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the UAV. The speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment.

[0062] In one example, the scene image data, UAV pose data, and target direction data in the first navigation data are input into the obstacle avoidance model using imitation learning. The model outputs a multi-mode first speed direction index and its corresponding probability prediction value through operations such as feature extraction, feature downsampling, feature fusion, and feature decoding, and locates the target expert trajectory segment to be executed in the expert trajectory library using the first speed direction index with the highest probability prediction value.

[0063] Step S106: Control the UAV to navigate according to the navigation points and their corresponding expert speed directions in the target expert trajectory segment to be executed, and continue to collect new first navigation data during the UAV navigation process until the UAV reaches the end position.

[0064] In one example, during the process of controlling the UAV navigation according to the navigation points and their corresponding expert speed directions within the target expert trajectory segment to be executed, new first navigation data is continuously collected, and a multi-mode new first speed direction index is output based on the new first navigation data through an obstacle avoidance algorithm based on imitation learning, so as to locate a new target expert trajectory segment, and then control the UAV navigation according to the new target expert trajectory segment. The above process is repeated until the UAV reaches the end position.

[0065] The obstacle avoidance method based on imitation learning provided by the embodiments of the present invention combines expert experience and deep learning technology. Through an obstacle avoidance model based on imitation learning, a multi-mode first speed direction index is output according to the scene image data, UAV pose data, and target direction data in the first navigation data, and then the UAV navigation is controlled according to the target expert trajectory segment located by the first speed direction index. The present invention realizes the decision-making process of imitating human or algorithm experts, not only has high robustness and flexibility, but also has a good obstacle avoidance effect.

[0066] To enable the obstacle avoidance model based on imitation learning to have a good obstacle avoidance effect, it is necessary to pre-train the obstacle avoidance model based on imitation learning. The embodiments of the present invention provide a specific implementation manner for training the obstacle avoidance model based on imitation learning. See the following steps 1 to 4:

[0067] Step 1: Obtain multiple pre-planned expert trajectories. Among them, the expert trajectories can be divided into expert trajectories based on the flight of the pilot, expert trajectories based on polynomial fitting, and expert trajectories based on traditional obstacle avoidance algorithms.

[0068] Step 2: Control the UAV navigation according to the expert trajectory in scenarios of different scales, and collect second navigation data during the UAV navigation. The second navigation data includes time stamps, scene image data, UAV pose data, and target direction data.

[0069] In one example, the embodiments of the present invention can control the UAV navigation according to the expert trajectory in large-scale scenarios and small-scale scenarios respectively, such as scenarios of different scales such as woods, buildings, geometric arrays, road signs, and electric towers. Data is collected during the UAV navigation. The data collection is in rounds, and each round of flight lasts for 1 minute. The time stamps, scene image data, UAV pose data (position, quaternion, speed, acceleration, angular velocity), target direction data, etc. during the flight process are recorded, and a total of 13 km of data is collected.

[0070] Step 3: Segment the expert trajectory based on the time stamp to construct an expert trajectory library. Specifically, it includes the following (a) to (d):

[0071] (a)Determine multiple groups of start timestamps and end timestamps according to the preset planning time to obtain multiple timestamp intervals to be intercepted. Among them, the planning time is used to limit the duration of the intercepted expert trajectory segment. For example, the planning time is set to 4s. In one example, the timestamp corresponding to a frame of the scene image data can be used as the start timestamp, and the timestamp 4s later can be used as the end timestamp, so as to obtain the timestamp interval; in the subsequent process, the timestamp corresponding to the second frame of scene image data can be used as the new start timestamp.

[0072] (b)Intercept the initial expert trajectory segment within the timestamp interval from the expert trajectory.

[0073] (c)Based on the navigation points included in the initial expert trajectory segment, perform trajectory reconstruction to obtain a smooth expert trajectory segment. Considering that the number of navigation points included in the initial expert trajectory segment is affected by the flight speed of the UAV. For example, when the UAV flight speed is relatively fast, the number of navigation points included in the initial expert trajectory segment is relatively small, and when the UAV flight speed is relatively slow, the number of navigation points included in the initial expert trajectory segment is relatively large, resulting in the initial expert trajectory segment being irregular. Therefore, it is necessary to reconstruct it. In one example, fitting can be performed based on the navigation points included in the initial expert trajectory segment to achieve trajectory reconstruction and obtain a smooth expert trajectory segment. The fitting algorithm can adopt a third-order B-spline fitting algorithm.

[0074] (d)Sample multiple navigation points from the smooth expert trajectory segment. The sampled navigation points and their corresponding expert speed directions constitute the target expert trajectory segment. For example, 10 navigation position points and their corresponding expert speed directions can be evenly sampled from the smooth expert trajectory segment to obtain the target expert trajectory segment.

[0075] Step 4: Use the expert trajectory library, as well as the scene image data, UAV pose data, and target direction data in the second navigation data, and combine with the multi-modal loss function to train the obstacle avoidance model based on imitation learning. Specifically, it includes the following (1) to (4):

[0076] (1)Through the obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data, and target direction data in the second navigation data, output the multi-mode second speed direction index.

[0077] In one implementation, extract the first feature information of the scene image data, and extract the second feature information of the UAV pose data and the target direction data; perform downsampling processing and fusion processing on the first feature information and the second feature information to obtain the target feature information; based on the target feature information, output the multi-mode first speed direction index and its corresponding probability prediction value.

[0078] The architecture of the obstacle avoidance model based on imitation learning is as follows: To reduce the training complexity and suppress the problem of lingering around obstacles, a speed encoding and multi-mode output scheme is adopted. Specifically, refer to Figure 2 The structural schematic diagram of an obstacle avoidance model based on imitation learning shown in Figure 2 , includes: the backbone network of the image part and the backbone network of the state part. The backbone network of the image part is used to extract the first feature information of the scene image, and the backbone network of the state part is used to extract the second feature information of the UAV pose data and the target direction data; it also includes multiple convolutional layers (Conv1D), which are used to perform convolution operations after downsampling and fusion processing on the first feature information and the second feature information, and finally output the first speed direction index of the top three modes with the highest probability prediction value and its corresponding probability prediction value after decoding processing.

[0079] (2) Based on the timestamps in the second navigation data, extract the first target expert trajectory segment that matches the second navigation data from the expert trajectory library. In an example, since the expert trajectory executed when the UAV collects the second navigation data is known, multiple target expert trajectory segments intercepted from this expert trajectory can be screened out from the expert trajectory library, and then, according to the timestamps in the second navigation data, further screen out the first target expert trajectory segment that matches the second navigation data from the screened target expert trajectory segments.

[0080] (3) According to the multi-modal loss function, based on the multi-mode second speed direction index and the first target expert trajectory segment, determine the classification loss value and the speed constraint loss value, and fuse the classification loss value and the speed constraint loss value to obtain the multi-modal loss value.

[0081] In an example, the process of determining the classification loss value (Classifictaion Loss) is as follows:

[0082] (31-1) In the expert trajectory library, locate the second target expert trajectory segment corresponding to the second speed direction index of each mode. (31-2) Based on the first target expert trajectory segment, determine the label corresponding to each second target expert trajectory segment. The label is used to represent whether the second target expert trajectory segment is the first target expert trajectory segment. For example, if the second target expert trajectory segment is the first target expert trajectory segment, the label of the second target expert trajectory segment can be determined as 1, otherwise, the label of the second target expert trajectory segment is determined as 0. (31-3) According to the probability prediction value corresponding to the second speed direction index of each mode and the label corresponding to each second target expert trajectory segment, determine the classification loss value. The classification loss value can be determined according to the following formula:

[0083] ;

[0084] Among them, is the classification loss value, is the number of second navigation data, is the number of patterns, is for the th second navigation data output of the th pattern of the second speed direction index, is the second speed direction index corresponding to the label of the second target expert trajectory segment, is the second speed direction index corresponding to the probability prediction value of the second target expert trajectory segment.

[0085] In one example, the process of determining the velocity constraint loss value is as follows:

[0086] (32 - 1) Decode the second speed direction index of each pattern to obtain the speed direction prediction value corresponding to the second speed direction index of each pattern. Exemplarily, the second speed direction index can be in numerical form, and based on the mapping relationship between the second speed direction index and the turning angle of the UAV, the second speed direction index can be decoded to obtain the corresponding turning angle, and then the corresponding speed direction prediction value can be solved according to the turning angle.

[0087] The decoding formula is as follows:

[0088] ;

[0089] Among them, is the turning angle, is for the th second navigation data output of the th pattern of the second speed direction index, is the total number of target expert trajectory segments included in the expert trajectory library, is the visible range of the monocular fisheye camera, is half of the visible range of the monocular fisheye camera.

[0090] The process of determining the speed direction prediction value is as follows:

[0091] ;

[0092] Among them, is the speed direction prediction value corresponding to the second speed direction index .

[0093] (32 - 2) Determine the matching degree between the predicted speed direction corresponding to the second speed direction index of each pattern and the expert speed direction included in the first target expert trajectory segment. The matching degree can be determined according to the following formula:

[0094] ;

[0095] Among them, is the predicted speed direction value and the expert speed direction between the matching degree, is the predicted speed direction value and the expert speed direction between the included angle, is the second speed direction index corresponding predicted speed direction value, is the th expert speed direction included in the first target expert trajectory segment matched by the second navigation data, is a constant, a small value to prevent division by zero.

[0096] (32 - 3)Based on the matching degree, determine the speed constraint loss value. The speed constraint loss value can be determined according to the following formula:

[0097] ;

[0098] Among them, is the speed constraint loss value.

[0099] Based on this, the process of fusing the classification loss value and the speed constraint loss value to obtain the multi - modal loss value is as follows:

[0100] ;

[0101] Among them, is the multi - modal loss value, is the weight coefficient.

[0102] On the basis of the foregoing embodiments, the multi - modal loss value is determined according to the following formula:

[0103] ;

[0104] Through the multi - modal loss function, the obstacle - avoidance model based on imitation learning can effectively predict the speed direction index and improve the accuracy of obstacle - avoidance planning.

[0105] (4) Train the obstacle avoidance model based on imitation learning using the multi-modal loss value. In one example, the cosine annealing restarts strategy can be used to dynamically adjust the learning rate, and combined with the Adam optimizer to achieve efficient optimization of the model.

[0106] In summary, referring to Figure 3 the schematic architecture diagram of an obstacle avoidance method based on imitation learning shown in

[0107] After the model passes the verification pre-test, it can be put into the actual trajectory planning scenario. After collecting the scene image data, UAV pose data, and target direction data of the UAV, through the obstacle avoidance model based on imitation learning, according to the scene image data, UAV pose data, and target direction data, output the first speed direction index of multiple modes, including: extracting the first feature information of the scene image data, and extracting the second feature information of the UAV pose data and target direction data; performing downsampling processing and fusion processing on the first feature information and the second feature information to obtain target feature information; based on the target feature information, output the first speed direction index of multiple modes and its corresponding probability prediction value. Specifically, refer to the foregoing (1), and the embodiments of the present invention will not be elaborated herein. Locate the target expert trajectory segment corresponding to the first speed direction index with the highest probability prediction value from the expert trajectory library, send it to the UAV to control the UAV to navigate, and continue to collect new scene image data, UAV pose data, and target direction data during the navigation process, and repeat this process until the end position is reached.

[0108] Imitation learning in the embodiments of the present invention, as a technology that combines expert experience and deep learning, has significant advantages. By imitating the decision-making process of human or algorithm experts, efficient control strategies can be quickly generated.

[0109] Based on the foregoing embodiments, an obstacle avoidance device based on imitation learning is provided in an embodiment of the present invention. Refer to Figure 4 the structural schematic diagram of an obstacle avoidance device based on imitation learning shown in

[0110] A data acquisition module 402, configured to acquire first navigation data of the unmanned aerial vehicle. The first navigation data includes scene image data, unmanned aerial vehicle pose data, and target direction data. The target direction data is a vector pointing from the current position of the unmanned aerial vehicle to the end position;

[0111] An index output module 404, configured to output a multi-mode first speed direction index and its corresponding probability prediction value according to the scene image data, unmanned aerial vehicle pose data, and target direction data in the first navigation data through an obstacle avoidance model based on imitation learning; wherein, the multi-mode first speed direction index is used to locate a target expert trajectory segment to be executed in an expert trajectory library, and the probability prediction value is used to characterize the possibility that the target expert trajectory segment corresponding to the first speed direction index is called. The target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the obstacle avoidance model is trained by using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function. The classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-mode second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the unmanned aerial vehicle, and the speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment;

[0112] An unmanned aerial vehicle control module 406, configured to control the unmanned aerial vehicle to navigate according to the navigation points and their corresponding expert speed directions in the target expert trajectory segment to be executed, and continue to acquire new first navigation data during the navigation of the unmanned aerial vehicle until the unmanned aerial vehicle reaches the end position.

[0113] The obstacle avoidance device based on imitation learning provided in the embodiment of the present invention combines expert experience and deep learning technology. Through the obstacle avoidance model based on imitation learning, a multi-mode first speed direction index is output according to the scene image data, unmanned aerial vehicle pose data, and target direction data in the first navigation data, and then the unmanned aerial vehicle is controlled to navigate according to the target expert trajectory segment located by the first speed direction index. The present invention realizes the decision-making process of imitating human or algorithm experts, and not only has high robustness and flexibility, but also has a good obstacle avoidance effect.

[0114] In an implementation manner, it further includes a model training module, configured to:

[0115] Obtain multiple pre-planned expert trajectories;

[0116] Control the UAV to navigate according to the expert trajectory in scenarios of different scales, and collect second navigation data during the UAV navigation. The second navigation data includes a timestamp, scene image data, UAV pose data, and target direction data;

[0117] Intercept segments of the expert trajectory based on the timestamp to construct an expert trajectory library;

[0118] Use the expert trajectory library, as well as the scene image data, UAV pose data, and target direction data in the second navigation data, and combine with a multi-modal loss function to train an obstacle avoidance model based on imitation learning.

[0119] In one implementation, the expert trajectory library includes multiple target expert trajectory segments; the model training module is specifically used for:

[0120] Determine multiple groups of start timestamps and end timestamps according to a preset planning time to obtain multiple timestamp intervals to be intercepted;

[0121] Intercept the initial expert trajectory segments within the timestamp intervals from the expert trajectory;

[0122] Reconstruct the trajectory based on the navigation points included in the initial expert trajectory segments to obtain smooth expert trajectory segments;

[0123] Sample multiple navigation points from the smooth expert trajectory segments. The sampled navigation points and their corresponding expert speed directions form the target expert trajectory segments.

[0124] In one implementation, the model training module is specifically used for:

[0125] Through the obstacle avoidance model based on imitation learning, output a multi-mode second speed direction index according to the scene image data, UAV pose data, and target direction data in the second navigation data;

[0126] Based on the timestamp in the second navigation data, extract the first target expert trajectory segment that matches the second navigation data from the expert trajectory library;

[0127] According to the multi-modal loss function, based on the multi-mode second speed direction index and the first target expert trajectory segment, determine the classification loss value and the speed constraint loss value, and fuse the classification loss value and the speed constraint loss value to obtain the multi-modal loss value;

[0128] Use the multi-modal loss value to train the obstacle avoidance model based on imitation learning.

[0129] In one implementation, the model training module is specifically used for:

[0130] Locate the second target expert trajectory segment corresponding to the second speed direction index of each pattern from the expert trajectory library;

[0131] Determine the label corresponding to each second target expert trajectory segment based on the first target expert trajectory segment, where the label is used to characterize whether the second target expert trajectory segment is the first target expert trajectory segment;

[0132] Determine the classification loss value according to the probability prediction value corresponding to the second speed direction index of each pattern and the label corresponding to each second target expert trajectory segment.

[0133] In one implementation, the model training module is specifically configured to:

[0134] Interpret the second speed direction index of each pattern to obtain the speed direction prediction value corresponding to the second speed direction index of each pattern;

[0135] Determine the matching degree between the speed direction prediction value corresponding to the second speed direction index of each pattern and the expert speed direction included in the first target expert trajectory segment;

[0136] Determine the speed constraint loss value based on the matching degree.

[0137] In one implementation, the model training module is specifically configured to:

[0138] Determine the multi-modal loss value according to the following formula:

[0139] ;

[0140] where, is the multi-modal loss value, is the number of second navigation data, is the number of patterns, is for the th second navigation data, the th pattern's second speed direction index, is the second speed direction index corresponding to the label of the second target expert trajectory segment, is the second speed direction index corresponding to the probability prediction value of the second target expert trajectory segment, is the second speed direction index corresponding to the speed direction prediction value, is the expert speed direction included in the first target expert trajectory segment matched by the th second navigation data, is the speed direction prediction value and the expert speed direction The matching degree between is the weight coefficient, and

[0141] In one embodiment, the index output module 404 is specifically configured to:

[0142] Extract the first feature information of the scene image data, and extract the second feature information of the UAV pose data and the target direction data;

[0143] Perform downsampling processing and fusion processing on the first feature information and the second feature information to obtain target feature information;

[0144] Based on the target feature information, output the first speed direction index in multiple modes and its corresponding probability prediction value.

[0145] The device provided by the embodiment of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For a brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.

[0146] The embodiment of the present invention provides an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and the computer program executes the method described in any one of the above embodiments when being run by the processor.

[0147] Figure 5 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 50, a memory 51, a bus 52, and a communication interface 53. The processor 50, the communication interface 53, and the memory 51 are connected through the bus 52; the processor 50 is configured to execute an executable module stored in the memory 51, such as a computer program.

[0148] Among them, the memory 51 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 53 (which may be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0149] The bus 52 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0150] Among them, the memory 51 is used to store a program. After receiving an execution instruction, the processor 50 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to or implemented by the processor 50.

[0151] The processor 50 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 50 or the instructions in the form of software. The above-mentioned processor 50 may be a general-purpose processor, including a central processing unit (Central Processing Unit, CPU for short), a network processor (Network Processor, NP for short), etc.; it may also be a digital signal processor (Digital Signal Processing, DSP for short), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC for short), a field-programmable gate array (Field-Programmable Gate Array, FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being completed by the hardware decoding processor, or completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 51, and the processor 50 reads the information in the memory 51 and combines its hardware to complete the steps of the above method.

[0152] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments, which will not be elaborated here.

[0153] If the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0154] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solution of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An obstacle avoidance method based on imitation learning, characterized in that: include: Collecting first navigation data of the drone, the first navigation data including scene image data, drone posture data and target direction data, the target direction data being a vector pointing from the current position of the drone to the end position; Through an obstacle avoidance model based on imitation learning, a multi-mode first speed direction index and its corresponding probability prediction value are output according to the scene image data, the UAV posture data and the target direction data in the first navigation data; wherein the multi-mode first speed direction index is used to locate the target expert trajectory segment to be executed from the expert trajectory library, and the probability prediction value is used to characterize the possibility of calling the target expert trajectory segment corresponding to the first speed direction index, and the target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the obstacle avoidance model is trained using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function, and the classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-modal second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the UAV, and the speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment; The UAV is controlled to navigate according to the navigation points in the target expert trajectory segment to be executed and the corresponding expert speed directions, and new first navigation data are continuously collected during the UAV navigation process until the UAV reaches the terminal position.

2. The obstacle avoidance method based on imitation learning according to claim 1, characterized in that: Before outputting a multi-mode first speed direction index according to the scene image data, the drone posture data and the target direction data through an obstacle avoidance model based on imitation learning, the method further includes: Get pre-planned multiple expert trajectories; Controlling the UAV to navigate according to the expert trajectory in scenes of different scales, and collecting second navigation data during the navigation of the UAV, wherein the second navigation data includes a timestamp, the scene image data, the UAV posture data, and the target direction data; Segment-cutting the expert trajectory based on the timestamp to construct the expert trajectory library; The expert trajectory library, as well as the scene image data, the drone pose data and the target direction data in the second navigation data, are used in combination with a multimodal loss function to train an obstacle avoidance model based on imitation learning.

3. The obstacle avoidance method based on imitation learning according to claim 2, characterized in that: The expert trajectory library includes a plurality of target expert trajectory segments; The expert trajectory is segmented based on the timestamp to construct the expert trajectory library, including: Determine multiple groups of start timestamps and end timestamps according to the preset planning time to obtain multiple timestamp intervals to be intercepted; Extracting an initial expert trajectory segment located within the timestamp interval from the expert trajectory; Reconstructing the trajectory based on the navigation points included in the initial expert trajectory segment to obtain a smooth expert trajectory segment; A plurality of the navigation points are sampled from the smooth expert trajectory segment, and the sampled navigation points and their corresponding expert speed directions constitute the target expert trajectory segment.

4. The obstacle avoidance method based on imitation learning according to claim 2, characterized in that: Using the expert trajectory library, the scene image data, the drone pose data, and the target direction data in the second navigation data, combined with a multimodal loss function, an obstacle avoidance model based on imitation learning is trained, including: Outputting a multi-mode second speed direction index according to the scene image data, the drone posture data and the target direction data in the second navigation data through an obstacle avoidance model based on imitation learning; extracting, from the expert trajectory library, a first target expert trajectory segment matching the second navigation data based on the timestamp in the second navigation data; According to the multimodal loss function, based on the second speed direction index of the multimodal and the first target expert trajectory segment, a classification loss value and a speed constraint loss value are determined, and the classification loss value and the speed constraint loss value are fused to obtain a multimodal loss value; The multimodal loss value is used to train an obstacle avoidance model based on imitation learning.

5. The obstacle avoidance method based on imitation learning according to claim 4, characterized in that: Determining a classification loss value based on the second speed direction index and the first target expert trajectory segment of the multi-modal according to the multi-modal loss function includes: Locate the second target expert trajectory segment corresponding to the second speed direction index of each mode from the expert trajectory library; Determine a label corresponding to each second target expert trajectory segment based on the first target expert trajectory segment, wherein the label is used to indicate whether the second target expert trajectory segment is the first target expert trajectory segment; A classification loss value is determined according to the probability prediction value corresponding to the second speed direction index of each mode and the label corresponding to each second target expert trajectory segment.

6. The obstacle avoidance method based on imitation learning according to claim 4, characterized in that: Determining a speed constraint loss value based on the second speed direction index of the multi-mode and the first target expert trajectory segment according to the multi-modal loss function includes: Interpreting the second speed direction index of each mode to obtain a speed direction prediction value corresponding to the second speed direction index of each mode; Determine a degree of match between the speed direction prediction value corresponding to the second speed direction index of each mode and the expert speed direction included in the first target expert trajectory segment; A speed constraint penalty value is determined based on the degree of matching.

7. The obstacle avoidance method based on imitation learning according to claim 1, characterized in that: Through the obstacle avoidance model based on imitation learning, according to the scene image data, the drone posture data and the target direction data, a multi-mode first speed direction index and its corresponding probability prediction value are output, including: Extracting first feature information of the scene image data, and extracting second feature information of the drone pose data and the target direction data; Performing downsampling processing and fusion processing on the first feature information and the second feature information to obtain target feature information; Based on the target feature information, a multi-mode first speed direction index and its corresponding probability prediction value are output.

8. An obstacle avoidance device based on imitation learning, characterized in that: include: A data acquisition module, used to collect first navigation data of the UAV, wherein the first navigation data includes scene image data, UAV posture data and target direction data, wherein the target direction data is a vector pointing from the current position of the UAV to the end position; An index output module is used to output a multi-mode first speed direction index and its corresponding probability prediction value according to the scene image data, the UAV posture data and the target direction data in the first navigation data through an obstacle avoidance model based on imitation learning; wherein the multi-mode first speed direction index is used to locate the target expert trajectory segment to be executed from the expert trajectory library, and the probability prediction value is used to characterize the possibility of calling the target expert trajectory segment corresponding to the first speed direction index, and the target expert trajectory segment includes multiple navigation points and their corresponding expert speed directions; the obstacle avoidance model is trained using a multi-modal loss function, and the multi-modal loss function is obtained by fusing a classification loss function and a speed constraint loss function, and the classification loss function is used to describe the deviation between the second target expert trajectory segment corresponding to the multi-mode second speed direction index predicted by the obstacle avoidance model and the first target expert trajectory segment actually executed by the UAV, and the speed constraint loss function is used to describe the deviation between the speed direction prediction value corresponding to the second speed direction index and the expert speed direction included in the first target expert trajectory segment; The drone control module is used to control the drone navigation according to the navigation points and the corresponding expert speed directions in the target expert trajectory segment to be executed, and continue to collect new first navigation data during the drone navigation process until the drone reaches the terminal position.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous navigation and obstacle avoidance method based on deep reinforcement learning

    CN117193355A

  • Aircraft approach planning method and device based on generative adversarial imitation learning

    CN118411858A