Man-machine collaborative decision planning system and method based on environment adaptive trajectory optimization
Through the environmental adaptive trajectory optimization system, the drone speed parameters are optimized in combination with semantic information and depth information, the burden problem of pilots in complex environments is solved and navigation efficiency and safety is improved.
Patent Information
- Application Number
- CN202510420421.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-04
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing human-aircraft collaborative decision planning system, pilots have heavy cognitive and operational burdens in complex environments, and the drone trajectory planning speed adjustment is not flexible enough to provide sufficient decision response time, resulting in insufficient track safety.
Through the environmental adaptive trajectory optimization system, combined with the semantic information and depth information in the environment perception, the speed parameters of the planner are dynamically adjusted, and real-time state estimation and map construction are used to use the RGB-D camera and RTK high-precision positioning equipment to optimize the flight speed of the drone, combining topological path search, visibility detection and environmental adaptive strategies.
It effectively reduces the cognitive and operational burden of pilots, improves navigation efficiency and system safety, provides more decision-making response time, and reduces dependence on manual speed commands.
Smart Images

Figure CN120276422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of human-machine collaboration, specifically a human-machine collaborative decision-making and planning system and method based on environment-adaptive trajectory optimization. Background Art
[0002] Human-machine collaborative decision-making and planning is a core technology in non-cooperative target scenarios. One of the challenges of existing human-machine collaborative decision-making and planning is the relatively heavy cognitive and control burden on pilots. Due to the timeliness requirements of search and rescue missions, drones usually maintain a high cruising speed, resulting in insufficient decision-making reaction time for pilots when passing through intersections. In addition, when drones pass through obstacles with different densities, they rely on pilots to manually adjust the speed. Since pilots cannot accurately estimate the depth of obstacles, manual speed commands usually have delays and mutations, leading to an increase in the pilot's operation burden and insufficient track safety. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art in that the speed adjustment of the planned trajectory is not flexible enough, unable to provide sufficient decision-making reaction time for pilots, and relying on manual speed commands, the present invention proposes a human-machine collaborative decision-making and planning system and method based on environment-adaptive trajectory optimization. By considering semantic information and depth information in environmental perception, the speed parameters of the planner are autonomously and dynamically adjusted, effectively improving the adaptability of the planned trajectory to different scenarios, thereby reducing the cognitive and operation burden of pilots in complex environments and enhancing navigation efficiency and system safety.
[0004] The present invention is realized by the following technical solutions:
[0005] The present invention relates to a human-machine collaborative decision-making and planning system based on environment-adaptive trajectory optimization, including: a human-machine interaction module, a rule-based motion planning module, and a learning-based parameter generation module. Among them: the human-machine interaction module processes data such as environmental perception and state estimation according to the data collected by on-board sensors and remote controls, and obtains the input information required by the motion planning and parameter generation modules; the motion planning module performs topological path search, visibility detection, and trajectory optimization processing according to the local map and user instruction information, and obtains the planned trajectory for drone tracking control; the parameter generation module performs environment-adaptive strategy processing according to the distribution information of the trajectory in the environment, and obtains the speed constraint parameters for adjusting motion planning.
[0006] The described human-computer interaction module includes: on-board sensors, a state estimation unit, and an online mapping unit. Among them: The on-board sensors include a Real-time kinematic (RTK) high-precision positioning device and a RGB-D camera, which use GPS differential positioning information and depth point cloud spatial geometric transformation respectively for fuselage state estimation and online map construction; The state estimation unit performs pose solution operation processing based on the information collected by the on-board sensors to obtain the state information of the UAV; The online mapping unit performs spatial geometric transformation processing based on the UAV state and camera depth information to obtain a local grid map.
[0007] The described RGB-D camera feeds back environmental image information to the pilot through a video transmission device, and the pilot uses a remote control handle to generate control commands with two types of information: direction and speed based on this.
[0008] The described rule-based motion planning module includes: a topological path search unit, a fast visibility detection unit, a human-machine shared decision-making unit, and a trajectory optimization unit. Among them: The topological path search unit performs sampling and graph structure processing based on the local grid map information to extract potential candidate paths containing semantic information; The fast visibility detection unit marks the visible intervals according to whether there are candidate paths in the field of view of the current trajectory; The human-machine shared decision-making unit selects a guiding path from the candidate paths according to the user's direction command; The trajectory optimization unit performs multi-objective optimization solution based on the guiding path marked by the visible interval and the speed constraint parameters to generate a warm-up trajectory and a planned trajectory respectively.
[0009] The described learning-based parameter generation module includes: an observation space unit, an environment adaptive strategy pre-training unit, and an environment adaptive strategy fine-tuning unit. Among them: The observation space unit extracts the spatial distribution information of the warm-up trajectory and obstacles based on the local grid map to generate the input information of the policy network; The environment adaptive strategy pre-training unit performs interactive training in a simulation environment based on the priori knowledge reward designed manually to generate a speed policy network with environment adaptability; The environment adaptive strategy fine-tuning unit fine-tunes the policy network according to the artificial feedback reward composed of the user's speed command to generate adaptive speed parameters that meet the user's preferences, and sends them to the trajectory optimization unit for environment adaptive optimization. Technical effects
[0010] Through the visibility detection technology, the present invention can detect and identify potential passable areas in the field of view of an on-board camera, and mark the path intervals where the pilot may make decisions; through the environmental adaptive parameter generation algorithm, it evaluates the environmental complexity and automatically adjusts the flight speed of the drone according to different user preferences. Compared with the prior art, the present invention can provide more decision-making reaction time for the pilot, reduce the system's dependence on manual speed commands, and while reducing the pilot's cognitive and operation burdens, improve the navigation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the system of the present invention;
[0012] Figure 2 It is a flowchart of the topological path search algorithm in the embodiment;
[0013] In the figure: (a) is for constructing a local graph, (b) is for generating a shortest path schematic diagram, (c) is for classifying homotopy paths schematic diagram, and (d) is for merging homotopy paths schematic diagram;
[0014] Figure 3 It is a visibility detection graph adopted in the embodiment;
[0015] In the figure: (a) is a geometric schematic diagram of the perception cone, and (b) is a schematic diagram of visibility discrimination;
[0016] Figure 4 It is a schematic diagram of the human-machine shared decision-making mechanism adopted in the embodiment;
[0017] In the figure: (a) is for adopting autonomous decision-making schematic diagram, and (b) is for manual decision-making schematic diagram;
[0018] Figure 5 It is a schematic diagram of the environmental observation information adopted in the embodiment;
[0019] Figure 6 It is a flowchart of the embodiment;
[0020] Figure 7 It is a schematic diagram of the experimental scenario in the embodiment;
[0021] In the figure: (a) and (b) are two different test scenarios respectively;
[0022] Figure 8 It is a schematic diagram of the effect in the embodiment and a comparison diagram with the existing human-machine collaborative decision-making and planning method;
[0023] In the figure: (a)-(c) are the result comparisons of the comparison methods No-EA, EA and in scenario 1 respectively; (d)-(f) are the result comparisons of the comparison methods No-EA, EA and in scenario 2 respectively. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] AsFigure 1 As shown in the figure, this embodiment relates to a human - machine collaborative decision - making and planning method based on the above - mentioned system, including:
[0025] Step 1) State estimation and online mapping: Start the on - board processor and the RTK positioning device, and subscribe to the pose data published by the RTK through the Robot Operating System (ROS); start the RGB - D camera and subscribe to the depth image data through the ROS; the mapping (SLAM) algorithm inside the planner generates a local grid map according to the body pose and the depth image data.
[0026] The pose data refers to the position p = [x, y, z] on three axes in the Euclidean space T , and the attitude angles Θ = [φ, θ, ψ] T .
[0027] The size of the global grid map is 100×100m, the local mapping depth is 10m, and the grid resolution is 0.1m.
[0028] Step 2) User instruction generation: The UAV sends the image data collected by the RGB - D camera to the user - side image display device through the video transmission system, and the user generates user instructions using the remote control handle according to the observed environmental information.
[0029] The user instructions contain two parts of information: direction and speed. Among them: The direction information refers to: substituting u h into the kinematic model to generate the user instruction primitive γ h , and the path duration t ∈ [0, T] is set manually; the speed information refers to the speed instructions on different axes in u h
[0030] Step 3) Topological path search: As Figure 2 shown, search for non - homotopic candidate paths in the local map obtained in Step 1, which represent potential user decisions, specifically including:
[0031] 3.1 Construct a local graph, specifically including: Based on the local grid map The graph - structure data generated by iterative sampling Among them: The root node V 0 represents the current position of the UAV, and the Euclidean distance E i between the node V j and the node V i,j is defined as an edge.
[0032] 3.2 Generate the shortest path, specifically including: According to the local graph structure G = {V k , Ei,j}, the set of shortest paths Γ = {γ 0} from the root node V searched by the Dijkstra algorithm to the top node, i ∈ N i}, i ∈ N p .
[0033] 3.3 Classify homotopy paths, specifically including: distinguishing topological paths with different semantic features. In this embodiment, the discrete monotonic reparameterization mechanism is used to sample i points on the path γ j (s) and the path γ discrete points, and then determine whether the corresponding paths are homotopy according to whether the corresponding connection lines between all discrete points s ∈ d [0, N collide with obstacles.
[0034] 3.4 Merge homotopy paths, specifically including: selecting the one with the largest safety margin as the candidate path from the set of paths in the same class m ∈ N l . The safety margin is defined as the sum of the degrees of the nodes on the path: N vi represents the number of nodes on the path , degree(V k ) represents the number of neighboring nodes of the node V k . A lower degree(V k ) means that the node is close to the obstacle.
[0035] In this embodiment, the number of effective sampling nodes K = 100.
[0036] In this embodiment, the number of shortest paths N p = 10.
[0037] In this embodiment, the number of discrete points N d = 10.
[0038] In this embodiment, the number of candidate paths N l ≤ 3.
[0039] Step 4) Visibility detection: As Figure 3As shown, according to the current flight path and the candidate paths obtained in step 3, visibility detection is used to assist the user in judging whether there are potential candidate paths in the current field of view, and the potential decision-making intervals of the user in the current flight path are marked. Specifically: in this embodiment, the field of view of the on-board RGB-D camera is modeled as a perception cone, and the potential feasible region is approximated as the convex hull of the perception cone at the up-sampled points of the candidate path. By judging whether the perception cones at different positions of the current path intersect with the convex hull to adjust the flight speed of the UAV, thereby reducing the user's cognitive burden.
[0040] The described perception cone is composed of five intersecting half-planes: k = 1, 2…5}, (x, y, z) represents the coordinates in the internal space of the perception cone, is the normal vector of the k-th half-plane, b k represents the offset coefficient of the k-th half-plane.
[0041] The described convex hull S m is the candidate path formed by the perception cones at different sampling points on : i = 1, 2…N w .
[0042] In this embodiment, the number of sampling points N of the candidate path w = 5.
[0043] The judgment of whether the perception cone intersects with the convex hull can be solved by linear programming calculation.
[0044] Step 5) Human-machine shared decision-making: As Figure 4 shown, according to the candidate paths obtained in step 3, using the user instruction primitive γ generated in step 2 h to screen them, and select the guidance path γ g , specifically: when the user does not send a control instruction, the human-machine shared decision-making unit selects the one with the minimum control cost x from the candidate paths according to the current motion state v = [v y , v z T as the guidance path. When the user sends a control instruction, the human-machine shared decision-making unit evaluates the similarity between the candidate paths and the user instruction primitive, and selects the one with the highest similarity to the user instruction u as the guidance path γ h . . g .
[0045] The described control cost where: Indicates the instantaneous motion trend at the sampling point p(i) on the corresponding candidate path.
[0046] The similarity metric described above Where: user instruction u h Indicates the instantaneous motion trend expected by the user when sending the instruction. Other variable definitions are the same as above.
[0047] Step 6) Warm-up trajectory optimization: According to the guidance path γ generated by human-machine shared decision-making g , the trajectory optimization unit uses the visibility detection result To generate a warm-up trajectory ξ that satisfies multi-objective constraints w .
[0048] The visibility detection result described above There are two cases: when , there is no candidate path in the RGB-D camera's field of view. Therefore, the constraint speed parameter v * = v max To ensure navigation efficiency; when , there is a candidate path in the RGB-D camera's field of view. Therefore, the constraint speed parameter v * = v min To provide the user with more reaction time.
[0049] In this embodiment, the maximum flight speed v max = 2m / s, and the minimum flight speed v min = 0.5m / s.
[0050] The multi-objective constraints described above include: safety constraints, trajectory smoothness constraints, speed constraints, and trajectory shape constraints. The safety constraint is defined as the distance between the trajectory and surrounding obstacles is not less than the minimum collision distance d min = 1.0m; the trajectory smoothness constraint is defined as the integral of the norm of the jerk; the speed constraint is defined as the speed of the trajectory control points not exceeding the speed parameter; the trajectory shape constraint is defined as the distance between the control points of the trajectory and the corresponding guidance path should be as small as possible.
[0051] In this embodiment, the flight trajectory is fitted by a uniform B-spline curve.
[0052] Step 7) Generate the observation space: The observation space of the environment adaptation strategy π Includes the instantaneous motion trend v t And the environmental observation information ε t , v π Is the speed parameter output by the environment adaptation strategy network, which is used to adjust the flight aggressiveness of the UAV.
[0053] The environmental observation information ε described above t Refers to: the warm-up trajectory ξw Spatial distribution relationship with obstacles. For example Figure 5 As shown, to evaluate the collision risk of the UAV in a complex environment, discrete points are upsampled from the warm-up trajectory ξ w and rays in the fan-shaped area are generated in combination with the instantaneous motion trend v of the UAV. The distance d between the sampling points and the obstacles on the rays is recorded and used as the characteristic element of the environmental observation information ε t .
[0054] Step 8) Pre-train the environment adaptation policy network: In the pre-training stage, the environment adaptation policy π guides the network π ea to update the gradient according to the prior knowledge reward R ea and generates the velocity parameter v with preliminary environment adaptation from the observation space information o π .
[0055] The prior knowledge reward R mentioned above ea =[r speed ,r smoothing ,r risk pre-trains the policy network through a rule-based explicit function, including: navigation efficiency reward Smoothing reward Environment adaptation reward
[0056] In this embodiment, the sampling number N of the observation state during the calculation of the reward function o =20, and the number of rays in the environmental observation data is N ray =15;
[0057] In this embodiment, the weights of [r speed ,r smoothing ,r risk are in turn: 1.0, 1.0, 10.0;
[0058] In this embodiment, the Soft Actor-Critic reinforcement learning algorithm is used to train the policy network.
[0059] Step 9) Fine-tune the environment adaptation policy network: In the fine-tuning stage, the environment adaptation policy π guides the network π h to update the gradient according to the user feedback reward R h and generates the corrected velocity parameter Δv with user preferences from the observation space information o π . To improve the convergence speed of the policy network π h , the action search space of the network π h is constrained within the neighborhood of the network π ea .
[0060] The user feedback reward R mentioned above h =[rhuman ]Inject user preferences into the policy network through user feedback, and its main reward function term
[0061] The strategy network π h The action search space is constrained in the network π ea The neighborhood of π h Sampling in regions of invalid motion reduces the amount of user feedback required for fine-tuning.
[0062] In this embodiment, the strategy network π ea The neighborhood range is: [-0.5,0.5]m / s.
[0063] Step 10) Environment Adaptive Trajectory Optimization: The speed parameter v generated by the environment adaptive strategy network is π Alternative constraint speed parameter v * =v π To optimize the planned trajectory for the second time. Finally, the planned trajectory ξ is published to the flight control module for tracking control.
[0064] The multi-objective constraints in the secondary optimization are consistent with the preheating trajectory generation step, and only the parameters of the speed constraint optimization item are modified.
[0065] The main steps of this embodiment are as follows: Figure 6 shown.
[0066] The flight control module in this embodiment is open source PX4 firmware.
[0067] Through specific practical experiments, the effectiveness of the method of the present invention is evaluated in two artificially constructed scenarios. Scenario 1: Figure 7 As shown in (a), the user is required to guide the drone to the target ① as quickly as possible. If there is a decision intersection in the environment, the user guides the drone to the target ②. Figure 7 As shown in (b), the user is required to guide the drone through sparse obstacles ① and dense obstacles ② as quickly as possible. Under this scenario task setting, three novice users use the remote controller to evaluate the baseline method (No-EA), the existing method (EA), and the proposed method (HEA). Special note: The user is not aware of the experimental scenario in advance to avoid the influence of prior information on the experimental results.
[0068] In the embodiment of scenario 1: due to the unknown environment and the high cruising speed of the drone, it is usually difficult for novice users to make steering decisions in time. Figure 8 As shown in (a), the No-EA method completely relies on user instructions to adjust the flight speed. Therefore, the short flight time (1.6s) of the drone at the intersection is not enough to support the novice user to make a decision, resulting in an emergency stop. Figure 8As shown in (b), the existing EA method does not consider the semantic information in environmental perception, which also results in a short flight time (1.5 s) at intersections. As Figure 8 As shown in (c), the HEA method of the present invention extracts semantic information in the environment through topological path search, and combines visibility detection and speed constraint optimization to extend the flight time of the UAV at the intersection decision (3.4 s), thus alleviating the cognitive burden of users.
[0069] In the embodiment of Scenario 2: Due to inaccurate depth estimation of obstacles, novice users usually have difficulty accurately adjusting the flight speed of the UAV. As Figure 8 As shown in (d), the No-EA method completely relies on user instructions to adjust the flight speed, resulting in a collision when the UAV passes through the dense obstacle ②. As Figure 8 As shown in (e), the existing EA method can use the depth information in environmental perception to adjust the flight speed of the UAV, but the rule-based environmental adaptive planner cannot incorporate user feedback. Therefore, the active speed reduction of the UAV when passing through the sparse obstacle ① does not conform to user preferences. As Figure 8 As shown in (f), the HEA method of the present invention generates speed parameters through an environmental adaptive strategy that incorporates user feedback, and adaptively adjusts the flight speed of the UAV in combination with trajectory optimization, thereby alleviating the operation burden of users.
[0070] To further evaluate the statistical characteristics of the method of the present invention, 12 users with different proficiencies were required to guide the UAV to perform tasks in a simulation experiment. To avoid the influence of environmental prior information on the experimental results, each user tested each method 3 times in different environments, and the experimental results are shown in Table 1.
[0071] Table 1: Comparative experimental results of the method of the present invention and the prior art methods
[0072] The experimental results show that: compared with the prior art methods No-EA and EA, the method of the present invention can extend the average decision reaction time reserved for users at intersections by 40.43% and 103.97%, and reduce the average operation burden of users by 46.33% and 43.71%. In summary, the method of the present invention can effectively reduce the cognitive burden of users by delaying the decision reaction time, and the environmental adaptive mechanism significantly reduces the operation burden and operation burden indicators of users, and indirectly optimizes other efficiency indicators.
[0073] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments. All implementation solutions within its scope are subject to the present invention.
Claims
1. A human-machine collaborative decision-making and planning system based on environment-adaptive trajectory optimization, characterized in that, Including: A human-computer interaction module, a rule-based motion planning module, and a learning-based parameter generation module, where: The human-computer interaction module processes data such as environmental perception and state estimation based on the data collected by the on-board sensors and the remote controller, and obtains the input information required by the motion planning and parameter generation modules; The motion planning module performs topological path search, visibility detection, and trajectory optimization processing based on the local map and user instruction information, and obtains the planned trajectory for UAV tracking control; The parameter generation module performs environmental adaptive strategy processing based on the distribution information of the trajectory in the environment, and obtains the speed constraint parameters for adjusting the motion planning.
2. The human-machine collaborative decision-making and planning system based on environment-adaptive trajectory optimization according to claim 1, characterized in that, The described human-computer interaction module includes: on-board sensors, a state estimation unit, and an online mapping unit, where: The on-board sensors include a real-time kinematic high-precision positioning device and a color depth camera, which respectively use GPS differential positioning information and depth point cloud spatial geometric transformation for fuselage state estimation and online map construction; The state estimation unit performs pose solution operation processing based on the information collected by the on-board sensors, and obtains the state information of the UAV; The online mapping unit performs spatial geometric transformation processing based on the UAV state and camera depth information, and obtains a local grid map.
3. The human-machine collaborative decision-making and planning system based on environment-adaptive trajectory optimization according to claim 1, wherein, The described rule-based motion planning module includes: a topological path search unit, a fast visibility detection unit, a human-machine shared decision-making unit, and a trajectory optimization unit, where: The topological path search unit performs sampling and graph structure processing based on the local grid map information, and extracts potential candidate paths containing semantic information; The fast visibility detection unit marks the visible intervals according to whether there are candidate paths in the field of view of the current trajectory; The human-machine shared decision-making unit selects a guiding path from the candidate paths according to the user's direction instruction; The trajectory optimization unit performs multi-objective optimization solution based on the guiding path marked by the visible intervals and the speed constraint parameters, and generates a warm-up trajectory and a planned trajectory respectively.
4. The human-machine collaborative decision-making and planning system based on environment-adaptive trajectory optimization according to claim 1, characterized in that, The described learning-based parameter generation module includes: an observation space unit, an environmental adaptive strategy pre-training unit, and an environmental adaptive strategy fine-tuning unit, where: The observation space unit extracts the spatial distribution information of the warm-up trajectory and obstacles based on the local grid map, and generates the input information of the policy network; The environmental adaptive strategy pre-training unit performs interactive training in a simulation environment according to the artificially designed prior knowledge reward, and generates a speed policy network with environmental adaptability; The environmental adaptive strategy fine-tuning unit fine-tunes the policy network according to the artificial feedback reward composed of the user's speed instruction, generates adaptive speed parameters that conform to the user's preference, and sends them to the trajectory optimization unit for environmental adaptive optimization.
5. A human-machine collaborative decision-making and planning method based on the system described in any one of claims 1-4, characterized in that, Including: Step 1) State estimation and online mapping: Start the on-board processor and RTK positioning device, and subscribe to the pose data published by RTK through the Robot Operating System (ROS); Start the RGB-D camera, and subscribe to the depth image data through ROS; The mapping (SLAM) algorithm inside the planner generates a local grid map based on the fuselage pose and depth image data; The pose data mentioned above refers to the position p = [x, y, z] on three axes in the Euclidean space T , and the attitude angles Θ = [φ, θ, ψ] T ; The user instruction described above contains two parts of information: direction and speed, where: The direction information refers to: substituting u h into the kinematic model to generate the user instruction primitive γ h , where the path duration t ∈ [0, T] is set manually; The speed information refers to: u h The speed commands in different axial directions in Step 2) User instruction generation: The drone sends the image data collected by the RGB-D camera to the user-side image display device through the video transmission system. The user generates user instructions using the remote control handle according to the observed environmental information; Step 3) Topological path search: Search for non-homotopic candidate paths in the local map obtained in Step 1, which are potential user decisions; Step 4) Visibility detection: Based on the current flight path and the candidate paths obtained in Step 3, use visibility detection to assist the user in determining whether there are potential candidate paths in the current field of view, and mark the potential decision intervals of the user in the current flight path. Specifically: Model the field of view of the on-board RGB-D camera as a perception cone, approximate the potential feasible region as the convex hull of the perception cone at the upsampling points on the candidate path, and adjust the flight speed of the UAV by determining whether the perception cones at different positions on the current path intersect with the convex hull, thereby reducing the user's cognitive burden; to reduce the user's cognitive burden; Step 5) Human-machine shared decision-making: Based on the candidate paths obtained in Step 3, use the user instruction primitive γ generated in Step 2 h to screen them and select the guiding path γ g , specifically: When the user does not send a control instruction, the human-machine shared decision-making unit determines according to the current motion state of the UAV v = [v x , v y , v z T to select the one with the minimum control cost from the candidate paths as the guiding path. When the user sends a control instruction, the human-machine shared decision-making unit evaluates the similarity between the candidate paths and the user instruction primitive, and selects the one with the highest h similarity as the guiding path γ g ; Step 6) Preheating trajectory optimization: According to the guiding path γ generated by human-machine shared decision-making g , the trajectory optimization unit uses the visibility detection result to generate a preheating trajectory ξ that satisfies multi-objective constraints w ; The described visibility detection result There are two cases: When this is the case, there is no candidate path in the RGB-D camera's field of view, so the constrained velocity parameter v * = v max is used to ensure the navigation efficiency; When this is the case, there is a candidate path in the RGB-D camera's field of view, so the constrained velocity parameter v * = v min is used to provide the user with more reaction time; The multi-objective constraints include: safety constraints, trajectory smoothness constraints, speed constraints, and trajectory shape constraints. The safety constraint is defined as the distance between the trajectory and surrounding obstacles being not less than the minimum collision distance d min = 1.0 m; the trajectory smoothness constraint is defined as the integral of the norm of the jerk; the speed constraint is defined as the speed of the trajectory control points not exceeding the speed parameter; the trajectory shape constraint is defined as the distance between the control points of the trajectory and the corresponding guiding path being as small as possible; Step 7) Generate the observation space: the observation space of the environment adaptation strategy π includes the instantaneous motion trend v t and the environmental observation information ε t , v π is the speed parameter output by the environment adaptation strategy network, which is used to adjust the flight aggressiveness of the UAV; The environmental observation information ε t refers to: the preheating trajectory ξ w and the spatial distribution relationship with obstacles. As shown in Figure 5, in order to evaluate the collision risk of the UAV in a complex environment, discrete points are sampled from the preheating trajectory ξ w and rays in the fan-shaped area are generated in combination with the instantaneous motion trend v of the UAV. The distance d between the sampling points and the obstacles on the rays is recorded and used as the characteristic element of the environmental observation information ε t ; Step 8) Pre-train the environment adaptation policy network: In the pre-training phase, the environment adaptation policy π is guided by the prior knowledge reward R ea to guide the network π ea to perform gradient update, and generate the velocity parameter v with preliminary environment adaptation from the observation space information o π ; The prior knowledge reward R ea = [r speed , r smoothing , r risk is pre-trained by a rule-based explicit function to guide the policy network, including: navigation efficiency reward Smoothing reward Environment adaptation reward Step 9) Fine-tune the environment adaptation policy network: In the fine-tuning stage, the environment adaptation policy π is guided by the user feedback reward R h to guide the network π h to perform gradient update, generate a modified velocity parameter Δv with user preferences from the observation space information o π To improve the convergence speed of the policy network π j constrain the action search space of the network π h within the neighborhood of the network π ea ; The user feedback reward R h = [r human Inject user preferences into the policy network through user feedback, and its main reward function term Step 10) Environment-adaptive trajectory optimization: Replace the constrained velocity parameter v with the velocity parameter v generated by the environment-adaptive policy network to perform quadratic optimization on the planned trajectory. Finally, publish the planned trajectory ξ to the flight control module for tracking control. π Replace the constrained velocity parameter v * = v π to perform quadratic optimization on the planned trajectory. Finally, publish the planned trajectory ξ to the flight control module for tracking control.
6. The human-machine collaborative decision-making and planning method according to claim 5, wherein The said Step 3 specifically includes: (3.1) Construct a local map, specifically including: Based on the local grid map The graph structure data G = {V k , E i,j}, k ∈ [1, K], Where: The root node V 0 is the current position of the UAV, and the Euclidean distance E i between node V j and node V i , j is defined as an edge; (3.2) Generate the shortest path, specifically: According to the local graph structure G = {V k , E i,j} constructed in 3.1, search for the set of shortest paths Γ = {γ 0} between the root node V i and the top node through the Dijkstra algorithm, where i ∈ N p ; (3.3) Classify homotopy paths, specifically: through a discrete monotonic reparameterization mechanism for path γ i (s) and path γ j (s) sample discrete points, and then according to all discrete points s ∈ [0, N d , the corresponding connecting lines are used to determine whether the corresponding paths are homotopy by checking if they collide with obstacles; (3.4) Combine homotopy paths, specifically: select the one with the maximum safety margin as the candidate path from the path sets in the same class The safety margin is defined as the sum of the degrees of the nodes on the path: N vi is the path The number of nodes on, degree(V k ), for node V k The number of neighboring nodes of, a lower degree(V k ) means that the node is close to the obstacle.
7. The human-machine collaborative decision-making and planning method according to claim 5, characterized in that The described perception cone is five intersecting half-planes: (x, y, z) are the coordinates in the internal space of the perception cone, is the normal vector of the k-th half-plane, b k is the bias coefficient of the k-th half-plane; The convex hull S m is a candidate path formed by the perception cones at different sampling points 8. The human-machine collaborative decision-making and planning method according to claim 5, characterized in that The control cost described Wherein: is the instantaneous motion trend at the sampling point p(i) on the corresponding candidate path; The similarity measure described above where: user instruction u h is the instantaneous motion trend expected by the user when sending the instruction, and the definitions of other variables are the same as above.
9. The human-machine collaborative decision-making and planning method according to claim 5, characterized in that The above-mentioned strategy network π h with its action search space constrained within the network π ea in the neighborhood avoids sampling of π h in the invalid action area, thus reducing the amount of user feedback required for fine-tuning.
10. The human-machine collaborative decision-making and planning method according to claim 5, wherein The multi-objective constraints in the secondary optimization are consistent with the preheating trajectory generation step, and only the parameters of the speed constraint optimization term are modified.
Citation Information
Cited By
Data setting recovery method and device of remote controller
CN121187857A
Flight path planning and attitude control system of forest deratization unmanned aerial vehicle
CN121879405A
A flight trajectory planning and attitude control system of a forest rat-eliminating unmanned aerial vehicle
CN121879405B