Path planning method and system based on deep reinforcement learning, and electronic equipment

By designing a multi-objective optimization reward function and dynamic adjustment of weight coefficients, combined with the training and real-time update strategy of deep reinforcement learning models, the shortcomings of path planning in the existing technology in multi-objective optimization and dynamic environment are solved, and a better and more flexible path planning effect is achieved.

CN119984290AInactive Publication Date: 2025-05-13QINGDAO AUTOMATIC RES INST

Patent Information

Application Number
CN202510464767.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN119984290A_ABST
    Figure CN119984290A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of path planning scheme design based on deep reinforcement learning, in particular to a path planning method and system based on deep reinforcement learning and electronic equipment. According to the method, a reward function is designed by comprehensively considering path length, smoothness, feasibility and multi-objective optimization weight through a deep reinforcement learning model, the model is trained by using experience playback and exploring and utilizing a balance strategy, and the generated path is used for real-time planning after being subjected to smoothness and feasibility correction. The system comprises an environment perception module, an action definition module, a reward function design module, a model construction and training module, a path optimization module, a plan updating module and the like. The electronic device covers a processor, a memory, and input and output modules to implement real-time path planning. According to the method, the quality and efficiency of path planning can be improved, the method adapts to a complex dynamic environment, potential risks caused by frequent turning or sudden change are effectively avoided, the practicability and safety of the path are improved, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path planning scheme design based on deep reinforcement learning, and specifically to a path planning method and system and electronic equipment based on deep reinforcement learning. Background Art

[0002] With the rapid development of science and technology, path planning technology plays a vital role in many fields. From the navigation of self-driving cars, to the operation of robots in complex environments, to the application of drones in logistics and distribution, efficient and accurate path planning solutions are needed to achieve navigation and decision-making.

[0003] Traditional path planning algorithms, such as Dijkstra algorithm and A* algorithm, perform well in simple environments and small-scale problems. However, when faced with complex, changeable, and dynamic real environments, these traditional algorithms expose many limitations. They often rely on pre-set maps or environmental models, and it is difficult to adjust the path in real time when encountering dynamic obstacles, unknown areas, or environmental changes.

[0004] Path planning methods based on deep reinforcement learning have emerged. Deep reinforcement learning allows the agent to interact with the environment and learn the optimal strategy based on the reward signal fed back by the environment, which can adaptively handle the uncertainty in complex environments. However, existing path planning solutions based on deep reinforcement learning still face many challenges.

[0005] On the one hand, the design of the reward function is crucial to the effectiveness of path planning, but the existing solutions are difficult to balance the relationship between multiple objectives such as path length, smoothness, and feasibility when optimizing multiple objectives, resulting in the generated path possibly not meeting actual needs. On the other hand, in a dynamic environment, the model's real-time and adaptability are insufficient, making it difficult to quickly respond to sudden changes and uncertainties in the environment.

[0006] Therefore, the prior art needs to be further developed. Summary of the invention

[0007] The purpose of the present invention is to overcome the above-mentioned technical deficiencies and provide a path planning method and system, and electronic device based on deep reinforcement learning to solve the problems existing in the prior art.

[0008] To achieve the above technical objectives, according to a first aspect of the present invention, the present invention provides a path planning method based on deep reinforcement learning, comprising: S100, environmental perception and state representation: acquiring environmental information through sensors, and converting the environmental information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes current position, target position, and obstacle distribution information in the environment; S200, action space definition: defining the action space of the agent in path planning, wherein the action space includes movement actions of different directions and step lengths; S300, Reward function design: Design a reward function that includes path length, smoothness, feasibility, and multi-objective optimization weights, where the path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment; for multi-objective optimization, assign an independent weight coefficient to each objective, and dynamically adjust the weight coefficient according to task requirements to balance the optimization relationship between different objectives; S400, Model Building and Training: Build a deep reinforcement learning model, including a policy network and a value network, pre-train in a simulated environment, and optimize training using a balance strategy in combination with experience replay and exploration; S500, path optimization: after the model generates the initial path, the spline interpolation algorithm is used to smooth the path, the collision detection algorithm is used to ensure the feasibility of the path, and the path is corrected; S600, real-time planning and updating: In actual scenarios, the path planning results are generated based on the current environmental status input model, and the strategy is dynamically updated according to environmental changes.

[0009] Specifically, the sensor includes at least one of the following: LiDAR, cameras and ultrasonic sensors.

[0010] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: The smoothness weight is dynamically adjusted according to the curvature change and the number of inflection points of the path. The greater the curvature change or the greater the number of inflection points, the higher the smoothness weight.

[0011] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: Curvature change calculation: discretize the path and divide it into a series of line segments. Calculate the angle change between adjacent line segments as the local curvature change of the path. By integrating or summing the local curvature change, the curvature change index of the entire path is obtained. The larger the curvature change index, the higher the curvature degree of the path. Calculation of the number of inflection points: The inflection points are determined by analyzing the changes in the tangent direction of each point on the path. When the angle between the tangent directions of adjacent line segments exceeds a certain threshold, the point is considered to be an inflection point. The number of inflection points in the path is counted as another indicator to measure the smoothness of the path. Smoothness weight adjustment strategy: set a baseline curvature and a baseline number of inflection points. When the curvature change of the actual path exceeds the baseline curvature or the number of inflection points exceeds the preset baseline number of inflection points, increase the smoothness weight according to the preset ratio.

[0012] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: The feasibility weight is dynamically adjusted according to the density of obstacles in the environment. The denser the obstacles are, the higher the feasibility weight is.

[0013] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: Obstacle density assessment: Analyze the obstacle distribution in the environment, divide the environment into several small areas using the spatial division method, count the number of obstacles or the occupied area in each area as the obstacle density index of the area, and obtain the obstacle density assessment value of the entire environment by performing weighted average or summation operations on the obstacle density index of all areas; Feasibility weight adjustment function: Establish an adjustment function between the density of obstacles and the feasibility weight. This function is a monotonically increasing function. As the density of obstacles increases, the feasibility weight increases accordingly according to the law of the adjustment function. Real-time update mechanism: During the path planning process, the changes of obstacles in the environment are monitored in real time. When the position or number of obstacles changes, the obstacle density of the environment is re-evaluated and the feasibility weight is updated according to the adjustment function.

[0014] Specifically, in step S500, the spline interpolation algorithm uses a cubic spline interpolation method to generate a smooth path curve by performing interpolation calculations on key points on the path.

[0015] Specifically, in step S500, the collision detection algorithm adopts a collision detection method based on a distance field, and determines whether there is a collision risk by calculating the distance between each point on the path and the obstacle.

[0016] According to a second aspect of the present invention, there is provided a path planning system based on deep reinforcement learning, comprising: Environmental perception and state representation module: used to obtain environmental information through sensors and convert the environmental information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes the current position, the target position, and the obstacle distribution information in the environment; Action space definition module: used to define the action space of the agent in path planning, which includes movement actions of different directions and step lengths; Reward function design module: used to design a reward function that includes path length, smoothness, feasibility, and multi-objective optimization weights. The path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment. For multi-objective optimization, an independent weight coefficient is assigned to each objective, and the weight coefficient is dynamically adjusted according to task requirements to balance the optimization relationship between different objectives. Model building and training module: used to build deep reinforcement learning models, including policy networks and value networks, pre-train in a simulated environment, and optimize training using a balance strategy in combination with experience replay and exploration; Path optimization module: After the model generates the initial path, it uses the spline interpolation algorithm to smooth the path, uses the collision detection algorithm to ensure the feasibility of the path, and corrects the path; Real-time planning and updating module: used in actual scenarios to generate path planning results based on the current environmental status input model, and dynamically update the strategy according to environmental changes.

[0017] According to a third aspect of the present invention, there is provided an electronic device, comprising: a memory; and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the above-mentioned path planning method based on deep reinforcement learning is implemented.

[0018] Beneficial effects: The path planning solution based on deep reinforcement learning provided by the present invention has many significant beneficial effects compared with the existing technology.

[0019] In terms of path planning effect, this solution can generate a better path through a carefully designed reward function, taking into account factors such as path length, smoothness, feasibility, and multi-objective optimization weights. This not only shortens the path length and reduces unnecessary travel, but also ensures the smoothness of the path, effectively avoiding potential risks caused by frequent turns or mutations, and improving the practicality and safety of the path, which is of great significance in the fields of autonomous driving, robot navigation, etc.

[0020] Secondly, this solution has strong generalization capabilities. The deep reinforcement learning model can effectively learn and adapt in a variety of different environments and task scenarios without the need for large-scale retraining for each specific scenario. This enables the system to quickly adjust its strategy and generate reasonable path planning results when faced with new environments or task changes, greatly improving the versatility and flexibility of the system.

[0021] Furthermore, the experience replay mechanism and exploration and utilization balance strategy are used in the model training process to improve sample efficiency, accelerate the convergence speed of the model, and reduce the training cost. At the same time, it also performs well in real-time. Through the optimized model structure and algorithm, it can achieve fast and real-time path planning in resource-constrained environments such as mobile devices or embedded devices, meeting the low-latency requirements in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flow chart of a path planning method based on deep reinforcement learning provided in a specific embodiment of the present invention; Figure 2 It is a schematic diagram of the system composition of a path planning system based on deep reinforcement learning provided in a specific embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is clearly and completely described below in conjunction with the accompanying drawings of the present invention. Based on the embodiments in this application, other similar embodiments obtained by ordinary technicians in this field without making creative work should all fall within the scope of protection of this application. In addition, the directional words mentioned in the following embodiments, "up", "down", "left" and "right" in the preferred embodiments of the present invention are only reference directions of the accompanying drawings. Therefore, the directional words used are used to illustrate rather than limit the invention.

[0024] The present invention will be further described below in conjunction with the accompanying drawings and preferred embodiments.

[0025] See also Figure 1 The present invention provides a path planning method based on deep reinforcement learning, comprising: S100, environmental perception and state representation: Acquire environmental information through sensors, and convert the environmental information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes the current position, the target position, and obstacle distribution information in the environment.

[0026] Specifically, the sensor includes at least one of the following: LiDAR, cameras and ultrasonic sensors.

[0027] It should be noted here that the laser radar, camera and ultrasonic sensor are all installed on an intelligent body, and the intelligent body used in the present invention is an intelligent transport robot.

[0028] Furthermore, the specific method of the present invention for processing data collected by the laser radar includes: Model training steps: First, collect a large amount of lidar point cloud data in different environments and annotate them. The annotation information includes the location, size and shape of obstacles. Use the point cloud processing network in deep learning, such as PointNet or PointCNN, to train the annotated data and learn the mapping relationship between the feature representation of point cloud data and environmental information.

[0029] Related numerical settings: During data processing, the sampling frequency of the point cloud data is set to 10Hz-20Hz to ensure that environmental change information can be obtained in real time. For the filtering operation, the filter window size is set to a 5×5×5 cube to effectively remove noise points and outliers.

[0030] Related threshold setting: Set the distance threshold in the clustering algorithm. In the preferred embodiment of the present invention, the Euclidean distance threshold is 0.5 meters, which is used to cluster the point cloud data into different objects.

[0031] Specific steps of the algorithm: The original lidar point cloud data is normalized and the coordinate values ​​of the points are mapped to the [0,1] interval.

[0032] The statistical filtering method is used to remove outliers. The average distance from each point to its neighboring points is calculated. If the average distance is greater than the set threshold, the point is considered an outlier and deleted.

[0033] The clustered point cloud data is segmented using a density-based spatial clustering algorithm (DBSCAN), and the point cloud data is segmented into different objects according to a set distance threshold and a minimum number of points (10 points in the preferred embodiment of the present invention).

[0034] For each segmented object, its center of mass and bounding box features are calculated to represent the location and shape information of the obstacle.

[0035] Furthermore, the specific method of the present invention for processing the data collected by the camera includes: Model training steps: Since ultrasonic sensor data processing is more based on physical properties and simple signal processing, complex deep learning model training is usually not required, but some statistical models can be established to optimize data processing. In a preferred embodiment of the present invention, the distance measurement data of ultrasonic sensors to various common obstacles (such as walls and pillars) in different environments are collected, the distribution characteristics of measurement errors are analyzed, and an error compensation model is established.

[0036] Related numerical settings: The transmission frequency of the ultrasonic sensor is set to 40kHz-60kHz, which is a common and effective frequency range and can work well in different environments. The duration of the transmission and reception pulses is set to microseconds, 10μs-20μs in the preferred embodiment of the present invention, to ensure sufficient signal strength and resolution.

[0037] Related threshold setting: Set the confidence interval threshold of the distance measurement, which is ±0.1 meters in the preferred embodiment of the present invention. When the measured distance value is within the confidence interval with a high frequency, the measurement result is considered reliable; when it exceeds the interval by a large amount, it may be interfered or there is a measurement error, and the data needs to be corrected or re-measured.

[0038] Specific steps of the algorithm: Trigger the ultrasonic sensor to emit ultrasonic pulses and record the emission time.

[0039] Receive ultrasonic echo signals and record the echo arrival time.

[0040] The distance between the sensor and the obstacle is calculated based on the time difference between transmission and reception and the propagation speed of ultrasound in the air (about 340 m / s).

[0041] The distance values ​​obtained by multiple measurements are statistically analyzed to calculate the mean and standard deviation. If the standard deviation is less than the set threshold (0.05 meters in the preferred embodiment of the present invention), the mean is used as the final measurement result; otherwise, it is determined that there may be interference and the measurement is repeated.

[0042] Furthermore, the specific method of the present invention for processing data collected by the ultrasonic sensor includes: Model training steps: Since ultrasonic sensor data processing is more based on physical properties and simple signal processing, complex deep learning model training is usually not required, but some statistical models can be established to optimize data processing. In a preferred embodiment of the present invention, the distance measurement data of ultrasonic sensors to various common obstacles (such as walls and pillars) in different environments are collected, the distribution characteristics of measurement errors are analyzed, and an error compensation model is established.

[0043] Related numerical settings: The transmission frequency of the ultrasonic sensor is set to 40kHz-60kHz, which is a common and effective frequency range and can work well in different environments. The duration of the transmission and reception pulses is set to microseconds, 10μs-20μs in the preferred embodiment of the present invention, to ensure sufficient signal strength and resolution.

[0044] Related threshold setting: Set the confidence interval threshold of the distance measurement, which is ±0.1 meters in the preferred embodiment of the present invention. When the measured distance value is within the confidence interval with a high frequency, the measurement result is considered reliable; when it exceeds the interval by a large amount, it may be interfered or there is a measurement error, and the data needs to be corrected or re-measured.

[0045] Specific steps of the algorithm: Trigger the ultrasonic sensor to emit ultrasonic pulses and record the emission time.

[0046] Receive ultrasonic echo signals and record the echo arrival time.

[0047] The distance between the sensor and the obstacle is calculated based on the time difference between transmission and reception and the propagation speed of ultrasound in the air (about 340 m / s).

[0048] The distance values ​​obtained by multiple measurements are statistically analyzed to calculate the mean and standard deviation. If the standard deviation is less than the set threshold (0.05 meters in the preferred embodiment of the present invention), the mean is used as the final measurement result; otherwise, it is determined that there may be interference and the measurement is repeated.

[0049] It should be noted here that the present invention also adopts a multi-sensor fusion method, including: Model training steps: Use fusion learning algorithms, such as Kalman filter-based fusion models or neural network-based fusion models, to jointly train data from lidar, cameras, and ultrasonic sensors. Take sensor data and a common reference quantity (such as the position and posture of the vehicle) as input, and the fused accurate environmental information as output. Learn the association and information complementarity between different sensors through a large amount of training data, thereby optimizing the fusion effect.

[0050] Related numerical settings: In Kalman filter fusion, set the process noise covariance matrix and the measurement noise covariance matrix. The process noise covariance matrix is ​​set according to the characteristics of the sensor itself and the dynamics of the environment. In the preferred embodiment of the present invention, the corresponding covariance value will be appropriately increased for moving obstacles; the measurement noise covariance matrix is ​​set according to the measurement accuracy of each sensor. In neural network fusion, set the number of layers, number of neurons and activation function hyperparameters of the network. In the preferred embodiment of the present invention, three hidden layers are used, and the number of neurons in each layer is 64, 128 and 64 respectively, and the activation function uses the ReLU function.

[0051] Specific steps of the algorithm: The data of each sensor is preprocessed and feature extracted to obtain a unified data representation. In a preferred embodiment of the present invention, the laser radar point cloud data can be represented as a series of three-dimensional coordinate points, the camera image data is a feature vector obtained after feature extraction by a convolutional neural network, and the ultrasonic sensor data is represented by distance value and angle information.

[0052] In a preferred embodiment of the present invention, Kalman filter fusion is used to first predict the state at the current moment based on the state equation and observation equation of the system. The state equation describes the evolution of the system state over time, and the observation equation describes the relationship between the sensor's measurement value and the system state. Then, the state estimate is updated by the Kalman gain based on the difference between the actual measurement value and the predicted value.

[0053] In another preferred embodiment of the present invention, neural network fusion is adopted, and the preprocessed sensor data is used as the input of the neural network, and the fused environmental information is obtained by forward propagation calculation. During the training process, the weights and biases of the network are continuously adjusted using a back propagation algorithm and an optimizer (such as an Adam optimizer) to minimize the loss function (such as mean square error loss).

[0054] S200, action space definition: define the action space of the agent in path planning, where the action space includes moving actions in different directions and step lengths.

[0055] It is understandable that the action space may be specifically set by the user of the present invention according to actual needs, and the present invention does not limit this.

[0056] S300, Reward function design: Design a reward function that includes path length, smoothness, feasibility, and multi-objective optimization weights, where the path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment; for multi-objective optimization, assign an independent weight coefficient to each objective, and dynamically adjust the weight coefficient according to task requirements to balance the optimization relationship between different objectives.

[0057] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: The smoothness weight is dynamically adjusted according to the curvature change and the number of inflection points of the path. The greater the curvature change or the greater the number of inflection points, the higher the smoothness weight.

[0058] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: Curvature change calculation: discretize the path and divide it into a series of line segments. Calculate the angle change between adjacent line segments as the local curvature change of the path. By integrating or summing the local curvature change, the curvature change index of the entire path is obtained. The larger the curvature change index, the higher the curvature of the path.

[0059] Furthermore, the specific steps of curvature change calculation include: Model training steps: Since the curvature change is mainly based on geometric calculations, there is usually no need to train a specific model. However, the calculation method can be verified and optimized through a large amount of simulation data and actual path data. Prepare path data sets of different shapes and complexities, calculate the curvature changes of various paths, and compare them with the smoothness of manual annotations, and adjust the parameters of the calculation method to improve accuracy.

[0060] Related numerical settings: Divide the path into line segments with a length of 0.5m-1m to balance the calculation accuracy and efficiency. For calculating the local curvature change, set the number of adjacent line segments k=5, that is, consider the current line segment and the two line segments before and after it to calculate the curvature change.

[0061] Related threshold setting: Set the measurement index of curvature change, in the preferred embodiment of the present invention, the average absolute curvature change. Set a threshold T1. When the average absolute curvature change is less than T1, it is considered that the curvature change of the path is small; when it is greater than T1, it is considered that the curvature change is large. The value of T1 can be set according to the actual application scenario and experience. The present invention preferably sets it to 0.5 radians / meter.

[0062] Specific steps of the algorithm: The discretized path segments are numbered from 1 to n.

[0063] For line segment i (2<=i<=n-1), calculate the angle change θ_i between it and line segment i-1 and line segment i+1. The angle change can be calculated by vector cross product and dot product.

[0064] Calculate the curvature change index C=∑|θ_i| / k (i from i-2 to i+2), which is the average value of the sum of the angle changes of k adjacent line segments.

[0065] According to the comparison between the value of C and the threshold T1, the curvature change state of the current path segment is determined.

[0066] Calculation of the number of inflection points: The inflection points are determined by analyzing the changes in the tangent direction of each point on the path. When the angle between the tangent directions of adjacent line segments exceeds a certain threshold, the point is considered to be an inflection point. The number of inflection points in the path is counted as another indicator to measure the smoothness of the path.

[0067] Specifically, the specific method for calculating the number of inflection points includes: Model training steps: Calculation of the same curvature change, mainly verified through a large amount of simulation and actual data. Construct a path dataset containing different inflection points, use target detection or feature extraction algorithms to identify inflection points, manually verify and evaluate the accuracy of the detection results, and optimize the inflection point detection algorithm.

[0068] Related numerical settings: Set the angle threshold α of the change in the direction of the tangents of adjacent line segments to 30 degrees to determine whether it is an inflection point. During the detection process, a sliding window method is used with a window size of 3 line segments to ensure accurate identification of inflection points.

[0069] Related threshold setting: Set the threshold T2 of the number of inflection points. When the number of inflection points of the actual path exceeds T2, it is considered that the path has a large number of inflection points. The value of T2 can be set according to the specific application scenario. In the preferred embodiment of the present invention, for the scenario with high requirements for smoothness in narrow channels, T2 is preferably set to 3; for the path in the open area, T2 is preferably set to 5.

[0070] Specific steps of the algorithm: Number the path segments from 1 to n.

[0071] For line segment i (2<=i<=n-1), calculate its tangent direction vector v_i. The tangent direction vector can be obtained by the coordinate difference of two adjacent points.

[0072] Calculate the angle Δθ_i=arccos((v_i·v_{i+1}) / (|v_i|·|v_{i+1}|)) between the tangent direction vectors of adjacent line segments.

[0073] If Δθ_i>α, the line segment i is considered to be an inflection point.

[0074] Count the number of inflection points N in the path.

[0075] Smoothness weight adjustment strategy: set a baseline curvature and a baseline number of inflection points. When the curvature change of the actual path exceeds the baseline curvature or the number of inflection points exceeds the preset baseline number of inflection points, increase the smoothness weight according to the preset ratio.

[0076] Furthermore, the smoothness weight adjustment strategy specifically includes: Model training step: This step mainly determines the parameters of the adjustment strategy through experiments and verification. In different simulation scenarios and actual path data, according to the set curvature change threshold and inflection point number threshold, change the smoothness weight, observe the quality of the path planning results (such as path smoothness, feasibility indicators), and determine the best weight adjustment method through comparative analysis.

[0077] Related numerical settings: Set the basic smoothness weight w0=0.5. When the curvature change exceeds the threshold T1 or the number of inflection points exceeds the threshold T2, the increment of the smoothness weight is Δw=0.2. The upper limit of the weight adjustment is set to w_max=2.0 to prevent the weight from being too large, which will lead to over-emphasis on smoothness and neglect of other factors.

[0078] Specific steps of the algorithm: Calculate the curvature change C and the number of inflection points N of the current path.

[0079] If C>T1 or N>T2, update the smoothness weight w=min(w0+Δw,w_max); otherwise, keep w=w0.

[0080] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: The feasibility weight is dynamically adjusted according to the density of obstacles in the environment. The denser the obstacles are, the higher the feasibility weight is.

[0081] Specifically, the method for dynamically adjusting the weight coefficient in step S300 includes: Obstacle density assessment: Analyze the obstacle distribution in the environment, divide the environment into several small areas using the spatial division method, count the number of obstacles or the occupied area in each area as the obstacle density index of the area, and obtain the obstacle density assessment value of the entire environment by performing weighted averaging or summing the obstacle density indexes of all areas.

[0082] Furthermore, the specific method for evaluating the density of obstacles includes: Model training steps: mainly to establish and verify the accuracy of the obstacle density assessment model. Collect environmental data in different scenarios, mark the location and shape of obstacles, pre-process the data through clustering and statistical analysis methods, and obtain the obstacle distribution characteristics in different areas. Use multiple evaluation indicators (such as obstacle density and coverage) for training, and optimize the evaluation model by comparing the accuracy of different evaluation methods.

[0083] Related numerical settings: Divide the environment into grid cells of size 1 meter × 1 meter, and each grid cell is used as a basic evaluation unit. For each grid cell, calculate the area ratio occupied by obstacles as the obstacle density d of the cell. Set the minimum obstacle density d_min of the grid cell to 0.1, the maximum obstacle density d_max to 1, and normalize d so that its value range is between [0,1]. The normalization formula is: d_normalized=(d-d_min) / (d_max-d_min).

[0084] Related threshold setting: Set the obstacle density threshold T3. When d_normalized exceeds T3, it is considered that the obstacles in the environment are relatively dense. The value of T3 can be set according to the specific application scenario. In the preferred embodiment of the present invention, T3 can be set to 0.6 in narrow passages and 0.3 in open areas.

[0085] Specific steps of the algorithm: Obtain point cloud data or image data of the environment, and obtain the location and shape information of obstacles through target detection and segmentation algorithms.

[0086] The environment is divided into grid units, the number of obstacle pixels or points in each grid unit is counted, and the area ratio d occupied by the obstacle is calculated.

[0087] Normalize d to get d_normalized.

[0088] Feasibility weight adjustment function: Establish an adjustment function between the density of obstacles and the feasibility weight. This function is a monotonically increasing function. As the density of obstacles increases, the feasibility weight increases accordingly according to the law of the adjustment function.

[0089] Furthermore, the specific method of the feasibility weight adjustment function includes: Model training steps: Determine the appropriate weight adjustment function through experiments and data verification. Construct simulation environments with different obstacle densities, set different feasibility weight adjustment function forms (such as linear, exponential, logarithmic), run the path planning algorithm in each environment, compare the feasibility and planning efficiency of the path, and select the optimal adjustment function form and parameters.

[0090] Related numerical settings: Take the linear adjustment function as an example, the function form is w_f=a*d_normalized+b, where a and b are adjustable parameters. According to experiments and verification, when a=2.0 and b=0.5, better results can be achieved in different scenarios. That is, when d_normalized=0, w_f=0.5; when d_normalized=1, w_f=2.5.

[0091] Specific steps of the algorithm: The d_normalized value obtained according to the above steps is substituted into the preset weight adjustment function to calculate the feasibility weight w_f.

[0092] Real-time update mechanism: During the path planning process, the changes of obstacles in the environment are monitored in real time. When the position or number of obstacles changes, the obstacle density of the environment is re-evaluated and the feasibility weight is updated according to the adjustment function.

[0093] Furthermore, the design method of the real-time update mechanism includes: Model training steps: During actual operation, the effectiveness of the real-time update mechanism is verified by continuously monitoring environmental changes. Real-time data of environmental changes is collected, and the path planning effects under different update frequencies are compared to determine the optimal update frequency and strategy.

[0094] Related numerical settings: Set the period of environmental change monitoring to T=1 second, that is, re-evaluate the environment every 1 second.

[0095] Specific steps of the algorithm: During the path planning process, the environment data is re-acquired at every T time interval and the density of obstacles d_normalized is calculated.

[0096] If d_normalized changes and exceeds a certain change threshold (0.1 in the preferred embodiment of the present invention), the feasibility weight w_f is recalculated according to the adjustment function and applied to the path planning decision.

[0097] S400, Model Building and Training: Build a deep reinforcement learning model, including a policy network and a value network, pre-train in a simulated environment, and optimize training using a balance strategy combined with experience replay and exploration.

[0098] Specifically, the sampling probability of the experience replay mechanism is related to the priority of the sample, and samples with higher learning value are preferentially selected for training. The specific implementation method is as follows: Sample priority evaluation indicators: Model training steps: Through a large number of path planning experiments in simulated environments and actual scenarios, we collect learning effect data of different samples and analyze the impact of various factors (such as TD error, state characteristics, action selection and reward value) on the learning value of samples. By comparing the convergence speed and performance of the model under different evaluation indicator combinations, we determine the most suitable priority evaluation indicators and weight distribution methods.

[0099] Related numerical settings: Set the time difference error (TD error) weight w1 = 0.5, the state feature weight w2 = 0.2, the action selection weight w3 = 0.2, and the reward value weight w4 = 0.1. These weights can be adjusted according to specific application scenarios and experimental results.

[0100] Specific steps of the algorithm: 1. During the path planning process, for each sample (state-action-reward-next state), calculate its TD error TDi. The calculation formula of TD error is: TDi=ri+γV(si+1)−V(si), where ri is the reward value, γ is the discount factor (usually set to 0.9-0.99), and V(s) is the estimated value of the state value function.

[0101] 2. Extract the state feature Si of the sample, and quantize or encode the state feature. In the preferred embodiment of the present invention, a low-dimensional vector is obtained to represent the state feature after dimensionality reduction through principal component analysis (PCA).

[0102] 3. Encode the action selection Ai of the sample. In the preferred embodiment of the present invention, the selected action number can be used as a vector element.

[0103] 4. According to the weights set above, calculate the comprehensive priority of each sample Pi = w1×TDi+w2×Si+w3×Ai+w4×Ri.

[0104] Priority calculation method: Model training step: Based on the evaluation indicators and weights determined in the above steps, the calculation method is further optimized by comparing the impact of different priority calculation methods on the model training effect during the simulation training process. In a preferred embodiment of the present invention, different feature extraction methods, encoding methods and weight fusion strategies are tried to observe the performance indicators of the model on the validation set (such as loss function value, average return).

[0105] Related numerical settings: In the comprehensive priority calculation, the quantized and encoded state features Si, action selection Ai and reward value Ri must be normalized so that their value range is between [0,1], so as to perform effective weighted summation with the TD error TDi (usually also normalized).

[0106] Specific steps of the algorithm: 1. The state feature Si, action selection Ai and reward value Ri are normalized. In the preferred embodiment of the present invention, the minimum-maximum normalization method is used to map them to the [0,1] interval.

[0107] 2. According to the set weights w1, w2, w3 and w4, calculate the comprehensive priority Pi.

[0108] Sampling probability calculation and sample selection: Model training steps: Conduct multiple cycles of training experiments in a simulated environment to compare the effects of different sampling probability calculation methods and sampling ratios on model training results. Record the number of iterations required for model convergence and the final performance indicators (such as cumulative reward value and loss function value) achieved in each experiment. Analyze these data to optimize the sampling probability calculation method and select the appropriate sampling ratio to ensure that the model can learn efficiently.

[0109] Related numerical settings: Set the sampling ratio parameter α=0.1 - 0.3, which indicates the maximum proportion of the high learning value samples in each sampling. In a preferred embodiment of the present invention, when α=0.2, 100 samples are sampled from the experience pool each time, of which at most 20 samples are selected based on higher priority. In addition, a probability smoothing factor β=0.1 is set to adjust the smoothness of the sampling probability to avoid training instability caused by excessive sampling probability differences.

[0110] Specific steps of the algorithm: 1. Based on the above calculated sample comprehensive priority , calculate the sampling probability of each sample ,in It is the sum of the comprehensive priorities of all samples.

[0111] 2. Generate a random number that follows a uniform distribution for each sample .

[0112] 3. Select samples according to the following rules: First, all samples are sorted according to the sampling probability Sort from largest to smallest.

[0113] Then, select the front Samples( is the total number of samples in the experience pool, These samples are selected with a higher probability. Each of the samples ,if , then select this sample.

[0114] For the remaining samples are selected with uniform probability, that is, the probability of each sample being selected is . But when choosing, the probability smoothing factor is introduced , adjust the sampling probability to (for high probability samples that are not preferred) and (for the remaining samples) to avoid large differences in sampling probabilities.

[0115] S500, path optimization: After the model generates the initial path, the spline interpolation algorithm is used to smooth the path, the collision detection algorithm is used to ensure the feasibility of the path, and the path is corrected.

[0116] Specifically, in step S500, the spline interpolation algorithm uses a cubic spline interpolation method to generate a smooth path curve by performing interpolation calculations on key points on the path.

[0117] Furthermore, the method of generating a smooth path curve by performing interpolation calculation on key points on the path includes: Key point selection: Model training steps: Collect a large amount of path data of different types and complexity, including manually planned ideal paths and path data in actual environments. Label these path data, mark the locations of key points and some information describing the characteristics of key points (such as points with large curvature changes and inflection points). Use machine learning algorithms (such as clustering algorithms and feature-based classification algorithms) to verify and optimize the key point selection method, and determine the most suitable key point selection strategy by comparing the accuracy of different selection methods and their impact on path smoothing.

[0118] Related numerical settings: For the selected key points, set their distribution density on the path. In a preferred embodiment of the present invention, in a relatively flat straight line segment of the path, the spacing between key points can be set to 10%-20% of the path length; in a curved segment, especially in an area with a large curvature change, the key point spacing is set to 5%-10% of the path length. This ensures that the key points can effectively capture the shape characteristics of the path, while avoiding excessive calculations caused by too many key points.

[0119] A curvature change threshold Tcurvature is set to identify the key feature points of the curve segment. In a preferred embodiment of the present invention, when the curvature change of adjacent line segments exceeds Tcurvature (such as 0.2 radians / meter), the point is marked as a key point.

[0120] Specific steps of the algorithm: The preprocessed path is preliminarily segmented into several small segments. In a preferred embodiment of the present invention, the segmentation is performed according to the direction change or distance interval (such as every 1 meter) of the path.

[0121] Calculate the curvature change of each segment, which is determined by calculating the difference in curvature between adjacent segments. For each line segment i and i+1, calculate their curvatures ki and ki+1, and the curvature change Δki=|ki+1-ki|.

[0122] The key point is determined based on the curvature change threshold Tcurvature. If Δki>Tcurvature, the end point of segment i or the start point of segment i+1 is marked as a key point.

[0123] The starting point, end point, and location where the direction changes significantly (such as an angle change of more than 30 degrees) of the path are marked as key points.

[0124] Furthermore, the interpolation calculation and curve generation method includes: Related value settings: In actual calculations, in order to ensure the accuracy of interpolation, the maximum number of iterations of numerical calculations is set. In a preferred embodiment of the present invention If the convergence condition is still not met after reaching the maximum number of iterations (the difference between the results of two consecutive iterations is less than ), it is believed that there may be a problem with the interpolation calculation, and it is necessary to check the input data or adjust the algorithm parameters.

[0125] Set a minimum subinterval length In a preferred embodiment of the present invention meters. When the distance between adjacent key points is less than , the two points are considered too close, which may have an adverse effect on the interpolation calculation. In this case, you can consider merging the two points or taking other special processing methods.

[0126] Specific steps of the algorithm: For any given value, first determine the subinterval it is in , so that .

[0127] According to the coefficients calculated previously , substituting into the cubic polynomial , calculate the corresponding Value, that is, the point on the interpolated path curve .

[0128] Repeat steps 1 and 2 for all the paths that need to be interpolated. Value (the sampling interval can be set according to actual needs. In the preferred embodiment of the present invention, meters), and calculate the corresponding value to obtain a complete smooth path curve.

[0129] Specifically, in step S500, the collision detection algorithm adopts a collision detection method based on a distance field, and determines whether there is a collision risk by calculating the distance between each point on the path and the obstacle.

[0130] Specifically, the specific method of constructing the distance field includes: Distance Field Construction: Model training steps: Collect a large amount of data of different types and environments, including information about obstacles of various shapes and distributions. Use numerical simulation methods to generate distance field data in different scenarios, and compare and verify with the collision detection results in the real environment. Improve the accuracy and efficiency of the distance field by adjusting the parameters of the distance field construction algorithm and optimizing the calculation method.

[0131] Set a resolution parameter R of the distance field. In the preferred embodiment of the present invention, R=0.1 meter. This parameter determines the spacing between each discrete point in the distance field. A smaller resolution can improve the accuracy of the distance field, but it will increase the amount of calculation and storage space. Select a suitable resolution based on the actual application scenario and computing resources.

[0132] Related value settings: For the representation of obstacles, a bounding box or grid-based method is used. If a bounding box is used, the minimum size limit S_min of the bounding box is set. In the preferred embodiment of the present invention, S_min=0.5m×0.5m to ensure that smaller obstacles can be accurately represented. If a grid-based method is used, the size of the grid is set to G. In the preferred embodiment of the present invention, G=0.2m×0.2m to divide the environment into uniform grid units.

[0133] When calculating the distance field, a maximum distance threshold D_max is set. In the preferred embodiment of the present invention, D_max=10 meters. When the distance between the calculation point and the obstacle exceeds D_max, it is considered that there is no collision risk between the point and the obstacle, which can reduce unnecessary calculation.

[0134] Specific steps of the algorithm: For each obstacle, the space it occupies in the environment is calculated based on its representation (bounding box or meshing).

[0135] Taking a reference point in the environment (the origin or a fixed point in the preferred embodiment of the present invention) as the starting point, discrete sampling is performed in the environment according to a set resolution R to obtain a series of sampling points.

[0136] For each sampling point p, calculate its distance to all obstacles. If the obstacle is represented by a bounding box, the shortest distance calculation method from the point to the bounding box can be used; if the grid method is used, the distance can be calculated by traversing the grid cells where the obstacle is located.

[0137] Store the calculated distance information in a data structure to form a distance field. A three-dimensional array or other efficient data structure can be used to store the distance field data for subsequent queries and calculations.

[0138] Specifically, the specific method for calculating the distance of path points includes: Model training steps: Generate a large number of test scenarios with different path and obstacle distributions in a simulation environment to test and verify the path point distance calculation method. By comparing the distance values calculated by the algorithm with the true distance values (which can be obtained through precise geometric calculation methods), analyze the error sources and accuracy of the algorithm. According to the experimental results, optimize and improve the algorithm to improve the accuracy of distance calculation.

[0139] Set an error tolerance εd for distance calculation. In a preferred embodiment of the present invention, εd = 0.05 meters. During the algorithm training process, by adjusting the calculation parameters and optimizing the algorithm logic, control the calculation error of the algorithm within εd.

[0140] Related numerical settings: For a point p_path on the path, set a search radius r_search. In a preferred embodiment of the present invention, r_search = 2 meters. When calculating the distance, only consider the obstacles within the range of r_search from p_path to reduce the calculation amount. The selection of this search radius needs to be adjusted according to the actual environment and path characteristics, ensuring that possible collision obstacles can be detected while avoiding excessive invalid calculations.

[0141] When the distance d between the calculation point p_path and the obstacle satisfies d < r_search, it is considered that the obstacle has a potential collision risk for the path point p_path. At this time, further accurately calculate the value of d and compare it with other thresholds.

[0142] Specific algorithm steps: For each point p_path on the path, first determine its search range, that is, a spherical region (a circular region in two-dimensional cases) centered at p_path with a radius of r_search.

[0143] Query the distance information of the obstacles within the search range in the distance field. A spatial index structure (such as an octree, KD-tree) can be used to accelerate the query process and improve the calculation efficiency.

[0144] For each obstacle distance d queried, determine whether it satisfies d < D_max (the maximum distance threshold). If it satisfies, continue to the next step of calculation; otherwise, it is considered that the obstacle has no collision risk for the path point p_path, and skip the subsequent calculation of this obstacle.

[0145] Calculate the actual distance d_real from the path point p_path to the obstacle. This can be calculated based on the discrete data points in the distance field through interpolation methods such as bilinear interpolation and trilinear interpolation.

[0146] Compare d_real with the set collision threshold d_threshold (in the preferred embodiment of the present invention, d_threshold = 0.2 meters). If d_real < d_threshold, it is considered that there is a collision risk between the path point p_path and the obstacle; otherwise, this point is considered safe.

[0147] Specifically, the specific method for collision risk assessment and handling includes: Collision risk assessment and handling: Model training steps: Construct a data set containing different collision risk scenarios, including different degrees of collision risk situations (such as minor collision risk, critical collision risk, and severe collision risk). Use machine learning algorithms (such as support vector machines and decision trees) to train the data set to learn the evaluation features and patterns of collision risk. Optimize the model parameters through cross-validation and model evaluation metrics (such as accuracy and recall) to improve the accuracy of collision risk assessment.

[0148] Set a confidence threshold Cthreshold for risk assessment. In the preferred embodiment of the present invention, Cthreshold = 0.8. When the confidence of the evaluation result is lower than this threshold, the evaluation result is considered unreliable and further analysis or other measures need to be taken.

[0149] Relevant numerical settings: For the collision risk assessment of multiple consecutive path points, set a consecutive risk point threshold Nrisk. In the preferred embodiment of the present invention, Nrisk = 3. When Nrisk consecutive path points are all evaluated as having a collision risk, it is considered that there is a relatively large collision risk on the path.

[0150] When a collision risk is detected, set a risk handling priority index Prisk. In the preferred embodiment of the present invention, according to the degree of collision risk (such as the distance from the obstacle and the degree of damage that the collision may cause), a priority value is assigned to each risk point. The higher the priority value, the more serious the risk and the need for priority handling.

[0151] Specific steps of the algorithm: Perform a collision risk assessment on each point on the path. According to the distance information and threshold comparison results calculated in the above steps, determine whether each point has a collision risk and record the confidence C of the risk assessment.

[0152] If C < Cthreshold, further analyze and verify this point. In the preferred embodiment of the present invention, recalculate the distance using a more accurate geometric calculation method, or consider more environmental factors (such as the motion state of obstacles).

[0153] Check whether there are Nrisk consecutive path points evaluated as having a collision risk. If so, it is considered that the path has a relatively high collision risk and corresponding handling measures need to be taken.

[0154] Sort the path points with collision risks according to the risk handling priority index Prisk, and give priority to handling the risk points with higher priorities.

[0155] For the path points with collision risks, various handling measures can be taken, such as adjusting the path (avoiding obstacles by re-planning or locally adjusting the path), issuing an alarm (prompting the user or operator to pay attention to the collision risk). The specific handling method can be selected and designed according to the actual application scenario and requirements.

[0156] S600, Real-time Planning and Update: In the actual scenario, generate a path planning result based on the input model of the current environmental state, and dynamically update the strategy according to environmental changes.

[0157] Specifically, when the environmental change exceeds the preset threshold, trigger the model to be retrained or online adjust the strategy parameters to adapt to the new environment. The specific implementation method is as follows: Model Training Steps: Simulate various environmental change situations in the simulation environment, including the movement, addition or disappearance of obstacles, and the change of the target position. Test and verify the monitoring algorithm under different environmental change scenarios, and analyze the detection ability of the algorithm for different types and degrees of environmental changes. According to the experimental results, adjust the parameters and thresholds of the monitoring algorithm to improve the accuracy and reliability of environmental change monitoring.

[0158] Set a sampling period Tsmaple for environmental change monitoring. In the preferred embodiment of the present invention, Tsmaple = 1 second. This period determines the detection frequency of the monitoring algorithm for environmental changes and needs to be reasonably selected according to the actual application scenario and computing resources.

[0159] Related Numerical Settings: For the monitoring of the position change of obstacles, set a position change threshold Tpos. In the preferred embodiment of the present invention, Tpos = 0.5 meters. When the position change of the obstacle exceeds this threshold, it is considered that the position of the obstacle has changed significantly.

[0160] For the monitoring of the shape change of obstacles, a shape similarity measurement method (such as shape context, Hausdorff distance) is adopted. A shape similarity threshold Tshape is set. In the preferred embodiment of the present invention, Tshape = 0.8 (the value range is from 0 to 1, and 1 represents complete similarity). When the shape similarity of the obstacle is lower than this threshold, it is considered that the shape of the obstacle has changed significantly.

[0161] For the monitoring of the change of the target position, a target position change threshold Ttarget is set. In the preferred embodiment of the present invention, Ttarget = 1 meter. When the change of the target position exceeds this threshold, it is considered that the target position has changed significantly.

[0162] Specific steps of the algorithm: Sample the environment according to the set sampling period Tsmaple, and obtain the environmental information at the current moment, including the position, shape of the obstacle and the target position.

[0163] Compare and analyze the environmental information at the current moment with the environmental information at the previous moment.

[0164] For each obstacle, calculate its position change amount Δp and shape change amount Δs. The position change amount can be obtained by calculating the coordinate difference of the center position of the obstacle, and the shape change amount can be calculated by the shape similarity measurement method. If Δp > Tpos or Δs < Tshape, it is considered that the state of the obstacle has changed significantly.

[0165] Calculate the change amount Δt of the target position. If Δt > Ttarget, it is considered that the target position has changed significantly.

[0166] Comprehensively consider the change situations of all obstacles and the target position, and judge whether the environmental change exceeds the preset threshold. If there is at least one obstacle or the change of the target position that exceeds the preset threshold, it is considered that the environment has changed significantly.

[0167] Preset threshold setting: Model training steps: Through a large number of experiments and data analysis, study the influence of different environmental changes on the performance of the path planning algorithm. According to the experimental results, determine the influence degree of different types of environmental changes on path planning, and set reasonable preset thresholds based on this. In the preferred embodiment of the present invention, in some application scenarios with high requirements for path accuracy, the position change threshold and shape similarity threshold can be appropriately reduced to improve the sensitivity to environmental changes.

[0168] A dynamic range Rthreshold for threshold adjustment is set. In a preferred embodiment of the present invention, Rthreshold=[0.5,1.5]. This range indicates that the preset threshold can be tested in different simulation environments and actual scenarios to collect data on environmental changes and path planning effects. By analyzing these data, the relationship between the preset threshold and the performance of the path planning algorithm is studied. In a preferred embodiment of the present invention, when the threshold is too low, the model may be retrained or parameters may be adjusted too frequently, increasing the computational cost; and when the threshold is too high, some situations that need to be adjusted may be missed, affecting the accuracy of path planning. Based on these analysis results, the preset threshold is further optimized within the dynamic range Rthreshold to balance the computational cost and path planning performance.

[0169] Considering the dynamic characteristics of the environment and the task requirements, an adaptive adjustment mechanism is introduced to determine the preset threshold. In a preferred embodiment of the present invention, when the environment changes frequently or the task has high real-time requirements for path planning, the threshold can be appropriately relaxed to reduce unnecessary model retraining or parameter adjustment; when the environment is relatively stable or the task has strict requirements for the accuracy of path planning, the threshold can be tightened to ensure the accuracy of path planning.

[0170] Furthermore, the method for adjusting the preset threshold includes: Model training steps: In the actual path planning system, a real-time monitoring mechanism is established to obtain the latest information on environmental changes in real time and feed it back to the threshold judgment and collaborative control modules.

[0171] Design a dynamic adjustment strategy to dynamically adjust the threshold and strategy parameters according to the real-time monitored environmental change information and path planning effect. In a preferred embodiment of the present invention, when it is detected that the speed of environmental changes is accelerating or the complexity is increasing, the threshold is appropriately relaxed so that it can respond to environmental changes more promptly; when the environmental changes are gradually slowing down and the path planning effect is stable, the threshold can be appropriately tightened to improve the accuracy and efficiency of path planning.

[0172] Related numerical settings: Set an environmental change speed index v_change to measure the speed of environmental change. v_change can be obtained by calculating the ratio of the amplitude of environmental change at two consecutive moments to the time interval, that is, ,in is the difference between the environmental change complexity indexes at two adjacent moments, and Δt is the time interval. A speed threshold v_threshold is set. In the preferred embodiment of the present invention, ,When v_change>v_threshold, it means that the environment changes quickly and the threshold needs to be relaxed.

[0173] Set a path planning effect stability index E_stable to determine whether the path planning effect is stable. E_stable can be measured by calculating the standard deviation of the path planning effect index over a period of time. If the standard deviation is less than a certain threshold (such as 0.05), the path planning effect is considered to be relatively stable and the threshold can be appropriately tightened.

[0174] Specific steps of the algorithm: Real-time calculation of the environment change speed index v_change and the path planning effect stability index E_stable.

[0175] Compare v_change to the speed threshold v_threshold and compare E_stable to the stability threshold: If v_change>v_threshold or E_stable does not meet the stability condition (i.e., the standard deviation is greater than 0.05), the threshold is appropriately relaxed.

[0176] If v_change≤v_threshold and E_stable satisfies the stability condition, tighten the threshold appropriately.

[0177] For the amplitude of threshold adjustment, the present invention further refines the setting of relevant adjustment factors. Setting an amplification factor and reduction factor ,When the threshold needs to be relaxed, the new threshold ; When the threshold needs to be tightened, the new threshold According to experiments and analysis, we set The value range is [0.1,0.3], The value range of is [0.05, 0.2]. Specifically, each time the adjustment is made, the appropriate factor value can be selected for adjustment according to the text content, environmental changes and real-time feedback of the path planning effect.

[0178] Specifically, set an adaptive adjustment cycle In a preferred embodiment of the present invention seconds. That is, every time, recalculate the environmental change speed index and path planning effect stability index, and adjust the threshold once according to the above comparison results. Adaptive adjustment cycle The setting should take into account both real-time and stability. If the cycle is too short, it may lead to frequent adjustment of the threshold, increase computing overhead and may cause over-adjustment due to short-term fluctuations in environmental changes; if the cycle is too long, it may not be able to respond to rapid changes in the environment in a timely manner.

[0179] See also Figure 2The present invention provides another embodiment, which provides a path planning system based on deep reinforcement learning, and the path planning system based on deep reinforcement learning includes: Environment perception and state representation module 100: used to obtain environment information through sensors and convert the environment information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes the current position, the target position, and the obstacle distribution information in the environment; Action space definition module 200: used to define the action space of the agent in path planning, the action space includes movement actions of different directions and step lengths; Reward function design module 300: used to design a reward function including path length, smoothness, feasibility and multi-objective optimization weights, wherein the path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment; for multi-objective optimization, an independent weight coefficient is assigned to each objective, and the weight coefficient is dynamically adjusted according to task requirements to balance the optimization relationship between different objectives; Model building and training module 400: used to build a deep reinforcement learning model, including a policy network and a value network, pre-train in a simulation environment, and optimize the training by combining experience replay and exploration using a balance strategy; Path optimization module 500: used to smooth the path using a spline interpolation algorithm, ensure path feasibility using a collision detection algorithm, and correct the path after the model generates an initial path; Real-time planning and updating module 600: used to generate path planning results according to the current environment state input model in actual scenarios, and dynamically update the strategy according to environmental changes.

[0180] In a preferred embodiment, the present application further provides an electronic device, the electronic device comprising: A memory; and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the path planning method based on deep reinforcement learning is implemented. The computer device can be broadly a server, a terminal, or any other electronic device with necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, a memory, a network interface, and a communication interface connected via a system bus. The processor of the computer device can be used to provide necessary computing, processing and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and an internal memory. An operating system and a computer program may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the computer device can be used to connect and communicate with external devices through a network. When the computer program is executed by the processor, the steps of the method of the present invention are performed.

[0181] The present invention may be implemented as a computer-readable storage medium having a computer program stored thereon, which causes the steps of the method of an embodiment of the present invention to be executed when executed by a processor. In one embodiment, the computer program is distributed on a plurality of computer devices or processors coupled to a network so that the computer program is stored, accessed and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be performed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be performed by one or more computer devices or processors, and one or more other method steps / operations may be performed by one or more other computer devices or processors. One or more computer devices or processors may perform a single method step / operation, or perform two or more method steps / operations.

[0182] It will be appreciated by a person skilled in the art that the method steps of the present invention may be performed by instructing related hardware such as a computer device or a processor through a computer program, and the computer program may be stored in a non-temporary computer-readable storage medium, which causes the steps of the present invention to be performed when the computer program is executed. Depending on the circumstances, any reference to memory, storage, database, or other media herein may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk. Examples of volatile memory include random access memory (RAM), external cache memory.

[0183] The various technical features described above can be combined arbitrarily. Although all possible combinations of these technical features are not described, any combination of these technical features should be considered to be covered by this specification as long as there is no contradiction in such combination.

[0184] The specific implementation of the present invention described above does not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and modifications made based on the technical concept of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A path planning method based on deep reinforcement learning, characterized in that: The method comprises: S100, environmental perception and state representation: acquiring environmental information through sensors, and converting the environmental information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes current position, target position, and obstacle distribution information in the environment; S200, action space definition: defining the action space of the agent in path planning, wherein the action space includes movement actions of different directions and step lengths; S300, Reward function design: Design a reward function that includes path length, smoothness, feasibility, and multi-objective optimization weights, where the path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment; for multi-objective optimization, assign an independent weight coefficient to each objective, and dynamically adjust the weight coefficient according to task requirements to balance the optimization relationship between different objectives; S400, Model Building and Training: Build a deep reinforcement learning model, including a policy network and a value network, pre-train in a simulated environment, and optimize training using a balance strategy in combination with experience replay and exploration; S500, path optimization: after the model generates the initial path, the spline interpolation algorithm is used to smooth the path, the collision detection algorithm is used to ensure the feasibility of the path, and the path is corrected; S600, real-time planning and updating: In actual scenarios, the path planning results are generated based on the current environmental status input model, and the strategy is dynamically updated according to environmental changes.

2. The path planning method based on deep reinforcement learning according to claim 1, characterized in that: The sensor includes at least one of the following: LiDAR, cameras and ultrasonic sensors.

3. The path planning method based on deep reinforcement learning according to claim 1, characterized in that: The method for dynamically adjusting the weight coefficient in step S300 includes: The smoothness weight is dynamically adjusted according to the curvature change and the number of inflection points of the path. The greater the curvature change or the greater the number of inflection points, the higher the smoothness weight.

4. The path planning method based on deep reinforcement learning according to claim 3, characterized in that: The method for dynamically adjusting the weight coefficient in step S300 includes: Curvature change calculation: discretize the path and divide it into a series of line segments. Calculate the angle change between adjacent line segments as the local curvature change of the path. By integrating or summing the local curvature change, the curvature change index of the entire path is obtained. The larger the curvature change index, the higher the curvature degree of the path. Calculation of the number of inflection points: The inflection points are determined by analyzing the changes in the tangent direction of each point on the path. When the angle between the tangent directions of adjacent line segments exceeds a certain threshold, the point is considered to be an inflection point. The number of inflection points in the path is counted as another indicator to measure the smoothness of the path. Smoothness weight adjustment strategy: set a baseline curvature and a baseline number of inflection points. When the curvature change of the actual path exceeds the baseline curvature or the number of inflection points exceeds the preset baseline number of inflection points, increase the smoothness weight according to the preset ratio.

5. The path planning method based on deep reinforcement learning according to claim 1, characterized in that: The method for dynamically adjusting the weight coefficient in step S300 includes: The feasibility weight is dynamically adjusted according to the density of obstacles in the environment. The denser the obstacles are, the higher the feasibility weight is.

6. The path planning method based on deep reinforcement learning according to claim 5, characterized in that: The method for dynamically adjusting the weight coefficient in step S300 includes: Obstacle density assessment: Analyze the obstacle distribution in the environment, divide the environment into several small areas using the spatial division method, count the number of obstacles or the occupied area in each area as the obstacle density index of the area, and obtain the obstacle density assessment value of the entire environment by performing weighted average or summation operations on the obstacle density index of all areas; Feasibility weight adjustment function: Establish an adjustment function between the density of obstacles and the feasibility weight. This function is a monotonically increasing function. As the density of obstacles increases, the feasibility weight increases accordingly according to the law of the adjustment function. Real-time update mechanism: During the path planning process, the changes of obstacles in the environment are monitored in real time. When the position or number of obstacles changes, the obstacle density of the environment is re-evaluated and the feasibility weight is updated according to the adjustment function.

7. The path planning method based on deep reinforcement learning according to claim 1, characterized in that: In step S500, the spline interpolation algorithm uses a cubic spline interpolation method to generate a smooth path curve by performing interpolation calculations on key points on the path.

8. The path planning method based on deep reinforcement learning according to claim 1, characterized in that: In step S500, the collision detection algorithm uses a distance field-based collision detection method to determine whether there is a collision risk by calculating the distance between each point on the path and the obstacle.

9. A path planning system based on deep reinforcement learning, characterized in that: include: Environmental perception and state representation module: used to obtain environmental information through sensors and convert the environmental information into a state representation suitable for a deep reinforcement learning model, wherein the state representation includes the current position, the target position, and the obstacle distribution information in the environment; Action space definition module: used to define the action space of the agent in path planning, which includes movement actions of different directions and step lengths; Reward function design module: used to design a reward function that includes path length, smoothness, feasibility, and multi-objective optimization weights. The path length weight is used to encourage the agent to find a shorter path, the smoothness weight is used to punish turning points and mutations in the path, and the feasibility weight is used to ensure that the generated path does not conflict with obstacles in the environment. For multi-objective optimization, an independent weight coefficient is assigned to each objective, and the weight coefficient is dynamically adjusted according to task requirements to balance the optimization relationship between different objectives. Model building and training module: used to build deep reinforcement learning models, including policy networks and value networks, pre-train in a simulated environment, and optimize training using a balance strategy in combination with experience replay and exploration; Path optimization module: After the model generates the initial path, it uses the spline interpolation algorithm to smooth the path, uses the collision detection algorithm to ensure the feasibility of the path, and corrects the path; Real-time planning and updating module: used in actual scenarios to generate path planning results based on the current environmental status input model, and dynamically update the strategy according to environmental changes.

10. An electronic device, characterized in that: include: Memory; and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the path planning method based on deep reinforcement learning according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Mobile robot intelligent path planning method

    CN112631294A

  • Deep-sea mining robot path planning method based on deep reinforcement learning

    CN116339316A

  • Unmanned vehicle adaptive path planning method based on dynamic window method and near-end strategy

    CN116679719A

  • Mobile robot local path planning method based on value distribution deep reinforcement learning

    CN117470244A

  • Obstacle avoidance path planning method fusing artificial potential field method and D*Lite

    CN117490711A

Cited By

  • Emergency obstacle avoidance method and device of improved dynamic window method based on particle filtering

    CN120178936A

  • Multi-unmanned aerial vehicle autonomous cooperative flight path planning method based on improved FGO algorithm

    CN120406564A

  • A multi-UAV autonomous collaborative trajectory planning method based on improved FGO algorithm

    CN120406564B

  • Autonomous navigation path planning method and device for humanoid robot in complex environment

    CN120447563A

  • A method and device for autonomous navigation path planning of a humanoid robot in a complex environment

    CN120447563B