An obstacle avoidance path planning method and related equipment for a construction robot

Environmental features are extracted through Bayesian filtering and lightweight convolutional neural networks, and combined with decision-making models to generate action instructions, it solves the problem that building robots are difficult to efficiently avoid obstacles and path planning in complex dynamic environments, and realizes efficient and real-time path planning.

CN119665987BActive Publication Date: 2025-05-27湖南工商大学

Patent Information

Application Number
CN202510196212.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing construction robots are difficult to achieve efficient obstacle avoidance and path planning in complex dynamic environments, existing path planning algorithms are difficult to adapt to changes in dynamic environments, and traditional obstacle avoidance algorithms are prone to local minimum value problems, causing robots to fall into trouble.

Method used

By obtaining the environmental data of the target space, Bayesian filtering is used to fuse the status information of the building robot with the environmental data to generate an occupied raster map. Then, the lightweight convolutional neural network is input to the occupancy raster map for feature extraction to generate environmental feature data. Enter the environmental feature data with the current status of the building robot to make decisions, generate action instructions, and realize efficient obstacle avoidance and path planning for building robots in complex environments.

Benefits of technology

It improves the accuracy and robustness of environment perception, reduces the computational complexity, improves decision-making speed, meets real-time requirements, and realizes efficient obstacle avoidance and path planning in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119665987B_ABST
    Figure CN119665987B_ABST
Patent Text Reader

Abstract

The present invention provides an obstacle avoidance path planning method and related devices for a construction robot. By using Bayesian filtering to fuse the state information of the target construction robot with environmental data, an occupancy grid map is obtained; the occupancy grid map is input into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data; the environmental feature data and the current state of the target construction robot are input into a trained decision-making model for decision-making to obtain an action instruction. The target construction robot reaches the target position in the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot; by improving the accuracy and robustness of environmental perception, efficient obstacle avoidance and path planning in a complex environment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an obstacle avoidance path planning method and related equipment for a construction robot. Background Art

[0002] With the rapid development of the construction industry, the demand for automation and intelligence at the construction site is increasing day by day. Construction robots are increasingly widely used in material distribution, inspection, cleaning, etc. at the construction site. However, the construction site environment is complex and changeable, with a large number of dynamic and static obstacles, such as construction workers, mechanical equipment, material stacking, etc. At the same time, the construction site environment usually has problems such as insufficient light, a lot of dust, and many obstacles, resulting in it being difficult for construction robots based on visual sensors (such as cameras) to reliably perceive the environment.

[0003] Traditional path planning algorithms, such as Algorithm, Dijkstra algorithm, are mainly based on known static maps and are difficult to adapt to the changes in dynamic environments. In addition, these algorithms mainly aim at the shortest path length and do not fully consider the total time taken to complete the task. Traditional obstacle avoidance algorithms, such as the artificial potential field method, may have the problem of "local minimum", causing the robot to get stuck, and are prone to oscillation in complex environments, affecting the task efficiency.

[0004] Although the path planning method based on deep reinforcement learning has a certain degree of environmental adaptability, it usually does not specifically optimize the task completion time, which may lead to a long path or a large time consumption. The existing methods still have deficiencies in real-time performance, obstacle avoidance performance, and task efficiency, and cannot meet the needs of construction robots to complete tasks in the shortest time in complex dynamic environments. Summary of the Invention

[0005] The present invention provides an obstacle avoidance path planning method and related equipment for a construction robot, and its purpose is to achieve efficient obstacle avoidance and path planning in a complex environment.

[0006] To achieve the above object, the present invention provides an obstacle avoidance path planning method for a construction robot, including:

[0007] Step 1, obtaining environmental data of the target space;

[0008] Step 2, fusing the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map;

[0009] Step 3, inputting the occupancy grid map into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data;

[0010] Step 4: Input the environmental feature data and the current state of the target construction robot into the trained decision-making model for decision-making to obtain an action instruction. The target construction robot reaches the target position within the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot.

[0011] Furthermore, Step 1 includes:

[0012] Install a lidar and multiple ultrasonic sensors inside the target construction robot;

[0013] Obtain the three-dimensional point cloud data of the target space through the lidar;

[0014] Obtain the distance data between the target construction robot and the obstacles through multiple ultrasonic sensors;

[0015] Integrate the three-dimensional point cloud data and the distance data into the environmental data of the target space.

[0016] Furthermore, before fusing the state information of the target construction robot with the environmental data through Bayesian filtering, it also includes:

[0017] Preprocess the environmental data to obtain the preprocessed environmental data;

[0018] Fuse the state information of the target construction robot with the preprocessed environmental data through Bayesian filtering data.

[0019] Furthermore, Step 2 includes:

[0020] Obtain the position, movement speed, and action instruction of the target construction robot at the previous moment;

[0021] Based on the position, movement speed, and action instruction of the target construction robot at the previous moment, use the kinematic model to generate a prior state estimate, and the prior state estimate includes the current position, speed, and steering information of the target construction robot;

[0022] Fuse the preprocessed environmental data with the prior state estimate to obtain a posterior state estimate;

[0023] Map the posterior state estimate onto a two-dimensional grid map to obtain an occupancy grid map.

[0024] Furthermore, the calculation expression of the posterior state estimate is:

[0025]

[0026]

[0027]

[0028] Among them, represents the posterior state estimate, represents the prior state estimate, represents the Kalman gain, represents the time of the observation value, represents the observation matrix, represents the transpose matrix of represents the observation noise covariance matrix, represents the identity matrix, represents the posterior state estimate covariance matrix, represents the prior state covariance matrix.

[0029] Furthermore, the trained lightweight convolutional neural network includes:

[0030] An input layer, a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a fully connected layer, and an output layer connected in sequence;

[0031] The input layer is used to receive the occupancy grid map;

[0032] Both the first convolutional layer and the second convolutional layer are used to extract features from the occupancy grid map;

[0033] Both the first pooling layer and the second pooling layer are used to downsample the extracted features;

[0034] The fully connected layer is used to flatten the features output by the second pooling layer to obtain an environmental feature vector;

[0035] The output layer generates environmental feature data based on the environmental feature vector.

[0036] Furthermore, the trained decision-making model includes a policy network and a value network;

[0037] The policy network generates an action probability distribution based on the environmental feature data and the current state of the target construction robot;

[0038] The value network generates a state value estimate based on the environmental feature data and the current state of the target construction robot.

[0039] The present invention also provides an obstacle avoidance path planning device for a construction robot, including:

[0040] An acquisition module, configured to acquire environmental data of the target space;

[0041] A fusion module, configured to fuse the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map;

[0042] An extraction module, configured to input the occupancy grid map into the trained lightweight convolutional neural network for feature extraction to generate environmental feature data;

[0043] A decision-making module, configured to input the environmental feature data and the current state of the target construction robot into the trained decision-making model for decision-making to obtain an action instruction. The target construction robot reaches the target position in the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot.

[0044] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the obstacle avoidance path planning method of the construction robot is implemented.

[0045] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the obstacle avoidance path planning method of the construction robot is implemented.

[0046] The above solution of the present invention has the following beneficial effects:

[0047] The present invention obtains the environmental data of the target space; fuses the state information of the target construction robot and the environmental data through Bayesian filtering to obtain an occupancy grid map; inputs the occupancy grid map into the trained lightweight convolutional neural network for feature extraction to generate environmental feature data; inputs the environmental feature data and the current state of the target construction robot into the trained decision-making model for decision-making to obtain an action instruction. The target construction robot reaches the target position in the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot. Compared with the prior art, the present invention first uses Bayesian filtering to perform data fusion on the state information of the target construction robot and the environmental data, improving the accuracy and robustness of environmental perception. The lightweight convolutional neural network and the optimized reinforcement learning network are used to extract the features of the data and generate action instructions, reducing the computational complexity and improving the decision-making speed, meeting the real-time requirement. Finally, based on the action instruction, efficient obstacle avoidance and path planning in a complex environment are achieved.

[0048] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic flowchart of an embodiment of the present invention;

[0050] Figure 2 It is a schematic structural diagram of the obstacle avoidance path planning device in an embodiment of the present invention;

[0051] Figure 3 The structural schematic diagram of the terminal device provided by the embodiment of the present application. Detailed implementation manners

[0052] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0054] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0055] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0056] The target space mentioned in the embodiment of the present invention is an actual construction site environment, and the target position mentioned in the embodiment of the present invention is the construction point where the materials are transported. The construction site environment is complex, with dynamic obstacles (such as construction workers and vehicles with uncertain trajectories) and static obstacles (such as material piles, equipment, waste piles, etc.). The purpose of the obstacle avoidance path planning method for construction robots provided by the embodiment of the present invention is to plan a path that can safely avoid obstacles and complete the material distribution task in the shortest time.

[0057] The present invention provides an obstacle avoidance path planning method for construction robots and related devices for the above tasks.

[0058] As Figure 1 shown, the embodiment of the present invention provides an obstacle avoidance path planning method for construction robots, including:

[0059] Step 1, obtaining the environmental data of the target space;

[0060] Step 2: Integrate the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map;

[0061] Step 3: Input the occupancy grid map into the trained lightweight convolutional neural network for feature extraction to generate environmental feature data;

[0062] Step 4: Input the environmental feature data and the current state of the target construction robot into the trained decision-making model for decision-making to obtain an action instruction. The target construction robot reaches the target position within the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot.

[0063] Most preferably, Step 1 includes:

[0064] Install a lidar and multiple ultrasonic sensors inside the target construction robot;

[0065] Obtain the three-dimensional point cloud data of the target space through the lidar;

[0066] Obtain the distance data between the target construction robot and the obstacles through multiple ultrasonic sensors;

[0067] Integrate the three-dimensional point cloud data and the distance data into the environmental data of the target space.

[0068] In the embodiment of the present invention, the lidar (LiDAR) equipped on the target construction robot is a 16-line lidar, with the model number of Velodyne VLP-16. The horizontal scanning range is 360°, the vertical scanning range is 30°, the ranging range is from 0.2 meters to 100 meters, the horizontal angular resolution is 0.2°, and the vertical angular resolution is 1°. It is used to obtain the three-dimensional point cloud data of the surrounding environment of the target construction robot in the target space. The surrounding refers to the space with a distance greater than 0.2 meters and less than 100 meters from the target construction robot. Since there are also some dynamic obstacles in the target space, relying solely on the lidar to obtain environmental data is likely to result in low detection accuracy of close-range obstacles and the situation of "dark under the lamp".

[0069] Based on the above situation, in the embodiment of the present invention, multiple ultrasonic sensors are installed in the front and side of the target construction robot to obtain the distance data between the target construction robot and the obstacles to cover the close-range, making up for the deficiency of the lidar in detecting close-range obstacles. The close-range refers to the space with a distance greater than 0.02 meters and less than 5 meters from the target construction robot. The model number of the ultrasonic sensor is HC-SR04. At least three ultrasonic sensors are required in the embodiment of the present invention.

[0070] It should be noted that the target construction robot adopts a differential drive mode and has precise motion control capabilities.

[0071] Most preferably, before fusing the state information of the target construction robot with the environmental data through Bayesian filtering, it further includes:

[0072] Preprocess the environmental data to obtain preprocessed environmental data;

[0073] Fuse the state information of the target construction robot with the preprocessed environmental data through Bayesian filtering data. The state information of the target construction robot includes position, movement speed, and action instructions, and the action instructions include forward speed and steering angle.

[0074] In the embodiment of the present invention, the data preprocessing includes:

[0075] I. Preprocess the three-dimensional point cloud data and distance data respectively, specifically as follows:

[0076] For the three-dimensional point cloud data:

[0077] First, perform denoising processing using median filtering to obtain denoised point cloud data;

[0078] Then, use voxel grid filtering to downsample the denoised point cloud data to obtain downsampled point cloud data;

[0079] Then, convert the downsampled point cloud data from the sensor coordinate system to the robot coordinate system;

[0080] Finally, convert the point cloud data in the robot coordinate system to the global coordinate system to obtain the point cloud data in the global coordinate system;

[0081] For the distance data:

[0082] First, perform validity check on the distance data and retain the valid distance data;

[0083] Then, apply moving average filtering to filter the valid distance data to obtain filtered distance data;

[0084] Then, convert the filtered distance data from the sensor coordinate system to the robot coordinate system;

[0085] Finally, convert the distance data in the robot coordinate system to the global coordinate system to obtain the distance data in the global coordinate system;

[0086] II. Interpolate the three-dimensional point cloud data and distance data in the global coordinate system to ensure the synchronization and alignment of the data of the lidar and ultrasonic sensor under the same time reference and match the sampling rates of different sensors, so as to provide high-quality input for subsequent multi-sensor data fusion and environmental modeling.

[0087] Specifically, step 2 includes:

[0088] Obtain the position, movement speed, and action instructions of the target construction robot at the previous moment;

[0089] Based on the position, movement speed, and action instructions of the target construction robot at the previous moment, use the kinematic model to generate a prior state estimate, which includes the current position, speed, and steering information of the target construction robot;

[0090] Fuse the preprocessed environmental data with the prior state estimate to obtain a posterior state estimate;

[0091] Map the posterior state estimate onto a two-dimensional grid map to obtain an occupancy grid map.

[0092] In the embodiments of the present invention, the prediction and update steps of Bayesian filtering are used to estimate the environmental state, and the prediction step and the update step are related and associated in the Bayesian filtering process; the prediction step uses the known system model and control input to generate a prior estimate of the state at the current moment; the update step then uses the environmental data combined with the prior estimate to calculate a more accurate posterior state estimate.

[0093] Specifically, the prediction step provides the prior information required by the update step, and the update step corrects the prediction result to reflect the actually observed environmental changes. This iterative process enables the filter to continuously track and estimate the system state, improving the accuracy and reliability of environmental perception.

[0094] Specifically, the prediction step:

[0095] According to the position, speed, and control input of the target construction robot at the previous moment, construct a kinematic model to predict the environmental state at the current moment. The expression of the kinematic model is:

[0096]

[0097]

[0098] Among them, represents the prior state estimate at the current moment, including information such as position, speed, and steering, represents the state transition function, which is used to describe the state of the target construction robot from time evolves to time , based on the posterior state estimate and the control input predict the prior state estimate at time represents the posterior state estimate at the previous moment, and the previous moment is The moment, including the position, speed, and orientation angle of the construction robot at the previous moment, represents the control input of the target construction robot at the previous moment, which represents the control commands or actions externally applied to the construction robot. These commands directly affect the state change of the construction robot. The control input includes the specific actions executed by the construction robot, , represents the current moment of the control input, represents the construction robot at the moment of the acceleration command, represents the construction robot at the moment of the steering angle change command, represents the prior estimate covariance matrix, represents the linear approximation of the state transition matrix, represents the transpose function of represents the process noise covariance matrix.

[0099] It should be noted that in the Bayesian filtering of the embodiment of the present invention, the state information and control input of the construction robot are represented separately, which can more accurately model and predict the dynamic behavior of the construction robot; while the state information such as the position and speed of the construction robot is integrated in the same state vector, because the position and speed jointly describe the dynamic state of the robot at the current moment, reflecting the specific position of the robot in the target space and its motion state (such as speed, acceleration, etc.). The posterior state estimate usually includes all variables that can describe the current state of the robot, that is:

[0100]

[0101] Among them, represents the posterior estimate of the position coordinates of the construction robot at the previous moment, represents the posterior estimate of the speed of the construction robot at the previous moment, represents the posterior estimate of the orientation angle of the construction robot at the previous moment.

[0102] Combining the above four variables in a state vector helps to uniformly process and update the overall state of the robot in the prediction and update steps of Bayesian filtering.

[0103] Specifically, the update step includes:

[0104] Fuse the environmental data with the prior state estimate to update the estimate of the environmental state. Calculate the Kalman gain by obtaining the environmental data at the current moment. The calculation expression of the Kalman gain is:

[0105]

[0106] Update the posterior state estimate based on the Kalman gain. The update expression for the posterior state estimate is as follows:

[0107]

[0108] Update the posterior covariance matrix based on the Kalman gain and the observation matrix. The expression is:

[0109]

[0110] Where, represents the posterior state estimate, including the new orientation angle , represents the prior state estimate, i.e., the state predicted according to the model, represents the Kalman gain, represents the time of the observation value, represents the observation matrix, represents the transpose matrix of represents the observation noise covariance matrix, represents the identity matrix, represents the posterior state estimate covariance matrix, represents the prior state covariance matrix.

[0111] Specifically, map the posterior state estimate to a two-dimensional grid map to obtain the occupancy grid map. The process is as follows:

[0112] Ensure that all environmental data are processed under a unified reference frame;

[0113] Divide the continuous environmental space into two-dimensional grids of a fixed size. Each grid cell represents a small area in the environment;

[0114] For each grid cell, calculate the occupancy probability value of the cell by analyzing the point cloud density and distance information in the corresponding area of the environmental data. Usually, Bayesian update is used to combine the environmental data with the prior state estimate to dynamically adjust the occupancy state of each grid cell;

[0115] To improve the accuracy and reliability of the map, use Kalman filtering technology to further process the occupancy probability of each grid cell and remove misjudgments caused by sensor noise;

[0116] In addition, combine time series data and use a dynamic update mechanism to reflect moving obstacles and environmental changes in real time to ensure that the occupancy grid map always accurately reflects the current environmental state;

[0117] Finally, a two-dimensional occupancy grid map with a size of 200×200 is generated, providing a reliable environmental representation basis for subsequent path planning and obstacle avoidance decision-making.

[0118] In the embodiment of the present invention, the size of the occupancy grid map is determined according to the working area of the construction robot, the resolution is set to 0.05 meters, and each grid stores information on whether the corresponding position is occupied by an obstacle.

[0119] Most preferably, the trained lightweight convolutional neural network includes:

[0120] An input layer, a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a fully connected layer, and an output layer connected in sequence;

[0121] The input layer is used to receive the occupancy grid map;

[0122] Both the first convolutional layer and the second convolutional layer are used to extract features from the occupancy grid map;

[0123] Both the first pooling layer and the second pooling layer are used to perform downsampling on the extracted features;

[0124] The fully connected layer is used to flatten the features output by the second pooling layer to obtain an environmental feature vector;

[0125] The output layer generates environmental feature data based on the environmental feature vector.

[0126] In the embodiment of the present invention, the first convolutional layer uses 32 3×3 convolutional kernels, the stride is 1, and the activation function is the ReLU function. The second convolutional layer uses 64 3×3 convolutional kernels, the stride is 1, and the activation function is the ReLU function. Both the first pooling layer and the second pooling layer are max pooling layers, the window size is 2×2, and the stride is also 2. The fully connected layer is used to flatten the features output by the convolutional layer and output a feature vector of a fixed length.

[0127] The convolutional neural network is pre-trained to optimize the parameters and can efficiently extract the spatial features of the environment.

[0128] Furthermore, the trained decision-making model includes a policy network and a value network;

[0129] The policy network generates an action probability distribution based on the environmental feature data and the current state of the target construction robot;

[0130] The value network generates a state value estimate based on the environmental feature data and the current state of the target construction robot.

[0131] In the embodiments of the present invention, the PPO algorithm is used to train the decision-making model. The input end of the decision-making model is the environmental feature data and the current state of the construction robot, and the output is the action probability distribution. The functions of the action probability distribution are as follows:

[0132] ① Define the policy and action selection:

[0133] In reinforcement learning, the policy defines the probability distribution of taking an action under a given state . The policy network makes clear the possibility of selecting each action in each state by outputting this probability distribution;

[0134] The robot samples an action from the output action probability distribution for execution. This way introduces randomness, enabling the robot to take different actions in the same state, which helps to explore the environment and avoid getting stuck in local optimal solutions;

[0135] ② Balance exploration and exploitation:

[0136] Explore new policies: Through the probability distribution, the policy network can try various action combinations during the training process to explore new policy spaces;

[0137] ③ Policy update and optimization:

[0138] Calculate the policy gradient: In the PPO algorithm, the calculation of the policy gradient depends on the policy probability ratio;

[0139] Calculation of the advantage function:

[0140] The calculation of the action probability distribution is closely related to the calculation of the advantage function. The advantage function measures the degree of advantage of taking an action in a state relative to the average policy;

[0141] Guide policy improvement: By calculating the advantage function, the policy network can identify which actions are more beneficial under the current policy, and thus tend to increase the selection probability of these actions during training.

[0142] Therefore, this action probability distribution defines the probabilities of various possible actions being selected given the environmental characteristics and the robot's state, guiding the robot on how to sample and execute actions. By outputting a probability distribution, the policy network introduces randomness, enabling the robot to take different actions in the same state, thus facilitating exploration of the environment and preventing the policy from falling into local optima. At the same time, the action probability distribution is the basis for calculating the policy gradient and is used in policy optimization to calculate the policy probability ratio and construct the clipped loss function, ensuring that the magnitude of policy updates is limited and maintaining the stability of training. In addition, the action probability distribution allows the robot to flexibly adjust action selection in dynamic and complex environments, improving adaptability and robustness. Especially in continuous action spaces, it can achieve fine-grained action control. In summary, the action probability distribution in the policy network not only defines the policy and guides action selection but also plays an important role in policy optimization, environment exploration, and adapting to complex task requirements. It is the key to achieving efficient obstacle avoidance path planning.

[0143] Specifically, the PPO (Proximal Policy Optimization) algorithm is a reinforcement learning algorithm that addresses the step-size sensitivity problem of Policy Gradient. Before training the decision model based on the PPO algorithm, it also includes:

[0144] Building a reinforcement learning environment

[0145] Build a simulation environment similar to the actual construction site in simulation software (such as Gazebo), including dynamic obstacles (such as moving human models) and static obstacles (such as stacked material models);

[0146] Designing a reward function

[0147] The reward function is:

[0148]

[0149] Where, represents the total reward, represents the time efficiency reward, represents the goal-oriented reward, represents the obstacle avoidance penalty;

[0150] Since in the path planning of construction robots, safety (obstacle avoidance) takes precedence over efficiency (time taken) and goal orientation, therefore, the weight of the obstacle avoidance penalty is set to a relatively high value to ensure that the robot gives priority to avoiding collisions. According to the rule of thumb, the initial weight coefficient is set as , , .

[0151] The specific definitions of each reward part are as follows:

[0152] Time efficiency reward: ;

[0153] represents the time increment consumed by the robot in the current decision. This reward is directly related to the time consumption and encourages the robot to reduce the time consumption to complete the task at the fastest speed;

[0154] Goal-oriented reward: ;

[0155] represents the Euclidean distance between the current position of the robot and the target position. This reward encourages the robot to approach the target position. The closer the distance, the smaller the reward (the absolute value of the negative value);

[0156] Obstacle avoidance penalty: ;

[0157] represents the collision penalty constant, which takes a positive value. If the robot collides, a fixed negative reward is given to punish the dangerous behavior and ensure safety.

[0158] The training process of the decision-making model based on the PPO algorithm in the embodiments of the present invention is as follows:

[0159] Train the policy network. The input of the policy network is the environmental feature data and the state of the construction robot at the current moment, and the output is the action probability distribution, including the forward speed and the steering angle ;

[0160] Use the simulated environment to collect a large number of state, action, and reward sequences, and adopt the clipping loss function of the PPO algorithm to update the parameters of the policy network and the value network. During the training process, adjust the hyperparameters such as the learning rate and batch size to ensure the stability and convergence speed of the model. Specifically:

[0161] Sampling: In the simulated environment, collect a large number of state, action, and reward sequences according to the current policy;

[0162] Use the sequence data obtained by sampling to calculate the advantage function estimation at each time step , which provides a basis for the update of the policy network. Use the Generalized Advantage Estimation (GAE) method to calculate the advantage function estimation .

[0163] First, perform the return calculation. The discounted cumulative reward (return) starting from time , and the calculation method is: ;

[0164] ;

[0165] Among them, is the discount factor, , is the immediate reward obtained at time , is the time span for return calculation, usually taken until the end of the episode or the set maximum number of steps, is the value estimation of the value network for state ;

[0166] Secondly, value estimation is performed. The value network is used to estimate the value of each state . The value network has a similar structure to the policy network but is used to estimate the value of the state rather than the action probability distribution;

[0167] Finally, the advantage function is estimated, and the calculation formula is:

[0168]

[0169] This estimated value represents the difference between the actual return obtained in state and the prediction of the value network, reflecting the degree of advantage of taking action .

[0170] Policy update:

[0171] Using the estimated advantage function obtained, the clip loss function of the PPO algorithm is adopted to update the parameters of the policy network and the value network, improving the performance of the policy and the stability of the model.

[0172] Policy network update:

[0173] First, calculate the policy probability ratio: ;

[0174] Among them, is the probability that the robot takes action in state under the policy with parameter , representing the probability distribution of the current policy to select actions, is the probability that the robot takes action in state under the policy with parameter , representing the probability distribution of the current policy to select actions, is the action taken by the robot at time step , including the forward speed and the steering angle , is the time when the robot is at The state, including information about the environment and itself. In the embodiments of the present invention, the environmental features refer to the environmental feature representations extracted from sensor data through Bayesian filtering and lightweight CNN, reflecting information such as the positions and distributions of obstacles in the current environment.

[0175] Secondly, construct the clipped loss function:

[0176]

[0177] Wherein, are the policy network parameters, is the advantage function estimate, representing the advantage of the action relative to the average level; is the clipping function, restricting the amplitude of policy update, is the clipping range, generally taking 0.1 or 0.2, represents the expected value at the time step ;

[0178] Finally, update the value network parameters: Use gradient descent or other optimization algorithms to minimize the loss function and update the parameters of the value network;

[0179]

[0180] is the learning rate of the policy network.

[0181] Value network update:

[0182] The training process of the value network is synchronized with the policy network, aiming to improve the value estimation of each state in the environment, thereby supporting the optimization of the policy network; specifically, the training process of the value network includes the following steps:

[0183] First, construct the loss function of the value network. The mean squared error (MSE) is used as the loss function of the value network, and the crossbeam value network predicts and the actual return The difference between them:

[0184]

[0185] Wherein, this loss function is the mean squared error (MSE), measuring the difference between the value network prediction and the actual return ;

[0186] is the batch size, that is, the number of data samples used for one parameter update;

[0187] Secondly, update the value network parameters using gradient descent or other optimization algorithms to minimize the loss function and update the parameters of the value network :

[0188]

[0189] is the learning rate of the value network.

[0190] Synchronously train the policy network and the value network:

[0191] The updates of the policy network and the value network are usually carried out synchronously on the same batch of data to ensure the accuracy of the advantage function estimation and the consistency of the policy update.

[0192] Hyperparameter tuning:

[0193] According to the training effect, appropriately adjust the learning rate , , discount factor , clipping parameter and other hyperparameters. Adjust the batch size N and the number of training epochs to balance the training speed and the model performance.

[0194] Repeat iteration:

[0195] Repeat the three steps of sampling, advantage function calculation, and policy update until the performance of the policy network and the value network meets the preset requirements, such as converging to a stable reward value or achieving the expected task success rate.

[0196] In the embodiment of the present invention, environmental data of the target space is obtained; the state information of the target construction robot is fused with the environmental data through Bayesian filtering to obtain an occupancy grid map; the occupancy grid map is input into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data; the environmental feature data and the current state of the target construction robot are input into a trained decision-making model for decision-making to obtain an action instruction, and the target construction robot reaches the target position in the shortest time based on the action instruction. The path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot; compared with the prior art, in the embodiment of the present invention, first, the state information of the target construction robot is fused with the environmental data through Bayesian filtering, which improves the accuracy and robustness of environmental perception. A lightweight convolutional neural network and an optimized reinforcement learning network are used to extract the features of the data and generate action instructions, which reduces the computational complexity, improves the decision-making speed, meets the real-time requirements, and finally realizes efficient obstacle avoidance and path planning in a complex environment based on the action instructions.

[0197] Corresponding to the obstacle avoidance path planning method of the construction robot described in the above embodiments, as Figure 2 shown, the present invention also provides an obstacle avoidance path planning device 100 for a construction robot. The obstacle avoidance path planning device 100 includes:

[0198] An acquisition module 101, configured to acquire environmental data of a target space;

[0199] A fusion module 102, configured to fuse the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map;

[0200] An extraction module 103, configured to input the occupancy grid map into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data;

[0201] A decision module 104, configured to input the environmental feature data and the current state of the target construction robot into a trained decision model for decision-making to obtain an action instruction. The target construction robot reaches the target position in the shortest time based on the action instruction, and the path passed by the target construction robot to the target position is the obstacle avoidance path of the target construction robot.

[0202] It should be noted that for the information interaction, execution process, etc. between the above devices / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0203] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0204] The present invention also provides a terminal device, as Figure 3 shown, the terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 3Only one processor is shown), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the obstacle avoidance path planning method of the above-mentioned construction robot is implemented.

[0205] The terminal device D10 may be a computing device such as a desktop computer, a notebook, a palm computer, a server, a server cluster, and a cloud server. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art can understand that Figure 3 These are merely examples of the terminal device D10 and do not constitute a limitation on the terminal device D10. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0206] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and the processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), off-the-shelf programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0207] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk equipped on the terminal device D10, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0208] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0209] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details are not described herein again.

[0210] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the obstacle avoidance path planning method for a construction robot.

[0211] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0212] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for obstacle avoidance path planning of a construction robot, characterized in that: include: Step 1: Obtain the target space environment data, including: Setting up a laser radar and multiple ultrasonic sensors inside the target construction robot; Acquire three-dimensional point cloud data of the target space through the laser radar; Acquiring distance data between the target construction robot and obstacles through a plurality of ultrasonic sensors; Integrating the three-dimensional point cloud data and the distance data into environmental data of the target space; Step 2, fusing the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map, including: Obtaining the position, movement speed and action instruction of the target construction robot at the last moment; In the prediction step of the Bayesian filter, a kinematic model is used to generate a priori state estimation based on the position, motion speed and action instruction of the target construction robot at the previous moment, wherein the priori state estimation includes the current position, speed and steering information of the target construction robot; In the updating step of the Bayesian filter, the environmental data is fused with the prior state estimate to obtain a posterior state estimate; Mapping the posterior state estimate onto a two-dimensional grid map to obtain an occupancy grid map; Step 3: Input the occupancy grid map into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data. The trained lightweight convolutional neural network includes: The input layer, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the fully connected layer and the output layer are connected in sequence; The input layer is used to receive the occupancy grid map; The first convolutional layer and the second convolutional layer are both used to extract features from the occupancy grid map; The first pooling layer and the second pooling layer are both used to downsample the extracted features; The fully connected layer is used to flatten the features output by the second pooling layer to obtain an environment feature vector; The output layer generates environmental feature data based on the environmental feature vector; Step 4, inputting the environmental feature data and the current state of the target construction robot into the trained decision model for decision making, obtaining an action instruction, wherein the target construction robot reaches the target position in the shortest time based on the action instruction, and the path taken by the target construction robot to the target position is the obstacle avoidance path of the target construction robot; The trained decision model includes a policy network and a value network; The policy network generates an action probability distribution based on the environmental feature data and the current state of the target construction robot; The value network generates a state value estimate based on the environmental characteristic data and a current state of the target construction robot.

2. The obstacle avoidance path planning method for a construction robot according to claim 1, characterized in that: Before fusing the state information of the target construction robot with the environmental data through Bayesian filtering, the method further includes: Preprocessing the environmental data to obtain preprocessed environmental data; The state information of the target construction robot is fused with the preprocessed environmental data through Bayesian filtering data.

3. The obstacle avoidance path planning method for a construction robot according to claim 1, characterized in that: The calculation expression of the posterior state estimation is: in, represents the posterior state estimate, represents the prior state estimate, represents the Kalman gain, Indicates time The observed value of represents the observation matrix, express The transposed matrix of represents the observation noise covariance matrix, represents the identity matrix, represents the posterior state estimate covariance matrix, represents the prior state covariance matrix.

4. An obstacle avoidance path planning device for a construction robot, characterized in that: include: The acquisition module is used to obtain the environmental data of the target space, including: Setting up a laser radar and multiple ultrasonic sensors inside the target construction robot; Acquire three-dimensional point cloud data of the target space through the laser radar; Acquiring distance data between the target construction robot and obstacles through a plurality of ultrasonic sensors; Integrating the three-dimensional point cloud data and the distance data into environmental data of the target space; A fusion module is used to fuse the state information of the target construction robot with the environmental data through Bayesian filtering to obtain an occupancy grid map, including: Obtaining the position, movement speed and action instruction of the target construction robot at the last moment; Based on the position, movement speed and action instruction of the target construction robot at the previous moment, a priori state estimation is generated using a kinematic model, wherein the priori state estimation includes the current position, speed and steering information of the target construction robot; fusing the environmental data with the prior state estimate to obtain a posterior state estimate; Mapping the posterior state estimate onto a two-dimensional grid map to obtain an occupancy grid map; An extraction module is used to input the occupancy grid map into a trained lightweight convolutional neural network for feature extraction to generate environmental feature data, wherein the trained lightweight convolutional neural network includes: The input layer, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the fully connected layer and the output layer are connected in sequence; The input layer is used to receive the occupancy grid map; The first convolutional layer and the second convolutional layer are both used to extract features from the occupancy grid map; The first pooling layer and the second pooling layer are both used to downsample the extracted features; The fully connected layer is used to flatten the features output by the second pooling layer to obtain an environment feature vector; The output layer generates environmental feature data based on the environmental feature vector; A decision module, used for inputting the environmental feature data and the current state of the target construction robot into the trained decision model to make a decision and obtain an action instruction. The target construction robot reaches the target position in the shortest time based on the action instruction, and the path taken by the target construction robot to the target position is the obstacle avoidance path of the target construction robot; The trained decision model includes a policy network and a value network; The policy network generates an action probability distribution based on the environmental feature data and the current state of the target construction robot; The value network generates a state value estimate based on the environmental characteristic data and a current state of the target construction robot.

5. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the obstacle avoidance path planning method for the construction robot as described in any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the obstacle avoidance path planning method for a construction robot as described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Path planning method fusing global artificial potential field and local reinforcement learning

    CN117539241A

  • Dynamic obstacle avoidance system and method for animal house cleaning robot based on laser sensing

    CN118760195A

Cited By

  • Building construction robot walking system based on intelligent obstacle avoidance and path planning algorithm

    CN120891826A

  • Construction robot walking system based on intelligent obstacle avoidance and path planning algorithm

    CN120891826B