Synchronous localization and mapping method based on PID (Proportion Integration Differentiation) real-time correction fusion vision
By introducing a synchronous positioning and map construction method based on PID real-time correction fused vision in the robot control system, combining SAC-PID controller and ORB-SLAM technology, the problem of insufficient accuracy in complex environments of traditional PID control methods is solved, and the real-time performance and map consistency of visual SLAM technology are improved, achieving more efficient and stable robot navigation and path tracking.
Patent Information
- Application Number
- CN202510299835.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional PID control methods are difficult to maintain high accuracy when facing complex and dynamically changing environments. Especially when path curves are complex and environment changes violently, PID controllers are prone to error accumulation or excessive correction, resulting in robot path deviation; while existing visual SLAM technologies still have challenges in real-time performance and map consistency, especially when long-term operation, the processing capabilities of loopback detection and map optimization are insufficient.
The synchronous positioning and map construction method based on PID-based real-time correction and fused vision are adopted. By collecting the robot's environment data in real time, the direction angle of each feature point is extracted and calculated, the state space and action space are defined, the control signal of the robot is adjusted using the SAC-PID controller, and the visual SLAM outputs a global consistent map through ORB-SLAM.
It improves the accuracy and stability of robot path tracking, enhances the adaptability of the controller and optimizes the environment perception technology, makes up for the shortcomings of insufficient control accuracy and poor map consistency in the existing technology, and improves the robot's navigation capabilities in complex dynamic environments.
Smart Images

Figure CN120176653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of automatic control and environmental perception, and particularly to a method for simultaneous localization and mapping (SLAM) based on real-time PID correction and integrated vision.
[0002] With the rapid development of intelligent robot technology, robots have become increasingly prominent in various application fields. Especially in automatic navigation and positioning, mobile robots rely on efficient control systems and precise environmental perception technologies to achieve autonomous driving, task execution, and environmental interaction. Traditional path tracking control methods, such as PID control (Proportional-Integral-Derivative control), have been widely applied in robot control. The advantage of PID control lies in its simple structure, easy implementation, and ability to provide relatively stable control effects. However, traditional PID control often has certain limitations when facing complex environments and dynamic changes. Especially when the path error is large or the robot is in a non-linear motion state, traditional PID controllers often struggle to achieve optimal real-time adjustment, resulting in a decrease in accuracy and limited control performance.
[0003] Although existing visual SLAM and PID control methods have been widely applied in many fields, there are still certain deficiencies. Traditional PID control algorithms are difficult to maintain high precision when facing complex and dynamically changing environments. Especially in situations where the path curve is complex and the environment changes violently, the PID controller is prone to error accumulation or overcorrection, leading to deviation of the robot path. While existing visual SLAM technologies can provide relatively accurate environmental mapping, they still face challenges in terms of real-time performance and consistency. Especially during long-term operation, the processing capabilities of loop detection and map optimization are often insufficient to handle the long-term autonomous operation of robots in complex scenarios. In the existing technology, the combination of the control system and the environmental perception system is often relatively loose, lacking an intelligent adjustment mechanism that can real-time coordinate and optimize the data fusion and control signals between the two. Figure 1 Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for simultaneous localization and mapping based on real-time PID correction and integrated vision to solve the problem that existing PID control methods often show poor adaptability when facing non-linear changes in dynamic environments, while visual SLAM technology has problems with insufficient loop detection and map optimization accuracy.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a method for simultaneous localization and mapping based on real-time PID correction and integrated vision, which includes: Collect the environmental data of the robot in real time and perform preprocessing; Extract and calculate the direction angle of each feature point based on the environmental data and calculate the path information; Define the state space and action space based on the path information, and adjust the control signal of the robot through the SAC-PID controller; Use ORB-SLAM for visual SLAM to output a globally consistent map; The environmental data includes point cloud data, RGB images, acceleration, and angular velocity data.
[0007] As a preferred solution of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision according to the present invention, wherein: the real-time collection of the environmental data of the robot includes: Select a lidar, GPS, RGB camera, and inertial measurement unit as the collection devices, determine the target area and install the collection devices on the robot, set the scanning frequency of the collection devices, collect the environmental data in real time, set the time synchronization window, and synchronize all the environmental data according to a unified timestamp by using linear interpolation.
[0008] As a preferred solution of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision according to the present invention, wherein: the extraction and calculation of the direction angle of each feature point based on the environmental data includes: Use the FAST algorithm to calculate the response value of each pixel point by judging the brightness change of the surrounding area of all pixel points in the grayscale image; According to the response value of each pixel point in the grayscale image, set the response threshold. If the point has a response value greater than the response threshold, then mark the point as a feature point , otherwise, do not perform any operation; Calculate the direction angle of each feature point. The formula is: , where, represents the direction angle of the feature point p, and represent the coordinates of the feature point and its neighboring point respectively.
[0009] As a preferred solution of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision according to the present invention, wherein: the calculation of the path information includes: Convert the grayscale image into a binary image, use the Canny edge detection algorithm to extract the edge of the path area in the binary image, and extract the complete path image through morphological operations; By performing a vertical projection on the binary image, the projection value of each column is obtained, and the upper and lower boundaries of the path are determined. The center line of the path is obtained by calculating the vertical center positions of the upper and lower boundaries of the path. According to the current position of the robot, the pixel coordinates of the current position of the robot at different times in the binary image are determined. Based on the center line of the path, the horizontal distance between the current position of the robot at different times and the center line of the path is calculated. Set the position coordinates of two adjacent points in the path as (x1, y1) and (x2, y2), and calculate the direction angle of the path through the angle between the two adjacent points in the path. , the formula is: , where atan2(y, x) is to calculate the direction angle from the x-axis to the tangent direction of the path. Based on the direction angle of the path, calculate the angular change rate of the path through the time difference. The formula is: , where and represent the direction angles of the path at the current frame t and the previous frame t - 1 respectively. According to the angular change rate of the path, calculate the curvature of the path. The formula is: , where represents the curvature of the path at the current frame t. Based on the direction angle of the path, calculate the lateral offset error of each feature point. The formula is: d error,p = d p ·sin(Δθ p ), where d error,p represents the lateral offset error of the feature point, d p represents the Euclidean distance between the feature point p and the current position of the robot, represents the angular error between the direction angle of the feature point p and the direction angle of the path. Perform a weighted average on the lateral offset errors of all feature points to obtain the final lateral offset error. Calculate the angular error between the direction angle of each feature point and the direction angle of the path and perform a weighted average to obtain the local positioning angle. Define the global orientation angle, which represents the deviation angle between the robot and the global target direction. By using the cosine values of the local positioning angle and the global orientation angle, calculate the orientation angle between the robot and the path direction. The formula is: where, represents the orientation angle between the current direction of the robot and the path direction, represents the cosine value of the global orientation angle, represents the cosine value of the local positioning angle.
[0010] As a preferred solution of the simultaneous localization and mapping method based on PID real-time correction and fusion vision according to the present invention, wherein: the defining the state space and the action space based on the path information includes: Normalize the path information and define the state space ; Take the three core gains of the PID controller as the variables of the action space ; Determine the range of the PID gains and use the range of the gains as the continuous action space. Model the actions using the Gaussian distribution and initialize the gain parameters.
[0011] As a preferred solution of the simultaneous localization and mapping method based on PID real-time correction and fusion vision according to the present invention, wherein: the adjusting the control signal of the robot by the SAC-PID controller includes: Define the total error. The formula is: , where, represents the total error, and represent the weighting factors; Measure the control effort by calculating the output increment of the PID controller. The formula is: , where, represents the control effort value, represents the perpendicular distance from the position coordinate of the robot at the current frame t to the midline of the path, represents the infinitesimal variable of the time integration variable t; Design the speed optimization reward term and give rewards based on the linear velocity. The formula is: , where, represents the speed optimization reward value, represents the speed weight; Construct the comprehensive reward function. The formula is: , Among them, represents the comprehensive reward value; Initialize the SAC algorithm of the reinforcement learning model, including the value network, Q-value network, and policy network; According to the current state space , generate the action space through the policy network, and update the policy network according to the comprehensive reward function. During the training process, the SAC algorithm continuously updates the policy network parameters to maximize the cumulative reward of each action space . Evaluate the quality of the current policy through the Q-value network, calculate the Q-value of each state-action pair using the Q-learning method, estimate the expected reward of each state space using the value network, provide the future expected value estimate by combining the outputs of the Q-value network and the policy network, use the target network to update the parameters of the Q-value network and the value network in a soft update manner, and store the quadruple of the state space, action space, reward value, and next state; Update through randomly sampling the experience samples in the replay pool, using the reward value and the value of the next state. Update the Q-value network parameters by minimizing the loss function, update the value network by combining the Q-value output by the current policy and the entropy of the policy, generate actions using the mean and covariance of the Gaussian distribution, update the policy network using the gradient descent method to maximize the policy objective, repeat the training steps, and verify the path tracking accuracy, loop detection accuracy, and control stability through simulation tests and actual experiments; Through the reinforcement learning model, based on the state space obtain the PID gain values , , of the current frame. According to the gain values, adjust the PID controller and calculate the PID increment. The formula is: , Among them, represents the PID increment at the current frame , and respectively represent the total errors of the previous frame and the previous two frames ; Generate a control signal according to the PID increment, and output the control signal to the drive system of the robot to control the angular velocity and linear velocity of the robot's turning. The formula is: , , Among them, represents the control quantity generated by the PID controller, directly controlling the angular velocity of the robot's turning, and represent the minimum and maximum values of the linear velocity control line; According to the current state space and the calculated PID increment, the reinforcement learning model updates the PID gain parameters in real time and adjusts the control signal accordingly.
[0012] As a preferred solution of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision according to the present invention, wherein: the use of ORB-SLAM for visual SLAM to output a globally consistent map includes: Obtain a real-time image stream from an RGB camera, perform image preprocessing, use ORB-SLAM to extract feature points from the image and perform matching, use three-dimensional reconstruction technology, and realize environmental modeling through feature matching and visual inertial odometry, and output the pose estimation of the robot and the current local map; During the process of constructing the local map, use the loop detection module of ORB-SLAM to analyze the similarity between the current position of the robot and the positions in the known map. When a robot loop is detected, use a deep learning model to judge the matching situation between the current position and the historical position. After determining the loop, feedback the pose information of the loop position to the SLAM system to trigger the optimization of the global map; Construct a graph optimization problem through the pose information obtained by loop detection, define nodes and constraints, use the Gauss-Newton method for nonlinear optimization, adjust the pose values of the nodes, minimize the cost function, and perform iterative convergence of the graph until the cost function reaches the minimum value to obtain an optimized globally consistent map.
[0013] As a preferred solution of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision according to the present invention, wherein: the preprocessing includes: Perform voxel filtering on the point cloud data, convert the RGB image to a grayscale image and perform Gaussian filtering, and use low-pass filtering on the acceleration and angular velocity data to remove high-frequency noise.
[0014] In a third aspect, the present invention provides a computer device, including a memory and a processor, where: the memory stores a computer program, wherein: when the computer program is executed by the processor, it realizes any step of the method for simultaneous localization and mapping based on PID real-time correction and fusion of vision as described in the first aspect of the present invention.
[0015] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, any step of the simultaneous localization and mapping method based on PID real-time correction and fusion of vision as described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: through the intelligent optimization and adjustment of the SAC-PID controller combined with reinforcement learning, the robot path is dynamically corrected, thereby improving the accuracy and stability of path tracking. Through the ORB-SLAM technology, the construction and optimization of a globally consistent map are carried out, further enhancing the robot's navigation ability in complex dynamic environments, enhancing the self-adaptability of the controller and optimizing the environmental perception technology, and making up for the deficiencies in control accuracy and poor global consistency in the prior art. Figure 1 defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the simultaneous localization and mapping method based on PID real-time correction and fusion of vision in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the above-mentioned objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification.
[0020] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments.
[0022] Embodiment 1, referring to Figure 1 , is the first embodiment of the present invention. This embodiment provides a simultaneous localization and mapping method based on PID real-time correction and fusion of vision, including the following steps: S1. Collect the environmental data of the robot in real time and perform preprocessing; Specifically, the real-time collection of the environmental data of the robot includes: According to the requirements of map construction, select a lidar (LiDAR), GPS, RGB camera, and inertial measurement unit (IMU) as the collection devices. According to the actual requirements, determine the target area and install the collection devices on the robot. Set the scanning frequency of the collection devices, collect the environmental data in real time, set the time synchronization window, and use linear interpolation to synchronize all the environmental data with a unified timestamp.
[0023] In the present invention, by integrating multi-source sensors such as LiDAR, GPS, RGB camera, and IMU, high-precision collection of the environmental data of the robot is achieved. Compared with the traditional single-sensor solution, this method can provide more comprehensive perception information in a complex environment, improve the robustness of positioning and map construction, perform time alignment on various types of data through the time synchronization window and the linear interpolation method, ensure the consistency and real-time performance of data fusion, thereby reducing the positioning error caused by the asynchronous sensor timing, improving the accuracy and stability of the SLAM system, and enabling the robot to more accurately perceive and map in a dynamic environment.
[0024] S2. Extract and calculate the direction angle of each feature point based on the environmental data and calculate the path information; Specifically, the extraction and calculation of the direction angle of each feature point based on the environmental data include: Use the FAST (Features from Accelerated Segment Test) algorithm. By judging the brightness change of the surrounding areas of all pixel points in the grayscale image, calculate the response value of each pixel point. The formula is: , where S p represents the response value of the point in the grayscale image, represents the number of neighboring points of the point , represents the pixel value of the point in the grayscale image at the current frame t, represents the neighboring points around the point ; According to the response value of each pixel point in the grayscale image, set a response threshold. If the response value of the point is greater than the response threshold, then mark the point as a feature point , otherwise, do not perform any operation; Calculate the direction angle of each feature point. The formula is: , Among them, represents the orientation angle of the feature point of, and respectively represent the feature point and its neighborhood points coordinates of.
[0025] By using the FAST algorithm to calculate the response value of each pixel point and extract feature points, significant features in the image can be effectively identified. This method calculates the brightness change in the neighborhood around the pixel points in the image and sets a response threshold to filter out points with strong local features, thereby improving the extraction efficiency and accuracy of feature points. By calculating the orientation angle of each feature point, accurate direction information can be provided for subsequent path tracking and map construction. Compared with traditional feature extraction methods, the FAST algorithm has obvious advantages in processing speed and computational efficiency, and can quickly respond in real-time applications to ensure that the robot accurately locates and navigates in a dynamic environment.
[0026] Furthermore, the calculation of path information includes: Preset a fixed threshold to convert the grayscale image into a binary image. Select the optimal fixed threshold through experiments. Set the pixel value of the pixel points in the grayscale image greater than or equal to the fixed threshold to 1 (white), otherwise, set it to 0 (black). White represents the path area, and black represents the background. Use the Canny edge detection algorithm to extract the edge of the path area in the binary image. Perform morphological operations (dilation and erosion) on the binary image to extract the complete path image. Apply the erosion operation to remove the noise in the binary image, and restore the continuity of the path through the dilation operation; By performing vertical projection on the binary image, that is, summing each column in the binary image to obtain the projection value of each column and determine the upper and lower boundaries of the path. The upper and lower boundaries of the path are usually the start and end positions of the high projection value area; Obtain the path center line by calculating the vertical center position of the upper and lower boundaries of the path; The path center line scans the vertical center position of the path in the binary image and calculates the path center line along the vertical projection. If the path is curved or has a complex shape, interpolation methods (such as linear interpolation, spline interpolation) are used to refine and correct the path center line to make it more suitable for the actual situation of the robot; According to the current position of the robot, determine the pixel coordinates of the current position of the robot at different times in the binary image. Based on the path center line, calculate the horizontal distance between the current position of the robot at different times and the path center line; Set the position coordinates of two adjacent points on the path as (x1, y1) and (x2, y2), and calculate the direction angle of the path through the angle between two adjacent points on the path. , the formula is: , where atan2(y, x) is the direction angle calculated from the x-axis to the tangent direction of the path; Based on the direction angle of the path, calculate the angular change rate of the path through the time difference, and the formula is: , where, and represent the direction angles of the path at the current frame t and the previous frame t - 1 respectively; According to the angular change rate of the path, calculate the curvature of the path, and the formula is: , where, represents the curvature of the path at the current frame t; Based on the direction angle of the path, calculate the lateral offset error of each feature point, and the formula is: d error,p = d p ·sin(Δθ p ) , where, d error,p represents the lateral offset error of the feature point , d p represents the Euclidean distance between the feature point and the current position of the robot, represents the angular error between the direction angle of the feature point and the direction angle of the path; Perform weighted averaging on the lateral offset errors of all feature points to obtain the final lateral offset error; Calculate the angular error between the direction angle of each feature point and the direction angle of the path and perform weighted averaging to obtain the local positioning angle; Define the global orientation angle, which represents the deviation angle between the robot and the global target direction (such as the north direction or the reference direction), and obtain it by extracting the global orientation angle through GPS; Through the cosine values of the local positioning angle and the global orientation angle, calculate the direction angle between the robot and the path direction, and the formula is: where, Represents the angular direction between the current direction of the robot and the path direction. The alignment degree between the robot and the path is measured by cosine similarity. Minimizing this angle means making the robot as close as possible to the target path. Represents the cosine value of the global orientation angle. Represents the cosine value of the local positioning angle.
[0027] Based on the above path information calculation method, the present invention effectively improves the accuracy and stability of robot path tracking through precise image processing and mathematical modeling. By converting the grayscale image into a binary image and combining Canny edge detection with morphological operations, the accurate extraction of the path area is ensured, while noise is removed and path continuity is restored. Based on the longitudinal projection method and the calculation of the path center line, the boundaries of the path are effectively identified and corrected, and then the position of the robot on the path is accurately determined. By calculating the path direction angle and curvature, the path adjustment of the robot is further optimized, and a more stable positioning angle is provided through the weighted average of the lateral offset error and the path direction error. Combining the cosine similarity calculation of the global orientation angle and the local positioning angle can accurately evaluate the alignment degree between the robot and the path, ensuring that the path correction and positioning of the robot are more accurate and flexible. This method significantly improves the path tracking accuracy of the robot in a dynamic environment, reduces deviations and errors, and improves the robustness and adaptability of path planning.
[0028] S3. Define the state space and action space based on the path information, and adjust the control signal of the robot through the SAC-PID controller; Specifically, defining the state space and action space based on the path information includes: Normalize the path information and define the state space , and the formula is: , Among them, respectively represent the position coordinates of the robot at the current frame t, represents the angular direction between the robot and the path direction at the current frame t, represents the perpendicular distance from the position coordinates of the robot at the current frame t to the path center line, which is determined by calculating the minimum distance between the robot position and each line segment of the path, represents the lateral offset error of the robot at the current frame t, and respectively represent the linear velocity and angular velocity of the robot at the current frame t. The linear velocity is obtained by integrating the acceleration measured by the accelerometer of the IMU, and the angular velocity is directly measured from the IMU. represents the curvature of the path at the current frame t; Take the three core gains of the PID controller as variables in the action space , and the formula is: , where , , represent the proportional gain, integral gain, and derivative gain respectively; Determine the range of the PID gains and use the range of the gains as the continuous action space. Model the actions using a Gaussian distribution (normal distribution). Initialize the gain parameters according to the requirements of the control system, and the initial values are obtained through empirical control experiments; The range of the PID gains includes setting value ranges for , , according to actual needs. If the proportional gain is too large, it will cause system oscillation, and if it is too small, the control effect will not be obvious, such as 0 to 10. The integral gain is usually small and is mainly used to compensate for long-term errors. If it is too large, it will cause the response of the controller to be unstable, such as 0 to 1. The derivative gain helps the system cope with rapidly changing errors and avoid overshoot, and its range is usually small, such as 0 to 2.
[0029] The method of defining the state space and action space based on path information accurately describes key factors such as the current position, orientation angle, offset error, speed, and path curvature of the robot by normalizing the path information, effectively improving the accuracy of the robot's state representation. Using the three core gains of the PID controller as variables in the action space and modeling through a Gaussian distribution ensure that the PID gains can be optimized and adjusted within a reasonable range, avoiding system oscillation caused by too large gains or insignificant control effects caused by too small gains. This method enables the control system to adaptively adjust the PID parameters, improves the path tracking accuracy, system stability, and adaptability in a dynamic environment, and optimizes the path tracking performance of the robot in a complex environment.
[0030] Furthermore, adjusting the control signal of the robot through the SAC-PID controller includes: When the robot deviates from the path, we need to design a negative reward to punish this deviation. Define the total error, and the formula is: , where represents the total error, and represent the weighting factors, which are adjusted according to actual needs; Excessive adjustment of the controller will affect system stability. It is necessary to design a penalty term to punish excessive control effort. Control effort is the response intensity of the PID controller to the error, and the control effort is measured by calculating the output increment of the PID controller. The formula is: , where, represents the control effort value, represents the vertical distance from the position coordinate of the robot at the current frame t to the midline of the path, represents the tiny variable of the time integration variable t; To ensure that the robot can quickly complete the path tracking task, it is necessary to design a speed optimization reward term. The greater the linear speed of the robot during path tracking, the faster the task is completed. A reward is given based on the linear speed. The formula is: , where, represents the speed optimization reward value, represents the speed weight, which is used to adjust the weight of the speed reward; Construct a comprehensive reward function. The formula is: , where, represents the comprehensive reward value; Initialize the SAC algorithm of the reinforcement learning model, including the value network, Q-value network, and policy network; According to the current state space , generate the action space through the policy network (Actor), that is, the PID gains , , , and update the policy network according to the comprehensive reward function. During the training process, the SAC algorithm continuously updates the policy network parameters to maximize the cumulative reward of each action space . Evaluate the quality of the current policy through the Q-value network (Critic), calculate the Q-value of each state-action pair using the Q-learning method, estimate the expected reward of each state space using the value network (Value), provide an estimated value of the future expectation by combining the outputs of the Q-value network and the policy network, and update the parameters of the Q-value network and the value network in a soft update manner using the target network (TargetNetworks), and store the quadruple of the state space, action space, reward value, and next state; By randomly sampling the experience samples in the replay pool and using the reward value and the value of the next state for update, which represents the expected long-term reward under a given state-action pair. By minimizing the loss function (such as TD error), the Q-value network parameters are updated. Combining the Q-value output by the current policy and the entropy of the policy, the value network is updated by minimizing the loss function. The action is generated using the mean and covariance of the Gaussian distribution, and the policy network is updated using the gradient descent method (such as the Adam optimizer) to maximize the policy objective. The training steps are repeated, including experience sampling, Q-value update, value network update, and policy optimization for training, gradually improving the control accuracy. Through simulation tests and actual experiments, the path tracking accuracy, loop detection accuracy, and control stability are verified; The policy refers to the rule or function by which an agent (robot) selects the action space under a given state space. The policy defines how the agent decides the specific action to execute according to the current state, and through the policy, the agent can take appropriate actions under different environmental states to maximize the expected cumulative reward; Through the reinforcement learning model, based on the state space Obtain the PID gain value of the current frame 、 、 , the PID gain output by the reinforcement learning model will be adjusted according to the state of the robot and can flexibly change under different environmental conditions to reduce the path tracking error. The PID controller is adjusted according to the gain value, and the PID increment is calculated. The formula is: , Among them, represents the PID increment at the current frame , and respectively represent the total errors of the previous frame and the frame before the previous frame ; Generate a control signal according to the PID increment , and output the control signal to the drive system of the robot to control the angular velocity of the robot's turning and the linear velocity , , Among them, represents the control quantity generated by the PID controller, directly controlling the angular velocity of the robot's turning, and represent the minimum and maximum values for controlling the linear velocity, The smaller it is, the slower the robot will move, ensuring that the robot will not move too fast when approaching the target path. The larger it is, the faster the robot will move. According to the current state space and the calculated PID increment, the reinforcement learning model updates the PID gain parameters in real time and adjusts the control signal accordingly.
[0031] By introducing the reinforcement learning method of the SAC-PID controller, the present invention significantly improves the accuracy and stability of robot path tracking. By means of the negative reward mechanism, path deviation is punished to ensure that the robot can maintain accurate path tracking. The penalty term of the control effort is designed to effectively prevent over-adjustment from affecting the system stability. The speed optimization reward term promotes the robot to complete tasks efficiently, improving the task execution speed. The design of the comprehensive reward function enables the controller to flexibly adjust the PID gain to achieve precise control of the path. The reinforcement learning model dynamically optimizes the PID gain through the combination of the policy network, Q-value network and value network, enabling the robot to adaptively adjust under different environmental conditions, thereby minimizing the path error and improving the control accuracy. Through this innovative mechanism, the present invention can achieve efficient and stable path tracking, and has a strong self-optimization ability, enhancing the navigation performance and real-time adaptation ability of the robot in complex environments.
[0032] S4. Use ORB-SLAM for visual SLAM to output a globally consistent map. Specifically, using ORB-SLAM for visual SLAM to output a globally consistent map includes: ORB-SLAM is a visual SLAM algorithm that can construct a local map of the environment where the robot is currently located based on the image data of the camera, obtain a real-time image stream from the RGB camera, and perform image preprocessing (such as denoising, grayscale conversion, binarization). Use ORB-SLAM to extract feature points (such as corner points, edges) from the image and perform matching. Use 3D reconstruction technology to realize environmental modeling through feature matching and visual inertial odometry (VIO), and output the pose estimation of the robot and the current local map (usually a 3D point cloud or 2D grid map). During the process of constructing the local map, the loop detection module of ORB-SLAM is used to analyze the similarity between the current position of the robot and the positions in the known map. When a loop of the robot is detected, a deep learning model (such as CNN) is used to judge the matching situation between the current position and the historical position. After determining the loop, the pose information of the loop position is fed back to the SLAM system to trigger the optimization of the global map. Construct a graph optimization problem based on the pose information obtained through loop detection, define nodes and constraints, use the Gauss-Newton method for nonlinear optimization, adjust the pose values of the nodes, minimize the cost function, and perform iterative convergence of the graph until the cost function reaches the minimum value to obtain an optimized globally consistent map; The definition of nodes and constraints means defining each node and edge in the graph through g2o, where the nodes represent the poses of the robot and the edges represent the constraints between the nodes (such as motion constraints or loop constraints).
[0033] The present invention can effectively construct a globally consistent map by combining the ORB-SLAM algorithm and loop detection technology. ORB-SLAM constructs a high-precision local map by real-time extracting and matching feature points in images, and further optimizes pose estimation through Visual-Inertial Odometry (VIO) technology. The loop detection function can identify whether the robot has returned to a known position and uses a deep learning model to improve the accuracy of loop matching. When a loop occurs, the pose information is fed back to the SLAM system, triggering the graph optimization process. Through the g2o graph optimization algorithm, the robot pose is adjusted to minimize the cost function, thereby generating a globally consistent optimized map. This method significantly improves the accuracy and consistency of map construction, solves common problems in loop detection and global optimization, and provides strong support for the precise navigation and positioning of robots in complex environments.
[0034] In addition, the preprocessing includes: Perform voxel filtering on the point cloud data, convert the RGB image to a grayscale image and perform Gaussian filtering, and use low-pass filtering on the acceleration and angular velocity data to remove high-frequency noise.
[0035] This preprocessing step effectively reduces the noise and redundant information in the environmental data by performing voxel filtering on the point cloud data, converting the RGB image to a grayscale image and performing Gaussian filtering, and performing low-pass filtering on the acceleration and angular velocity data. Voxel filtering can reduce the computational complexity of the point cloud data, remove unnecessary details, and make map construction more efficient; grayscale conversion and Gaussian filtering contribute to the smoothing and noise suppression of image data, improving the accuracy of image feature extraction; low-pass filtering ensures the accurate estimation of the robot's motion state by removing high-frequency noise in the acceleration and angular velocity data. The combination of these processing means significantly improves the accuracy and stability of subsequent data fusion, path calculation, and map construction, and enhances the robustness of the system in complex environments.
[0036] This embodiment also provides a computer device, which is applicable to the case of the simultaneous localization and mapping method based on real-time PID-corrected fused vision, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the simultaneous localization and mapping method based on real-time PID-corrected fused vision proposed in the above embodiment.
[0037] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0038] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the simultaneous localization and mapping method based on real-time PID-corrected fused vision proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0039] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A synchronous positioning and map construction method based on PID real-time correction fusion vision, characterized by: include: Collect the robot's environmental data in real time and pre-process it; Extract and calculate the direction angle of each feature point based on the environmental data and calculate the path information; Define the state space and action space based on the path information, and adjust the robot's control signal through the SAC-PID controller; Use ORB-SLAM for visual SLAM to output a globally consistent map; The environmental data includes point cloud data, RGB images, acceleration and angular velocity data.
2. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 1, characterized in that: The real-time collection of robot environmental data includes: Select LiDAR, GPS, RGB camera and inertial measurement unit as acquisition devices, determine the target area and install the acquisition devices on the robot, set the scanning frequency of the acquisition devices, collect environmental data in real time, set the time synchronization window, and use linear interpolation to synchronize all environmental data according to a unified timestamp.
3. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 2, characterized in that: The step of extracting and calculating the direction angle of each feature point based on the environmental data comprises: Using the FAST algorithm, the response value of each pixel is calculated by judging the brightness changes in the area around all pixels in the grayscale image. According to the response value of each pixel in the grayscale image, the response threshold is set. If the response value is greater than the response threshold, the mark point For feature points , otherwise, no operation is performed; Calculate the direction angle of each feature point using the formula: , in, Representing feature points The direction angle, and Represents the feature points Its field points The coordinates of .
4. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 3 is characterized by: The calculation path information includes: Convert the grayscale image into a binary image, use the Canny edge detection algorithm to extract the path area edge in the binary image, and extract the complete path image through morphological operations; By performing longitudinal projection on the binary image, the projection value of each column is obtained and the upper and lower boundaries of the path are determined. The center line of the path is obtained by calculating the vertical center positions of the upper and lower boundaries of the path. According to the current position of the robot, the pixel coordinates of the current position of the robot at different times in the binary image are determined, and based on the center line of the path, the lateral distance between the current position of the robot at different times and the center line of the path is calculated; Set the position coordinates of two adjacent points in the path to (x1, y1) and (x2, y2), and calculate the direction angle of the path through the angle between the two adjacent points in the path. , the formula is: , Among them, atan2(y,x) is used to calculate the direction angle from the x-axis to the tangent direction of the path; Based on the direction angle of the path, the angle change rate of the path is calculated by the time difference. The formula is: , in, and Respectively represent the direction angles of the path at the current frame t and the previous frame t-1; According to the angle change rate of the path, the curvature of the path is calculated using the formula: , in, Represents the curvature of the path at the current frame t; The lateral offset error of each feature point is calculated based on the direction angle of the path. The formula is: d error,p =d p ·sin(Δθ p ) , Among them, d error,p Representing feature points The lateral offset error, d p Representing feature points The Euclidean distance from the robot's current position, Representing feature points The angular error between the direction angle of and the direction angle of the path; The lateral offset errors of all feature points are weighted averaged to obtain the final lateral offset error; Calculate the angle error between the direction angle of each feature point and the direction angle of the path and take the weighted average to obtain the local positioning angle; Define the global orientation angle, which represents the deviation angle between the robot and the global target direction; The direction angle between the robot and the path direction is calculated by the cosine value of the local positioning angle and the global orientation angle. The formula is: in, Indicates the direction angle between the robot's current direction and the path direction. represents the cosine of the global orientation angle, Indicates the cosine of the local positioning angle.
5. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 4 is characterized by: Defining the state space and the action space based on the path information includes: Normalize the path information and define the state space ; The three core gains of the PID controller are used as variables in the action space ; Determine the range of PID gain and use the range of gain as a continuous action space, use Gaussian distribution to model the action, and initialize the gain parameters.
6. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 5, characterized in that: The control signal of the robot is adjusted by the SAC-PID controller, comprising: Define the total error as: , in, represents the total error, and represents the weighting factor; The control effort is measured by calculating the output increment of the PID controller, as follows: , in, represents the control effort value, Indicates the vertical distance from the robot's position coordinates at the current frame t to the center line of the path, A tiny variable representing the time-integrated variable t; Design speed optimization bonus items, give bonuses based on line speed, the formula is: , in, represents the speed optimization reward value, represents the speed weight; Construct a comprehensive reward function, the formula is: , in, Indicates the comprehensive reward value; Initialize the SAC algorithm of the reinforcement learning model, including the value network, Q-value network, and policy network; According to the current state space , generating action space through policy network , and updates the policy network according to the comprehensive reward function. During the training process, the SAC algorithm continuously updates the policy network parameters to maximize each action space The cumulative reward is used to evaluate the quality of the current strategy through the Q-value network, the Q-value of each state-action pair is calculated using the Q-learning method, and the value network is used to estimate the Q value of each state space The expected reward is combined with the output of the Q-value network and the policy network to provide an estimate of the expected value in the future. The target network is used to update the parameters of the Q-value network and the value network through soft updates, and the four-tuple of state space, action space, reward value and next state is stored. By randomly sampling experience samples in the playback pool, using the reward value and the value of the next state to update, by minimizing the loss function, updating the Q-value network parameters, combining the Q-value output of the current strategy and the entropy of the strategy, updating the value network by minimizing the loss function, using the mean and covariance of the Gaussian distribution to generate actions, using the gradient descent method to update the policy network, maximizing the policy goal, repeating the training steps, and verifying the path tracking accuracy, loop detection accuracy and control stability through simulation tests and actual experiments; Through the reinforcement learning model, based on the state space Get the PID gain value of the current frame , , , adjust the PID controller according to the gain value and calculate the PID increment. The formula is: , in, Indicates that in the current frame The PID increment at and Respectively represent the previous frame and the first two frames The total error of Generate control signal based on PID increment , and output the control signal to the robot's drive system to control the robot's steering angular velocity and line speed , the formula is: , , in, It represents the control quantity generated by the PID controller, which directly controls the angular velocity of the robot's steering. and Indicates the minimum and maximum values of the controlled line speed; Based on the current state space and the calculated PID increments, the reinforcement learning model updates the PID gain parameters in real time and adjusts the control signal accordingly.
7. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 6, characterized in that: The use of ORB-SLAM to perform visual SLAM to output a globally consistent map includes: Obtain real-time image stream from RGB camera and perform image preprocessing. Use ORB-SLAM to extract feature points from the image and match them. Use 3D reconstruction technology to achieve environment modeling through feature matching and visual inertial odometer, and output the robot's pose estimation and current local map. In the process of building a local map, the loop detection module of ORB-SLAM is used to analyze the similarity between the current position of the robot and the position in the known map. When a robot loop is detected, the deep learning model is used to determine the match between the current position and the historical position. After the loop is determined, the posture information of the loop position is fed back to the SLAM system to trigger the optimization of the global map. The pose information obtained through loop closure detection is used to construct a graph optimization problem, define nodes and constraints, use the Gauss-Newton method for nonlinear optimization, adjust the pose values of the nodes, minimize the cost function, and perform convergence iterations of the graph until the cost function reaches the minimum value, thus obtaining an optimized global consistent map.
8. The synchronous positioning and mapping method based on PID real-time correction and fusion vision as claimed in claim 7, characterized in that: The pre-processing comprises: Voxel filtering is performed on the point cloud data, the RGB image is converted to a grayscale image and Gaussian filtering is performed, and low-pass filtering is used to remove high-frequency noise on the acceleration and angular velocity data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the synchronous positioning and mapping method based on PID real-time correction and fusion vision described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the synchronous positioning and mapping method based on PID real-time correction and fusion vision described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Cross slope offset suppression method and device, electronic equipment and storage medium
CN121133693A
Visual positioning navigation method and device based on robot
CN121432499A
Power transmission line cableway stock yard point pattern recognition method and system based on gradient compensation and road width dynamic correction, storage medium and computing equipment
CN121958981A