Sweeping robot control system based on reinforcement learning
Through the combination of multi-sensor combined positioning algorithm and stable target guidance deep Q learning, the problems of low environmental perception accuracy, insufficient positioning drift and adaptability in the sweeping robot control system are solved, and high-precision positioning and dynamic adaptability are improved.
Patent Information
- Application Number
- CN202510686685.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing sweeping robot control system has problems such as low environmental perception accuracy, insufficient positioning drift and adaptability.
The multi-sensor joint positioning algorithm is adopted to integrate the data of lidar, IMU and vision sensors, and the error and noise impact are reduced through multi-increment IMU pre-integration algorithm and least squares method optimization. At the same time, use stable targets to guide deep Q learning to generate optimal control strategies to achieve dynamic adaptation.
It improves the positioning accuracy and dynamic adaptability of the sweeping robot in complex environments, avoids positioning drift, enhances the adaptability of the robot under environmental changes, and improves the efficiency of cleaning path planning and obstacle avoidance strategies.
Smart Images

Figure CN120233780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a control system for a floor cleaning robot based on reinforcement learning. Background Art
[0002] With the rapid development of artificial intelligence and robotics technologies, floor cleaning robots, as a type of home automation device, have gradually become popular and play an increasingly important role in home cleaning. However, there are still some deficiencies in the existing control systems for floor cleaning robots. Firstly, although the existing systems are equipped with multiple sensors such as lidar, IMU, and vision sensors to improve environmental perception capabilities, the measurement errors and data noises of different sensors still have a negative impact on the positioning and navigation accuracy of the system. In particular, due to the noise and bias inherent in the IMU sensor itself, the estimation error of the robot's motion state will accumulate over time. Especially in long-term operation or complex environments, the robot is prone to positioning drift. Secondly, although the existing systems use artificial intelligence methods such as reinforcement learning to achieve dynamic decision-making, in practical applications, the robot often relies on pre-trained models and has weak adaptability in changing environments. Summary of the Invention
[0003] The present invention provides a control system for a floor cleaning robot based on reinforcement learning, aiming to solve the problems of low environmental perception accuracy, positioning drift, and insufficient adaptability existing in the prior art. Specifically, the environmental perception module of the present invention adopts a multi-sensor joint positioning algorithm to fuse the data of lidar, IMU, and vision sensors to estimate the motion state of the floor cleaning robot in real time. By introducing a multi-incremental IMU pre-integration algorithm and a least squares optimization method, the influence of errors and noises between sensors on the positioning accuracy is effectively reduced. Especially in long-term operation or complex environments, the problem of positioning drift caused by IMU data deviation is avoided. This algorithm can continuously provide accurate robot motion state information in a dynamic environment to ensure the real-time positioning accuracy of the robot. In the robot control module, the present invention uses stable target-guided deep Q-learning to generate an optimal control strategy. By obtaining the state information of the robot in real time and combining it with the deep Q-learning model, the robot can adaptively adjust its behavior decision according to environmental changes, thereby achieving more efficient cleaning path planning and obstacle avoidance strategies.
[0004] The present invention provides a control system for a floor cleaning robot based on reinforcement learning, which includes an environmental perception module, a robot control module, a floor cleaning robot, and an action execution module. The environmental perception module constructs a multi-sensor joint positioning algorithm through a multi-incremental IMU pre-integration algorithm, a least squares method, and a compensation method, and estimates the motion state of the floor cleaning robot in real time through the multi-sensor joint positioning algorithm to generate the state information of the real-time robot. A robot control module takes the real-time state information of the robot as input, and through stable target-guided deep Q-learning, generates an optimal control strategy; An action execution module generates control instructions according to the optimal control strategy, controls the behavior actions of the sweeping robot, generates feedback information, and feeds it back to the environment perception module.
[0005] Furthermore, the process of the environment perception module generating the real-time state information of the robot specifically includes the following: Step B1: Collect point cloud data through a lidar to construct an environmental map to obtain position information, collect IMU data through an IMU, and generate environmental information data; Step B2: Combine the position information and IMU data, and through the multi-incremental IMU pre-integration algorithm and the least squares method, obtain the robot motion state information; Step B3: According to the robot motion state information, use a compensation method to compensate for the rotational distortion of the IMU data, and use a constant velocity model to compensate for the translational distortion of the point cloud data to obtain corrected environmental information data; Step B4: Combine the robot motion state information and the corrected environmental information data to obtain the real-time state information of the robot.
[0006] Furthermore, Step B2 specifically includes: performing time integration on the IMU data through the multi-incremental IMU pre-integration algorithm to obtain motion state increments, constructing a measurement model based on the motion state increments, and using the least squares method to iteratively align the motion state increments and the position information to estimate the robot's motion state and obtain the robot motion state information; the motion state increments include displacement increments, velocity increments, rotation increments, attitude increments, and friction increments.
[0007] Furthermore, the process of the robot control module generating the optimal control strategy specifically includes the following steps: Step S1: Initialize the Q network, experience replay pool, and parameters of the stable target-guided deep Q-learning. The Q network includes an online Q network and a target Q network. The online Q network is responsible for calculating state-action values and is used for decision-making, and the target Q network is used to provide a stable training signal for the online Q network; Step S2: According to the real-time state information of the robot, initialize the current state of the robot, calculate the Q value of the robot action in the current state of the robot through the online Q network, and use the ε-greedy strategy to select the robot action, execute the robot action, generate an experience tuple and store it in the experience replay pool; the robot actions include moving direction, accelerating, decelerating, performing a cleaning operation, obstacle avoidance, and switching modes; Step S3: Set the batch size, sample mini-batch data from the experience replay pool according to the batch size, and calculate the target value of the mini-batch data; Step S4: Update the parameters of the online Q-network using the gradient descent algorithm according to the target value; update the parameters of the target Q-network using the asymmetric gradient target tracking method; Step S5: To avoid drifting too far and failing to follow the changes in the parameters of the online Q-network when updating the parameters of the target Q-network using the asymmetric gradient target tracking method, set an update period and periodically copy the parameters of the online Q-network to the target Q-network for periodic synchronization; this helps reduce the training oscillation caused by the instability of the target Q-network and makes the calculation of the target value smoother; Step S6: Iterate steps S2 - S5, gradually optimize the estimation of the state-action value by the online Q-network through the experience replay pool and gradient update, and ensure the stable convergence of the learning process of the online Q-network through the target Q-network, obtain the trained online Q-network, and use the trained online Q-network to generate an optimal control strategy for the robot to execute tasks.
[0008] Further, step S4 specifically includes: calculating the loss function of the online Q-network and the loss function of the target Q-network according to the target value; and using the gradient descent algorithm to update the parameters of the online Q-network for more accurate estimation of the state-action value; using the asymmetric gradient target tracking method to update the parameters of the target Q-network for smooth update and stable target value of the online Q-network; the asymmetric gradient target tracking method gradually updates the parameters of the target Q-network by calculating the difference between the online Q-network and the target Q-network; this method makes the update of the target Q-network smoother, avoids the target Q-network following the update of the online Q-network too quickly, and provides a stable target value to help the online Q-network for training.
[0009] Adopting the above solution, the beneficial effects of the present invention are as follows: The present invention realizes the combination of the multi-sensor joint positioning algorithm and the stable target-guided deep Q-learning, improving the positioning accuracy and dynamic adaptability of the sweeping robot in complex environments; firstly, by fusing the data of lidar, IMU, and vision sensors, the multi-sensor joint positioning algorithm of the present invention effectively estimates the motion state of the sweeping robot in real time; this technology can reduce the influence of sensor noise and errors on the positioning accuracy, especially solves the problem of error accumulation of the IMU sensor during long-term operation, thus effectively avoiding the common positioning drift phenomenon in traditional control systems; ensuring that the sweeping robot always maintains high-precision positioning in a dynamic environment, thereby improving the stability and reliability of its navigation and path planning; Secondly, the present invention generates a more efficient and self-learning control strategy by introducing stable target-guided deep Q-learning, which solves the limitations of traditional sweeping robot control systems in adapting to environmental changes. Traditional control methods often rely on pre-trained models and lack the ability to adapt to environmental changes in real time. In contrast, the present invention combines deep Q-learning with the real-time motion state information of the robot, enabling the robot to continuously optimize its decision-making strategy according to the actual situation in different environments. This method enhances the robot's adaptive ability in complex and changing environments, ensuring that the sweeping robot can efficiently handle different obstacles, cleaning tasks, and environmental mode switches, thereby greatly improving the working efficiency of the robot in actual use. In summary, by combining multi-sensor joint positioning and deep Q-learning, the present invention enables the robot to accurately locate in a complex environment and adjust its decision-making strategy in real time, avoiding the problems of unstable path planning and low cleaning efficiency faced by traditional sweeping robots. It greatly improves the robot's environmental perception ability and autonomous decision-making level, further promoting the development of sweeping robot technology towards intelligence and automation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a schematic diagram of the modules of a sweeping robot control system based on reinforcement learning provided by the present invention. Figure 2 It is a schematic diagram showing the effect of IMU data rotation and translation distortion compensation in step B3 of the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] Embodiment 1. According to Figure 1 , the present invention provides a sweeping robot control system based on reinforcement learning, which includes an environmental perception module, a robot control module, a sweeping robot, and an action execution module. The environmental perception module constructs a multi-sensor joint positioning algorithm through a multi-incremental IMU pre-integration algorithm, a least squares method, and a compensation method, and estimates the motion state of the sweeping robot in real time through the multi-sensor joint positioning algorithm to generate real-time robot state information. The robot control module takes the real-time robot state information as input and generates an optimal control strategy through stable target-guided deep Q-learning. The action execution module generates control instructions according to the optimal control strategy, controls the behavior actions of the sweeping robot, generates feedback information, and feeds it back to the environmental perception module.
[0013] Embodiment 2. According to Figure 2 , this embodiment is based on Embodiment 1. In this embodiment, the process of the environmental perception module generating the state information of the real-time robot specifically includes the following content: Step B1: Collect point cloud data through a lidar to construct an environmental map to obtain position information, collect IMU data through an IMU, and generate environmental information data; Step B2: Combine the position information and IMU data, and obtain the robot motion state information through the multi-incremental IMU pre-integration algorithm and the least squares method; Step B3: According to the robot motion state information, use a compensation method to compensate for the rotational distortion of the IMU data, and use a constant velocity model to compensate for the translational distortion of the point cloud data to obtain corrected environmental information data; Step B4: Combine the robot motion state information and the corrected environmental information data to obtain the state information of the real-time robot.
[0014] Embodiment 3. This embodiment is based on Embodiment 1. In this embodiment, the process of the environmental perception module generating the state information of the real-time robot specifically includes the following content: Step E1: Collect point cloud data through a lidar to construct an environmental map to obtain position information, collect IMU data through an IMU, and generate environmental information data; Step E2: Combine the position information and IMU data, and obtain the robot motion state information through the combination of the Kalman filter algorithm and the particle filter algorithm; Step E3: According to the robot motion state information, compensate for the rotational distortion of the IMU data, and use a constant velocity model to compensate for the translational distortion of the point cloud data to obtain corrected environmental information data; Step E4: Combine the robot motion state information and the corrected environmental information data to obtain the state information of the real-time robot.
[0015] Embodiment 4. This embodiment is based on Embodiment 2. In this embodiment, Step B2 specifically includes: performing time integration on the IMU data through the multi-incremental IMU pre-integration algorithm to obtain the motion state increment, constructing a measurement model based on the motion state increment, and using the least squares method to iteratively align the motion state increment and the position information to estimate the motion state of the robot to obtain the robot motion state information; the motion state increment includes displacement increment, velocity increment, rotation increment, attitude increment, and friction increment; In the IMU pre-integration algorithm, the displacement increment formula: ; Among them, represents the frame index, represents the displacement increment from the th frame to the th frame, represents the rotation matrix, represents the rotation matrix of the th frame, represents the translation vector of the th frame, represents the translation vector of the th frame, represents the difference in translation vectors between the th frame and the th frame, represents the gravitational acceleration vector, represents the time interval, represents the square term of the time difference, represents the velocity vector of the th frame; In the IMU pre-integration algorithm, the velocity increment formula: ; Among them, represents the velocity increment from the th frame to the th frame, represents the rotation matrix of the th frame, represents the velocity vector of the th frame; Measurement model formula: ; Among them, represents the target variable, that is, the measurement value from the th frame to the th frame, represents the transformation matrix of the time interval, represents the external transformation matrix from the LiDAR coordinate system to the IMU coordinate system; represents the attitude increment, represents the friction increment; Least squares method formula: ; Among them, represents the total number of frames, represents the Jacobian matrix, represents the robot motion state information.
[0016] Embodiment 5. This embodiment is based on Embodiment 4. In this embodiment, the process of the robot control module generating the optimal control strategy specifically includes the following steps: Step S1: Initialize the Q-network, experience replay pool, and parameters of the stable target-guided deep Q-learning. The Q-network includes an online Q-network and a target Q-network. The online Q-network is responsible for calculating the state-action value and is used for decision-making, and the target Q-network is used to provide a stable training signal for the online Q-network; Step S2: According to the real-time state information of the robot, initialize the current state of the robot. Calculate the Q-value of the robot action in the current state of the robot through the online Q-network, and use the ε-greedy strategy to select the robot action. Execute the robot action, generate an experience tuple, and store it in the experience replay pool; The robot actions include moving direction, accelerating, decelerating, performing a cleaning operation, obstacle avoidance, and switching modes; Step S3: Set the batch size, sample mini-batch data from the experience replay pool according to the batch size, and calculate the target value of the mini-batch data; Step S4: According to the target value, use the gradient descent algorithm to update the parameters of the online Q-network; use the asymmetric gradient target tracking method to update the parameters of the target Q-network; Step S5: In order to avoid drifting too far and not being able to keep up with the change of the parameters of the online Q-network when updating the parameters of the target Q-network using the asymmetric gradient target tracking method, set an update period, and periodically copy the parameters of the online Q-network to the target Q-network for periodic synchronization; This helps to reduce the training oscillation caused by the instability of the target Q-network and makes the calculation of the target value smoother; Step S6: Iterate Steps S2 - S5, gradually optimize the estimation of the state-action value by the online Q-network through the experience replay pool and gradient update, and ensure the stable convergence of the learning process of the online Q-network through the target Q-network, obtain the trained online Q-network, and use the trained online Q-network to generate the optimal control strategy for the robot to execute tasks.
[0017] Embodiment 6. This embodiment is based on Embodiment 4. In this embodiment, the process of the robot control module generating the optimal control strategy specifically includes the following steps: Step R1: Initialize the Q-network, experience replay pool, and parameters of the stable target-guided deep Q-learning. The Q-network includes an online Q-network and a target Q-network; Step R2: According to the real-time state information of the robot, initialize the current state of the robot. Calculate the Q-value of the robot action in the current state of the robot through the online Q-network, and use the ε-greedy strategy to select the robot action. Execute the robot action, generate an experience tuple, and store it in the experience replay pool; The robot actions include moving direction, accelerating, decelerating, performing a cleaning operation, obstacle avoidance, and switching modes; Step R3: Set the batch size, sample mini-batch data from the experience replay pool according to the batch size, and calculate the target value of the mini-batch data; Step R4: Update the parameters of the online Q-network using the gradient descent algorithm according to the target value; Step R5: Regularly copy the parameters of the online Q-network to the target Q-network for synchronization; Step R6: Iterate Steps R2 - R5 to obtain the trained online Q-network, and use the trained online Q-network to generate an optimal control strategy for the robot to execute tasks.
[0018] Embodiment 7. This embodiment is based on Embodiment 5. In this embodiment, Step S4 specifically includes: calculating the loss function of the online Q-network and the loss function of the target Q-network according to the target value; and using the gradient descent algorithm to update the parameters of the online Q-network for more accurately estimating the state-action value; using the asymmetric gradient target tracking method to update the parameters of the target Q-network for smooth updating and stabilizing the target value of the online Q-network; the asymmetric gradient target tracking method progressively updates the parameters of the target Q-network by calculating the difference between the online Q-network and the target Q-network; this method makes the update of the target Q-network smoother, avoids the target Q-network following the update of the online Q-network too quickly, and provides a stable target value to help the online Q-network for training; ; Among them, represents the parameters of the online Q-network, represents the learning rate, represents the mini-batch data, represents the loss function of the online Q-network; represents the gradient of the parameter ; ; Among them, represents updating the parameters of the target Q-network, represents the current state of the robot, represents the action of the robot, represents the estimated value of the online Q-network for the state and the action , represents the estimated value of the target Q-network for the state and the action , represents the scaling function, which dynamically adjusts the weight of the target Q-network update according to the difference between the online Q-network and the target Q-network; represents the gradient of the target Q-network.
[0019] The above description of the present invention and its implementation manners is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design, without creative efforts, structural manners and embodiments similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A control system for a floor cleaning robot based on reinforcement learning, including a floor cleaning robot; characterized in that: The system further includes an environment perception module and a robot control module; The environment perception module constructs a multi-sensor joint positioning algorithm, and generates the state information of the real-time robot through the multi-sensor joint positioning algorithm; The robot control module takes the state information of the real-time robot as input, and generates an optimal control strategy through stable target-guided deep Q learning to control the sweeping robot.
2. The control system of a floor cleaning robot based on reinforcement learning according to claim 1, characterized in that: Construct a multi-sensor joint positioning algorithm through the multi-incremental IMU pre-integration algorithm, least squares optimization and compensation method.
3. The control system of a floor cleaning robot based on reinforcement learning according to claim 2, wherein: The process of the environment perception module generating the state information of the real-time robot specifically includes the following contents: Step B1: Collect point cloud data to obtain position information, collect IMU data, and generate environmental information data; Step B2: Combine the position information and IMU data, and obtain the robot motion state information through the multi-incremental IMU pre-integration algorithm and the least squares method; Step B3: According to the robot motion state information, use the compensation method to perform rotation distortion compensation and translation distortion compensation on the environmental information data to obtain the corrected environmental information data; Step B4: Combine the robot motion state information and the corrected environmental information data to obtain the state information of the real-time robot.
4. The control system of a floor cleaning robot based on reinforcement learning according to claim 3, characterized in that: Step B2 specifically includes: performing time integration on the IMU data through the multi-incremental IMU pre-integration algorithm to obtain the motion state increment, constructing a measurement model, and using the least squares method to iteratively align the motion state increment and the position information to obtain the robot motion state information.
5. The control system of a floor sweeping robot based on reinforcement learning according to claim 4, characterized in that: The motion state increment includes displacement increment, velocity increment, rotation increment, attitude increment and friction increment.
6. The control system of a floor-sweeping robot based on reinforcement learning according to claim 1, characterized in that: The process of the robot control module generating the optimal control strategy specifically includes the following steps: Step S1: Initialize the Q network, experience replay pool and parameters of the stable target-guided deep Q learning. The Q network includes an online Q network and a target Q network; Step S2: According to the state information of the real-time robot, initialize the current state of the robot, select and execute the robot action through the online Q network and the ε-greedy strategy, and generate an experience tuple and store it in the experience replay pool; Step S3: Sample mini-batch data from the experience replay pool and calculate the target value of the mini-batch data; Step S4: According to the target value, use the gradient descent algorithm to update the parameters of the online Q network; use the asymmetric gradient target tracking method to update the parameters of the target Q network; Step S5: Set the update period, and periodically copy the parameters of the online Q network to the target Q network for periodic synchronization; Step S6: Iterate steps S2 - S5 to obtain the trained online Q network, and use the trained online Q network to generate the optimal control strategy.
7. The control system of a floor-sweeping robot based on reinforcement learning according to claim 6, characterized in that: Use the asymmetric gradient target tracking method to update the parameters of the target Q network, and smoothly update and stabilize the online Q network.
Citation Information
Patent Citations
VIO rapid united initialization method based on monocular camera
CN108981693A
Pose estimation method based on RGB-D and IMU information fusion
CN109993113A
Electric power inspection robot positioning method based on multi-sensor fusion
CN111739063A
Mobile robot local path planning method based on value distribution deep reinforcement learning
CN117470244A
Financial commodity recommendation method and system based on big data analysis
CN118586988A