Cleaning robot obstacle avoidance decision-making method fusing deep reinforcement learning and fuzzy logic
By integrating deep reinforcement learning and fuzzy logic, a decision-making model based on multimodal convolutional neural network and long-term memory network is constructed, and obstacle avoidance strategies are optimized. The problem of inefficiency of traditional cleaning robots in complex environments is solved, and efficient and stable cleaning tasks are achieved.
Patent Information
- Application Number
- CN202510391051.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
When facing a complex and changing home environment, traditional cleaning robots find it difficult to adjust their movement strategies independently, resulting in inefficient cleaning and prone to trapping. Traditional motion control technology lacks self-learning and adaptability.
A clean robot obstacle avoidance decision-making method that integrates deep reinforcement learning and fuzzy logic, builds a decision model through multimodal convolutional neural network and long-term memory network, combines reinforcement learning mechanism and fuzzy logic module to optimize the decision model to adapt to changes in the dynamic environment and generate obstacle avoidance instructions.
Achieve efficient and stable cleaning tasks in complex environments, improving the obstacle avoidance and adaptability of the cleaning robot, ensuring that the robot can avoid obstacles in a timely manner and maintain stable movement.
Smart Images

Figure CN120276436A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot obstacle avoidance, and specifically to an obstacle avoidance decision-making method for a cleaning robot that integrates deep reinforcement learning and fuzzy logic. Background Art
[0002] With the continuous popularization of the concept of smart home, as an important part of home automation equipment, cleaning robots have gradually entered people's daily lives. With the characteristics of automated operation, cleaning robots not only greatly reduce people's housework burden, but also effectively improve the quality of life, and the market demand shows a rapid growth trend.
[0003] In the technical system of cleaning robots, motion control technology plays a decisive role in cleaning efficiency and cleaning effect. Early cleaning robots mainly adopted a random collision motion strategy, which was simple and rough, changing the motion direction by constantly colliding with obstacles. Although the cost was low, the cleaning path lacked planning, resulting in cleaning dead corners and extremely low cleaning efficiency.
[0004] To overcome the above problems, traditional motion control technologies emerged. Such technologies usually control the movement of cleaning robots based on preset rules, such as boundary detection, local path planning, etc. For example, by using infrared sensors and ultrasonic sensors to sense the surrounding environment, a simple obstacle avoidance function can be realized, and the cleaning path can be planned according to a predetermined algorithm. Although traditional motion control technologies can meet basic cleaning needs and significantly improve cleaning efficiency and effect, they still expose many limitations when facing complex home environments.
[0005] On the one hand, traditional methods are difficult to process complex and changing environmental information. In actual home scenarios, the shapes, sizes, and positions of obstacles are different and may change at any time. The fixed rules and limited sensor information relied on by traditional technologies are difficult to accurately and flexibly respond to these complex situations. On the other hand, traditional motion control technologies lack self-learning and adaptive capabilities. When the environment changes greatly or encounters new types of obstacles, the robot often cannot autonomously adjust its motion strategy, resulting in reduced cleaning efficiency and even possible entrapment.
[0006] With the rapid development of artificial intelligence technology, intelligent algorithms such as deep reinforcement learning and fuzzy logic provide new ideas for solving the motion control problem of cleaning robots. Deep reinforcement learning enables the robot to autonomously obtain the optimal motion strategy through continuous trial-and-error learning during the interaction with the environment. Fuzzy logic can process inaccurate and uncertain information, enabling the robot to make more reasonable decisions in complex environments. Integrating these two technologies is expected to significantly improve the adaptability and obstacle avoidance ability of cleaning robots in complex environments, thereby optimizing cleaning efficiency and effect. Summary of the invention
[0007] In response to the needs and shortcomings of current technological development, the present invention provides a cleaning robot obstacle avoidance decision-making method that integrates deep reinforcement learning and fuzzy logic.
[0008] The present invention provides a cleaning robot obstacle avoidance decision-making method integrating deep reinforcement learning and fuzzy logic, and the technical solution adopted to solve the above technical problems is as follows:
[0009] A cleaning robot obstacle avoidance decision method integrating deep reinforcement learning and fuzzy logic comprises the following steps:
[0010] S1. Model construction: The cleaning robot collects information about the surrounding environment through sensors and converts it into images or point cloud data with spatial structure; a decision model is constructed based on a multimodal convolutional neural network and a long short-term memory network. The decision model automatically extracts data features of the image or point cloud data and generates a preliminary obstacle avoidance strategy;
[0011] S2. Model optimization: With the help of reinforcement learning mechanism, the cleaning robot optimizes the decision model in the process of continuous interaction with the environment and continuous trial and error, so that it can adapt to the dynamically changing environment; at the same time, the fuzzy logic module is introduced to fuzzify the precise sensor data, and the fuzzy rule reasoning and defuzzification operations are used to further optimize the decision model and enhance the decision model's ability to cope with the uncertainty of environmental data;
[0012] S3. Model execution: The cleaning robot's sensors continuously collect surrounding environment data in real time and input it into a decision-making model that integrates deep reinforcement learning and fuzzy logic. The decision-making model generates obstacle avoidance instructions based on the latest environmental information, and controls the cleaning robot's execution components to avoid obstacles in time and maintain stable motion.
[0013] Optionally, the cleaning robot uses a variety of sensors to collect information about the surrounding environment and processes the information according to the characteristics of different types of sensor data, wherein:
[0014] The information obtained by the visual sensor is converted into a 224×224 pixel RGB image, which represents a rasterized environment map of the robot's surroundings;
[0015] The data obtained by the timing sensor is presented in the form of a vector and contains the current speed, heading angle, and the measurement value of the i-th distance sensor;
[0016] Subsequently, the collected data is preprocessed by normalization, noise filtering and data enhancement to facilitate use in subsequent decision-making models.
[0017] Further optionally, a multimodal convolutional neural network part for the decision model includes:
[0018] The convolution layer extracts the local features of obstacles and generates a feature map by sliding the convolution kernel over the 224×224 pixel RGB image data generated by the visual sensor.
[0019] The pooling layer is responsible for taking over the feature map output by the convolutional layer and performing dimensionality reduction on the data, thereby reducing the amount of computation while retaining key features.
[0020] The fully connected layer is responsible for integrating the output data of the pooling layer, and then deeply analyzing the integrated data to output the judgment results of the environment;
[0021] Based on the judgment results output by the fully connected layer, the decision model generates a preliminary obstacle avoidance strategy to determine whether there is an obstacle ahead and, if so, to infer the possible avoidance direction;
[0022] For the long short-term memory network part of the decision model, it processes the time series sensor data and mines the time series information and rules in the data;
[0023] The output results of the fully connected layer of the multimodal convolutional neural network and the output results of the long short-term memory network are fused, and the decision model outputs the action probability distribution. Based on this, a preliminary obstacle avoidance strategy is generated to determine whether there is an obstacle ahead and, if so, to infer the possible avoidance direction.
[0024] Optionally, the cleaning robot uses a reinforcement learning mechanism to optimize the decision model in the process of continuous interaction with the environment and continuous trial and error, so that it can adapt to the dynamically changing environment. This process specifically includes:
[0025] The cleaning robot relies on sensors to perceive the surrounding environment in real time, converts the collected information into a state vector, and incorporates it into the state space of the Markov decision process; the cleaning robot is regarded as an intelligent agent of reinforcement learning, and based on the current environment state vector, it selects an action to execute from the preset action space;
[0026] After the action is executed, the cleaning robot receives corresponding reward feedback: i) If it successfully avoids obstacles and completes the cleaning task efficiently, the cleaning robot will receive a positive reward; ii) If it collides with obstacles or gets into work difficulties, it will receive a negative reward. The reward feedback is quantified through the reward function of the Markov decision process, and the proximal strategy optimization algorithm is used to continuously adjust the parameters of the policy network in the decision model according to the reward feedback and advantage function, and gradually optimize the obstacle avoidance strategy.
[0027] With multiple interactions with the environment, the decision-making model of the cleaning robot continues to evolve, and it can make optimal action decisions when faced with various complex and changing environments, and adapt to the dynamically changing environment.
[0028] Further optionally, the involved state space is composed of position, angular velocity and linear velocity, and the obstacle distance vector, where the obstacle distance vector covers the 360° scanning data of the lidar, comprehensively describing the position, motion state of the robot in the environment and the distribution of surrounding obstacles;
[0029] The preset action space covers forward, left turn (α), right turn (α), and stop, where α is the turning angle of 0°-90°.
[0030] Further optionally, a fuzzy logic module is introduced to fuzzify the precise sensor data. Through fuzzy rule reasoning and defuzzification operations, the decision-making model is further optimized to enhance the ability of the decision-making model to handle the uncertainty of environmental data. This process specifically includes:
[0031] The cleaning robot collects precise data through sensors, and then maps the precise data to the corresponding fuzzy language variable space according to the preset membership function;
[0032] Based on the preset fuzzy rule base, reasoning is performed on the fuzzified results; among them, the preset fuzzy rule base is summarized according to the working experience of the robot and expert knowledge, covering the current state of the robot and the perceived environmental information;
[0033] The centroid method is used for defuzzification operation to convert the fuzzy reasoning result into a precise control action, and it is determined whether the cleaning robot should take emergency avoidance, deceleration avoidance, or continue to drive normally, realizing the further optimization of the decision-making model;
[0034] Through the above process, when facing uncertain environmental information, the cleaning robot can make reasonable and robust decisions based on reasonable inferences about the environmental information.
[0035] Preferably, the involved cleaning robot collects three kinds of precise data of distance, obstacle size, and its own current speed through sensors, and then maps the precise data to the corresponding fuzzy language variable space according to the preset membership function, where:
[0036] The cleaning robot maps the distance data to the sets of "near", "medium", and "far" according to the preset trapezoidal membership function;
[0037] The cleaning robot maps the obstacle size data to the sets of "small", "medium", and "large" according to the preset trapezoidal membership function;
[0038] The cleaning robot maps the current speed to the sets of "low", "medium", and "high" according to the preset trapezoidal membership function.
[0039] Optionally, the involved step S3 specifically includes:
[0040] When the cleaning robot is actually running, the sensors carried by it continuously and real-time collect the surrounding environment data;
[0041] The collected data is input into the optimized decision-making model. Based on the latest environmental information, the decision-making model quickly generates obstacle avoidance instructions;
[0042] The generated obstacle avoidance instructions are transmitted to the execution components of the cleaning robot to control the cleaning robot to avoid obstacles in time and maintain stable movement.
[0043] A method for obstacle avoidance decision-making of a cleaning robot that combines deep reinforcement learning and fuzzy logic according to the present invention has the beneficial effects compared with the prior art as follows:
[0044] The present invention constructs a decision-making model based on a multi-modal convolutional neural network and a long short-term memory network. By optimizing the decision-making model through reinforcement learning and fuzzy logic, the optimized decision-making model can output optimal obstacle avoidance instructions according to the latest environmental data, adjust the motion control parameters, and ensure that the robot can always efficiently and stably complete the cleaning task in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The appendix Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to make the technical solutions, the technical problems solved and the technical effects of the present invention clearer and more understandable, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.
[0047] Embodiment:
[0048] Combined with the appendix Figure 1 , this embodiment proposes a method for obstacle avoidance decision-making of a cleaning robot that combines deep reinforcement learning and fuzzy logic, which includes the following steps:
[0049] S1. Model construction: The cleaning robot collects the surrounding environment information through sensors and converts it into image or point cloud data with a spatial structure; constructs a decision-making model based on a multi-modal convolutional neural network and a long short-term memory network. The decision-making model automatically extracts the data features of the image or point cloud data and generates a preliminary obstacle avoidance strategy.
[0050] In this step, the cleaning robot uses a variety of sensors to collect information about the surrounding environment and processes it according to the characteristics of different types of sensor data. Specifically: for the information obtained by the visual sensor, it is converted into an RGB image of 224×224 pixels, representing the rasterized environmental map of the robot's surrounding environment; for the data obtained by the time-series sensor, it is presented in vector form and includes the current speed, orientation angle, and the measurement value of the i-th distance sensor. Subsequently, preprocessing operations such as normalization, noise filtering, and data augmentation are performed on the collected data for subsequent use by the decision-making model.
[0051] For the multi-modal convolutional neural network part of the decision-making model, it includes:
[0052] The convolutional layer, which scans the 224×224 pixel RGB image data generated by the visual sensor by sliding the convolutional kernel over the data and performs a convolutional operation with the data to extract the local features of the obstacle and generate a feature map;
[0053] The pooling layer is responsible for receiving the feature map output by the convolutional layer, reducing the dimensionality of the data, and retaining the key features while reducing the computational complexity;
[0054] The fully connected layer is responsible for integrating the output data of the pooling layer, then performing in-depth analysis on the integrated data, and outputting the judgment result of the environment;
[0055] Based on the judgment result output by the fully connected layer, the decision-making model generates a preliminary obstacle avoidance strategy, determines whether there is an obstacle ahead, and if there is an obstacle, infers the possible avoidance direction.
[0056] For the long short-term memory network part of the decision-making model, it processes the time-series sensor data and mines the time-series information and patterns in the data.
[0057] The output results of the fully connected layer of the multi-modal convolutional neural network and the output results of the long short-term memory network are fused. The decision-making model outputs the action probability distribution and generates a preliminary obstacle avoidance strategy based on this, determines whether there is an obstacle ahead, and if there is an obstacle, infers the possible avoidance direction.
[0058] S2. Model optimization: The cleaning robot uses the reinforcement learning mechanism to optimize the decision-making model in the process of continuously interacting with the environment and making mistakes, so that it can adapt to the dynamically changing environment; at the same time, a fuzzy logic module is introduced to fuzzify the precise sensor data, and through fuzzy rule reasoning and defuzzification operations, the decision-making model is further optimized to improve the ability of the decision-making model to cope with the uncertainty of environmental data. The specific description is as follows.
[0059] S2.1. The cleaning robot uses the reinforcement learning mechanism to optimize the decision-making model in the process of continuous interaction with the environment and continuous trial and error, so that it can adapt to the dynamically changing environment. This process specifically includes:
[0060] S2.1.1. The cleaning robot relies on sensors to perceive the surrounding environment information in real time, converts the collected information into a state vector, and incorporates it into the state space of the Markov decision process; the cleaning robot is regarded as an intelligent agent of reinforcement learning, and based on the current environment state vector, it selects an action to execute from the preset action space.
[0061] The state space is composed of position, angular velocity, linear velocity, and obstacle distance vector. The obstacle distance vector covers the 360° scanning data of the lidar, which comprehensively describes the robot's position, motion state, and surrounding obstacle distribution in the environment.
[0062] The preset action space A includes forward, left turn (α), right turn (α), and stop, where α is a turning angle of 0°-90°.
[0063] S2.1.2. After the action is executed, the cleaning robot obtains corresponding reward feedback: i) If it successfully avoids obstacles and completes the cleaning task efficiently, the cleaning robot will receive a positive reward; ii) If it collides with obstacles or gets into work difficulties, it will receive a negative reward; the reward feedback is quantified through the reward function of the Markov decision process, and the proximal policy optimization (PPO) algorithm is used to continuously adjust the parameters of the policy network in the decision model according to the reward feedback and advantage function, and gradually optimize the obstacle avoidance strategy.
[0064] S2.1.3. With multiple interactions with the environment, the decision-making model of the cleaning robot continues to evolve, and it can make optimal action decisions when facing various complex and changing environments, and adapt to the dynamically changing environment.
[0065] S2.2, introduce the fuzzy logic module to fuzzify the precise sensor data, and further optimize the decision model through fuzzy rule reasoning and defuzzification operations to enhance the ability of the decision model to cope with the uncertainty of environmental data. This process specifically includes:
[0066] S2.2.1. The cleaning robot collects precise data through sensors, and then maps the precise data to the corresponding fuzzy language variable space according to the pre-set membership function. Specifically, the cleaning robot collects three kinds of precise data, namely, distance, obstacle size, and current speed, through sensors, and then maps the precise data to the corresponding fuzzy language variable space according to the pre-set membership function, where:
[0067] The cleaning robot maps the distance data to the “near”, “middle” and “far” sets according to the pre-set trapezoidal membership function;
[0068] The cleaning robot maps the obstacle size data to the sets of "small", "medium", and "large" according to a preset trapezoidal membership function.
[0069] The cleaning robot maps the current speed to the sets of "low", "medium", and "high" according to a preset trapezoidal membership function.
[0070] S2.2.2. Infer the results after fuzzyfication based on a preset fuzzy rule base. The preset fuzzy rule base is summarized according to the working experience of the robot and expert knowledge, covering the current state of the robot and the perceived environmental information.
[0071] The specific preset fuzzy rule base is as follows:
[0072] Short distance + large obstacle + low speed → Emergency obstacle avoidance requires a sharp turn;
[0073] Short distance + large obstacle + medium speed → Quick response is required;
[0074] Approaching a large obstacle at high speed → Forced emergency stop or large-angle turn;
[0075] Approaching a medium obstacle at low speed → Medium detour;
[0076] Approaching a medium obstacle at high speed → Strong correction to avoid collision;
[0077] Small obstacle at low speed → Slight adjustment is sufficient;
[0078] Approaching a small obstacle at medium speed → Medium avoidance;
[0079] Approaching a small obstacle at high speed → Medium avoidance while taking efficiency into account;
[0080] Large obstacle at low speed in the middle distance → Anticipatory detour;
[0081] Dynamic large obstacle in the middle distance → Balance safety and efficiency;
[0082] Approaching a large obstacle in the middle distance at high speed → Strong correction;
[0083] Medium obstacle at low speed in the middle distance → Light avoidance;
[0084] Medium obstacle at medium speed → Standard obstacle avoidance strategy;
[0085] Approaching a medium obstacle at high speed → Decelerate in advance and detour;
[0086] Small obstacle at low speed in the middle distance → No major adjustment is required;
[0087] Approaching a small obstacle at medium speed → Slight avoidance;
[0088] Approaching a small obstacle at high speed in the medium distance → Medium avoidance;
[0089] Approaching a large obstacle at low speed in the long distance → No immediate response required;
[0090] Approaching a large obstacle at medium speed in the long distance → Light path planning;
[0091] Approaching a large obstacle at high speed in the long distance → Predictive adjustment;
[0092] Approaching a medium-sized obstacle at low speed in the long distance → Ignore or fine-tune;
[0093] Approaching a medium-sized obstacle at medium speed in the long distance → Maintain the original path;
[0094] Approaching a medium-sized obstacle at high speed in the long distance → Plan a detour in advance;
[0095] Approaching a small obstacle at low speed in the long distance → No avoidance required;
[0096] Approaching a small obstacle at medium speed in the long distance → Ignore;
[0097] Approaching a small obstacle at high speed in the long distance → Maintain the path, with minor correction.
[0098] S2.2.3. Use the centroid method for defuzzification operation to convert the fuzzy inference result into an exact control action, and determine whether the cleaning robot should take emergency avoidance, decelerate for avoidance, or continue to drive normally, so as to further optimize the decision-making model.
[0099] S2.2.4. Through the above process, when facing uncertain environmental information, the cleaning robot can make reasonable and robust decisions based on reasonable inferences about the environmental information.
[0100] S3. Model execution: The sensors of the cleaning robot continuously and real-time collect the surrounding environmental data, input it into the decision-making model that combines deep reinforcement learning and fuzzy logic. The decision-making model generates obstacle avoidance instructions according to the latest environmental information, and controls the execution components of the cleaning robot to avoid obstacles in time and maintain stable movement.
[0101] This step specifically includes:
[0102] S3.1. When the cleaning robot is actually running, the sensors it carries continuously and real-time collect the surrounding environmental data.
[0103] For example, the lidar scans 360° around the robot. With a high resolution of 1°, it carefully detects the distance information of obstacles in different directions. In an open indoor environment, the robot can accurately draw the environmental contour with the help of the lidar, providing basic data for subsequent decisions; the vision camera captures the video stream at a frame rate of 30fps, and through image recognition algorithms, quickly detects dynamic obstacles such as moving pets or pedestrians, supplementing the deficiencies of the lidar in detecting dynamic targets.
[0104] For the collected data, an Extended Kalman Filter (EKF) is used to align the distance data obtained by the lidar with the environmental map constructed by visual SLAM. This process can eliminate the errors between different sensor data and generate more accurate and comprehensive environmental information. Subsequently, it is input into the optimized decision-making model.
[0105] S3.2. The collected data is input into the optimized decision-making model, and the decision-making model quickly generates obstacle avoidance instructions based on the latest environmental information.
[0106] S3.3. The generated obstacle avoidance instructions are transmitted to the execution components of the cleaning robot to control the cleaning robot to avoid obstacles in time and maintain stable movement.
[0107] Once the sensor detects a sudden obstacle, the system immediately triggers local path replanning, updates the input state of the decision-making model, enables the robot to quickly adjust its action strategy, and continuously and safely completes the cleaning task.
[0108] In summary, by using the obstacle avoidance decision-making method for a cleaning robot that combines deep reinforcement learning and fuzzy logic of the present invention, the decision-making model constructed based on a multi-modal convolutional neural network and a long short-term memory network is optimized through reinforcement learning and fuzzy logic. The optimized decision-making model can output optimal obstacle avoidance instructions according to the latest environmental data and adjust the motion control parameters to ensure that the robot can always efficiently and stably complete the cleaning task in a complex environment.
[0109] The above application of specific examples has elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art of this technology without departing from the principle of the present invention shall fall within the patent protection scope of the present invention.
Claims
1. A collision avoidance decision-making method for a cleaning robot that integrates deep reinforcement learning and fuzzy logic, characterized in that The steps include: S1. Model construction: The cleaning robot collects information about the surrounding environment through sensors and converts it into images or point cloud data with spatial structure; a decision model is constructed based on a multimodal convolutional neural network and a long short-term memory network. The decision model automatically extracts data features of the image or point cloud data and generates a preliminary obstacle avoidance strategy; S2. Model optimization: With the help of reinforcement learning mechanism, the cleaning robot optimizes the decision model in the process of continuous interaction with the environment and continuous trial and error, so that it can adapt to the dynamically changing environment; at the same time, the fuzzy logic module is introduced to fuzzify the precise sensor data, and the fuzzy rule reasoning and defuzzification operations are used to further optimize the decision model and enhance the decision model's ability to cope with the uncertainty of environmental data; S3. Model execution: The cleaning robot's sensors continuously collect surrounding environment data in real time and input it into a decision-making model that integrates deep reinforcement learning and fuzzy logic. The decision-making model generates obstacle avoidance instructions based on the latest environmental information, and controls the cleaning robot's execution components to avoid obstacles in time and maintain stable motion.
2. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic according to claim 1, characterized in that The cleaning robot uses a variety of sensors to collect information about the surrounding environment and processes it according to the characteristics of different types of sensor data, including: The information obtained by the visual sensor is converted into a 224×224 pixel RGB image, which represents a rasterized environment map of the robot's surroundings; The data obtained by the timing sensor is presented in the form of a vector and contains the current speed, heading angle, and the measurement value of the i-th distance sensor; Subsequently, the collected data is preprocessed by normalization, noise filtering and data enhancement to facilitate use in subsequent decision-making models.
3. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic according to claim 2, characterized in that The multimodal convolutional neural network part for the decision model includes: The convolution layer extracts the local features of obstacles and generates a feature map by sliding the convolution kernel over the 224×224 pixel RGB image data generated by the visual sensor. The pooling layer is responsible for taking over the feature map output by the convolutional layer and performing dimensionality reduction on the data, thereby reducing the amount of computation while retaining key features. The fully connected layer is responsible for integrating the output data of the pooling layer, and then deeply analyzing the integrated data to output the judgment results of the environment; Based on the judgment results output by the fully connected layer, the decision model generates a preliminary obstacle avoidance strategy to determine whether there is an obstacle ahead and, if so, to infer the possible avoidance direction; For the long short-term memory network part of the decision model, it processes the time series sensor data and mines the time series information and rules in the data; The output results of the fully connected layer of the multimodal convolutional neural network and the output results of the long short-term memory network are fused, and the decision model outputs the action probability distribution. Based on this, a preliminary obstacle avoidance strategy is generated to determine whether there is an obstacle ahead and, if so, to infer the possible avoidance direction.
4. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic according to claim 1, characterized in that, The cleaning robot uses the reinforcement learning mechanism to optimize the decision-making model in the process of continuous interaction with the environment and continuous trial and error, so that it can adapt to the dynamically changing environment. This process specifically includes: The cleaning robot relies on sensors to perceive the surrounding environment information in real time, converts the collected information into a state vector, and incorporates it into the state space of the Markov decision process; regards the cleaning robot as an agent of reinforcement learning, and selects an action from the preset action space to execute based on the current environmental state vector; After the action is executed, the cleaning robot obtains corresponding reward feedback: i) If it successfully avoids obstacles and efficiently completes the cleaning task, the cleaning robot will receive a positive reward; ii) If it collides with an obstacle or gets stuck in a working dilemma, it will receive a negative reward; Quantify the reward feedback through the reward function of the Markov decision process, adopt the proximal policy optimization algorithm, and continuously adjust the parameters of the policy network in the decision-making model according to the reward feedback and the advantage function, gradually optimizing the obstacle avoidance strategy; With multiple interactions with the environment, the decision-making model of the cleaning robot continues to evolve, and it can make optimal action decisions when facing various complex and changing environments, realizing the adaptation to the dynamically changing environment.
5. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic, characterized in that, The state space is composed of position, angular velocity and linear velocity, and the obstacle distance vector, where the obstacle distance vector covers the 360° lidar scan data, comprehensively describing the position of the robot in the environment, its motion state and the distribution of surrounding obstacles; The preset action space covers forward, left turn (α), right turn (α), stop, where α is the turning angle from 0° to 90°.
6. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic according to claim 4, characterized in that, Introduce a fuzzy logic module to fuzzify the precise sensor data, and through fuzzy rule reasoning and defuzzification operations, further optimize the decision-making model and enhance the ability of the decision-making model to handle the uncertainty of environmental data. This process specifically includes: The cleaning robot collects precise data through sensors, and then maps the precise data to the corresponding fuzzy linguistic variable space according to the preset membership function; Based on the preset fuzzy rule base, reason about the fuzzified results; among them, the preset fuzzy rule base is summarized according to the working experience of the robot and expert knowledge, covering the current state of the robot and the perceived environmental information; Adopt the centroid method for defuzzification operation, convert the fuzzy reasoning result into a precise control action, and determine whether the cleaning robot should take emergency avoidance, decelerated avoidance, or continue to drive normally, realizing the further optimization of the decision-making model; Through the above process, when facing uncertain environmental information, the cleaning robot can make reasonable and robust decisions based on reasonable inferences about the environmental information.
7. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic, characterized in that, The cleaning robot collects three kinds of precise data: distance, obstacle size, and its own current speed through sensors, and then maps the precise data to the corresponding fuzzy linguistic variable space according to the preset membership function, where: The cleaning robot maps the distance data to the "near", "medium", and "far" sets according to the preset trapezoidal membership function; The cleaning robot maps the obstacle size data to the "small", "medium", and "large" sets according to the preset trapezoidal membership function; The cleaning robot maps the current speed to the "low", "medium", and "high" sets according to the preset trapezoidal membership function.
8. A method for obstacle avoidance decision-making of a cleaning robot integrating deep reinforcement learning and fuzzy logic according to claim 1, characterized in that The specific steps of step S3 include: When the cleaning robot is actually running, the sensors it carries continuously and real-time collect the surrounding environment data; The collected data is input into the optimized decision-making model, which quickly generates obstacle avoidance instructions based on the latest environmental information; The generated obstacle avoidance instructions are transmitted to the execution components of the cleaning robot to control the cleaning robot to avoid obstacles in time and maintain stable movement.