Cleaning method and device of robot and storage medium

By generating collection information of cleaning area and selecting cleaning area based on the congestion, the problem that existing cleaning robots are difficult to make the best decisions is solved, and more efficient cleaning efficiency is achieved.

CN120167848APending Publication Date: 2025-06-20HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317521.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing cleaning robots have difficulty making the best decisions based on various situations in the cleaning environment, resulting in poor cleaning efficiency.

Method used

By generating the collection information of the cleaning area, based on the congestion of the area to be cleaned, the current cleaning area is selected from each area to be cleaned, and the robot is controlled to clean the current cleaning area until all areas to be cleaned are cleaned.

Benefits of technology

It realizes adaptive selection of the current cleaning area according to the changes in the crowding degree of the cleaning environment, and responds to dynamic environmental changes in real time, improving cleaning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120167848A_ABST
    Figure CN120167848A_ABST
Patent Text Reader

Abstract

The invention discloses a robot cleaning method and device and a storage medium, and the method comprises the steps: responding to a robot to receive a cleaning task, and generating cleaning area set information, the cleaning area set information comprises a plurality of pieces of area information, and each piece of area information represents a to-be-cleaned area; based on the crowding degree of the to-be-cleaned areas, a current cleaning area is selected from the to-be-cleaned areas, and the robot is controlled to clean the current cleaning area; the step is executed again until all the to-be-cleaned areas are cleaned; wherein the congestion degree is obtained by analyzing a dynamic object in the to-be-cleaned area. By means of the mode, dynamic environment changes can be responded, and the cleaning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of robots, and particularly to a cleaning method, device and storage medium for a robot. Background Art

[0002] With the rapid development of domestic service robot technology and the increasing improvement of living standards, floor cleaning robots are increasingly used in daily household cleaning work, liberating people's hands to a certain extent.

[0003] However, at present, during the cleaning process, it is difficult for cleaning robots to make the best decisions on various situations of the cleaning environment and perform corresponding operations, resulting in poor cleaning efficiency. Summary of the Invention

[0004] To solve the above technical problems, the technical solution adopted in the present application is: to provide a cleaning method, device and storage medium for a robot, so as to at least solve the problem that in the related art, during the cleaning process, it is difficult for the robot to make the best decisions on various situations of the cleaning environment and perform corresponding operations, resulting in poor cleaning efficiency.

[0005] According to an embodiment of the present invention, a cleaning method for a robot is provided, including:

[0006] In response to the robot receiving a cleaning task, generating cleaning area set information, where the cleaning area set information includes a plurality of area information, and each piece of area information respectively represents a to-be-cleaned area;

[0007] Based on the congestion degree of the to-be-cleaned area, selecting the current cleaning area from each to-be-cleaned area, and controlling the robot to clean the current cleaning area; re-executing this step until all to-be-cleaned areas are cleaned;

[0008] Wherein, the congestion degree is obtained by analyzing dynamic objects in the to-be-cleaned area.

[0009] To solve the above technical problems, a technical solution adopted in the present application is: to provide an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the cleaning method for the robot in the above technical solution.

[0010] To solve the above technical problems, a technical solution adopted in the present application is: to provide a computer-readable storage medium, which is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the cleaning method for the robot in the above technical solution.

[0011] Through the above solution, the cleaning method of the robot provided by this application generates cleaning area set information including the areas to be cleaned in response to the robot receiving a cleaning task. Based on the congestion degree of the areas to be cleaned, where the congestion degree is obtained by analyzing dynamic objects in the areas to be cleaned, the current cleaning area is selected from each area to be cleaned, and the robot is controlled to clean the current cleaning area. The steps of repeatedly selecting the current cleaning area from each area to be cleaned based on the congestion degree of the areas to be cleaned and controlling the robot to clean the current cleaning area are repeated until all areas to be cleaned are completed. In this way, the robot can adaptively select the current cleaning area according to the change of the cleaning environment congestion degree to respond to the dynamic environment change in real time and improve the cleaning efficiency. Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0013] Figure 1 is a flowchart of an embodiment of the cleaning method of the robot provided by this application;

[0014] Figure 2 is a flowchart of an embodiment of the area cleaning process provided by this application;

[0015] Figure 3 is a flowchart of a specific embodiment of the cleaning method of the robot provided by this application;

[0016] Figure 4 is a schematic structural diagram of an embodiment of the electronic device provided by this application;

[0017] Figure 5 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed Embodiments

[0018] The following will further describe this application in detail in conjunction with the drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate this application, but do not limit the scope of this application. Similarly, the following embodiments are only partial embodiments of this application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0019] References to "embodiments" in this application mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0020] It should be noted that the terms "first", "second", etc. in this application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. can explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0021] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the cleaning method of the robot provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment includes:

[0022] S110: In response to the robot receiving a cleaning task, generate cleaning area set information.

[0023] The user issues a cleaning task according to the cleaning requirements, which usually includes mapless whole-house cleaning, map-based whole-house cleaning, zoned cleaning, and selected area cleaning. The robot receives the cleaning task issued by the user and generates cleaning area set information, and the cleaning area set information includes several area information, and each area information respectively represents a to-be-cleaned area.

[0024] In one embodiment, the area information of each to-be-cleaned area includes the contour information of the area, the cleaning status, the area attribute information, and the congestion information. Specifically:

[0025] (1) The contour information of the area consists of contour points, which refers to the shape of the physical area described by means of a set of coordinate points, polygon boundaries, grid data, etc. Exemplarily, for indoor environment cleaning, the contour information of the area can be described by the indoor environment map data constructed by SLAM (Simultaneous Localization and Mapping) technology. For outdoor environment cleaning, the contour information of the area can also be described by GPS coordinates.

[0026] (2) The cleaning status is used to record the cleaning treatment situation of the cleaning area. In one embodiment, it can be divided into four states: cleaned, unreachable, uncleaned, and skipped cleaning. In other embodiments, in addition to the aforementioned four states, it may also include status parameters such as the proportion of the cleaned area, the cleaning completion timestamp, and the cleaning task priority.

[0027] (3) The area attribute information is used to describe the physical characteristics and functional attributes of the area. In one embodiment, it may include area types such as kitchen, living room, bathroom, corridor, unknown, etc. In other embodiments, in addition to the area attribute information, it may also include floor material information such as carpet, tile, floor, etc., and obstacle type information included in the area such as furniture, steps, wires, etc.

[0028] (4) The crowding degree information reflects the real-time or short-term traffic status in the area. Usually, it can be obtained through sensor data or external system input, and may include dynamic parameters such as the flow density of people or pets, the moving trend of obstacles, or the probability of equipment traffic conflicts.

[0029] If the user issues a cleaning task for the first time and there is no historical cleaning area information in the cleaning system, the cleaning status of the area information of each area to be cleaned in the current time is set to uncleaned, the area attribute information is set to unknown, and the area crowding degree information is set to non-crowded state.

[0030] If the current cleaning task is not the first time and there is historical cleaning area information in the cleaning system, the area information in the current cleaning area set information is compared and integrated with the historical cleaning area information, so as to consider the historical cleaning task before the start of the current cleaning task, thereby improving the cleaning efficiency of the current cleaning task.

[0031] In one embodiment, historical cleaning area information is obtained. The historical cleaning area information includes the area information of each first historical cleaning area that the robot cleaned last time. For each area to be cleaned, the first historical cleaning area that is the same area as the area to be cleaned is found, and the area information of the found first historical cleaning area is used to update the area information of the area to be cleaned, where the update of the area information includes at least one of the cleaning status and the crowding degree.

[0032] Exemplarily, calculate the overlap degree between the contours of each area to be cleaned in the current cleaning area set information and the contours of each area in the historical cleaning area information. Among them, the area contour overlap degree can be comprehensively calculated from the coincidence point ratio of area contour points, area shape similarity, and intersection over union. If the overlap degree is greater than the preset contour overlap degree threshold, it is considered that the two cleaning areas being compared are the same area, and the historical cleaning area information is updated to the area information of the corresponding current area to be cleaned. The main updated area information is the reachability information in the cleaning status and the congestion information in the historical cleaning area in the previous cleaning.

[0033] In this way, considering multiple cleaning information during the cleaning process enables the robot to have the most prior decision-making during the current cleaning process, thereby improving the cleaning efficiency.

[0034] S120: Based on the congestion degree of the areas to be cleaned, select the current cleaning area from each area to be cleaned, and control the robot to clean the current cleaning area.

[0035] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the area cleaning process provided by this application. As Figure 2 shown, this embodiment includes:

[0036] S121: Analyze the cleaning status of each area to be cleaned in the current cleaning task, and determine the candidate cleaning areas.

[0037] The cleaning area set information generated based on the current cleaning task includes several area information, and each area information represents an area to be cleaned. Among them, the area information includes the cleaning status of the area to be cleaned, and the cleaning status mainly includes status information such as uncleaned, cleaned, skipped cleaning, reachable, and unreachable.

[0038] In one embodiment, at least one area to be cleaned is used as a cleanable area, and the cleaning status of each cleanable area is analyzed to determine the candidate cleaning areas or directly determine the current cleaning area. If there is a cleanable area with an uncleaned status, a candidate cleaning area is selected from the cleanable areas with an uncleaned status, and step S122 is further executed to determine the current cleaning area. Exemplarily, calculate the weighted result between the shortest planned distance from the robot to the cleanable area and the cleaning area of the cleanable area to obtain the comprehensive scheduling distance between the robot and each cleanable area with an uncleaned status, sort the cleanable areas with an uncleaned status according to the size of each comprehensive scheduling distance, and use the cleanable area at the preset position after sorting as the candidate cleaning area.

[0039] If the cleaning status of each cleanable area in the current cleaning task is the skipped cleaning status, then step S122 is executed to further determine the current cleaning area.

[0040] It should be noted that in other embodiments, a cleanable area can also be randomly selected from the areas to be cleaned, or all areas can be used as cleanable areas, or only the currently reachable areas can be used as cleanable areas. The specific selection can be made according to actual needs and is not limited here.

[0041] In one embodiment, if the cleaning status of each area to be cleaned includes information on whether the robot can currently reach, when at least one area to be cleaned is used as a cleanable area, the areas to be cleaned whose current cleaning status is not the cleaned status are found and used as the second target areas. For each second target area, using the connected domain map and the current pose of the robot, it is determined whether the robot can currently reach the second target area, and the second target areas that can be currently reached are used as cleanable areas.

[0042] In one implementation manner, after determining whether the robot can currently reach the second target area, in response to the second target area being currently unreachable and the cleaning status of the second target area not being the unreachable status, it means that the area was reachable in the historical cleaning task but the robot is unreachable for the current cleaning task. Therefore, in the current cleaning task, the cleaning status in the area information of the second target area can be updated to the unreachable status.

[0043] In response to the second target area being currently reachable and the cleaning status of the second target area being the unreachable status, it means that the target area was historically unreachable but the robot can reach it in the current cleaning task. Therefore, the cleaning status in the area information of the second target area is updated to the uncleaned status.

[0044] Further, the second target areas with the current cleaning status of the unreachable status are verified to update the cleaning status of the second target areas. If they are still unreachable, they are set to the cleaned status to avoid repeatedly judging whether the second target areas are reachable. In response to all the current cleaning statuses of the second target areas being the unreachable status, a regional exploration movement is performed on each second target area, and it is judged whether the second target area is reachable during the regional exploration movement. Exemplarily, the robot performs a regional exploration movement to the possible connected areas of the second target area. When it moves to the connected area and determines that the second target area is still unreachable, in response to determining that the judgment result of the second target area is unreachable, the area status of the second target area is set to the cleaned status. Otherwise, it is set to the uncleaned status.

[0045] After obtaining the current cleaning status of all the second target areas, the second target areas with the current cleaning status of the uncleaned status or the skipped cleaning status are used as cleanable areas.

[0046] S122: Select the current cleaning area from each area to be cleaned based on the congestion level of the area to be cleaned.

[0047] After determining the candidate cleaning areas through step S121, when the area information includes the congestion level, obtain the current congestion level of each candidate cleaning area. The congestion level is obtained by analyzing the dynamic objects in the area to be cleaned. Exemplarily, it can be obtained from the congestion level information of the same area in the historical area information, or can be obtained in real time through sensor data or input from an external system.

[0048] In one example, in response to the congestion level of the candidate cleaning area meeting the first congestion condition, for example, the congestion level value of the candidate cleaning area is less than the congestion threshold, then determine the candidate cleaning area as the current cleaning area. Exemplarily, control the robot to move to the candidate cleaning area, control the robot to perform area exploration movement, and collect the image data of the candidate cleaning area during the area exploration movement, analyze the dynamic objects in the image data to obtain the congestion level of the candidate cleaning area, and update the congestion level in the area information of the candidate cleaning area with the analyzed congestion level.

[0049] If the congestion level of the candidate cleaning area does not meet the first congestion condition, then update the cleaning status in the area information of the candidate cleaning area to the skipped cleaning status, and re - execute step S121 and its subsequent steps.

[0050] In one example, in response to the cleaning status of each cleanable area being the skipped cleaning status, select the current cleaning area from each cleanable area whose congestion level meets the second congestion condition. The second congestion condition is that the congestion level of the cleanable area is less than the congestion threshold or other preset congestion thresholds.

[0051] S123: Control the robot to move to the candidate cleaning area or the current cleaning area.

[0052] After determining the candidate cleaning area or the current cleaning area in the current cleaning task, control the robot to move to the candidate cleaning area or the current cleaning area, so as to control the robot to further determine the current cleaning area or clean the current cleaning area.

[0053] In one embodiment, a candidate cleaning area or the current cleaning area is used as the first target area, and an area scheduling point in the first target area is determined. Exemplarily, the point with the shortest planned distance from the current position of the robot in the first target area is used as the area scheduling point. Planning reference data is generated using the area scheduling point and the current pose of the robot, and the planning reference data is input into the first reinforcement learning network for path planning to obtain a planned path for the robot to move to the first target area, so as to control the robot to move to the first target area according to the planned path. Among them, the first reinforcement learning network can be an Actor-Critic (policy-value network) deep reinforcement learning network. Typical Actor-Critic series deep reinforcement learning networks include TRPO (Trust Region Policy Optimization), PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), etc. This network can be obtained through large-scale training in a simulation environment and transfer training in an actual environment. A specific embodiment is provided as follows:

[0054] S1231: After determining the area scheduling point of the first target area, obtain the global planning map of the current cleaning environment of the robot, the acquisition data of the sensors configured on the robot, and the current pose of the robot.

[0055] Among them, the global planning map can be directly obtained from the existing data of the system, or the SLAM algorithm can be used to explore and build a map in an unknown environment; the sensors configured on the robot include, but are not limited to, sensors configured on the robot device or in the environment, usually a depth camera, an RGB camera, and a radar, etc.; the current pose of the robot includes the current position of the robot and heading angle information, etc.

[0056] S1232: Generate a current static obstacle map according to the global planning map and the current pose of the robot; generate a current dynamic obstacle map using the acquisition data of the sensors and the current pose of the robot; generate a sub-area scheduling point based on the area scheduling point and the current position of the robot.

[0057] Among them, the current static obstacle map represents the obstacle information stored in the global planning map in the current cleaning environment; the current dynamic obstacle map represents the obstacle information in the current cleaning area obtained according to the acquisition data of each current sensor of the robot; the sub-area scheduling point is a series of sub-scheduling points for the robot to reach the area scheduling point from the current position.

[0058] S1233: Generate a scheduling plan graph using the global planning map, regional scheduling points, and sub-regional scheduling points. Extract features from the current static obstacle map, current dynamic obstacle map, and scheduling plan graph respectively to obtain the corresponding static obstacle features, dynamic obstacle features, and scheduling plan features. Integrate the static obstacle features, dynamic obstacle features, and scheduling plan features to obtain the first integrated feature as the planning reference data.

[0059] Among them, use a feature extraction network to extract features from the current static obstacle map, current dynamic obstacle map, and scheduling plan graph respectively. The feature extraction network includes, but is not limited to, convolutional neural networks such as ResNet and MobileNet, and can be specifically selected according to the hardware level; the static obstacle features, dynamic obstacle features, and scheduling plan features can be integrated by splicing, adding, or other methods.

[0060] S1234: Process the planning reference data using the first reinforcement learning network to output the type of local planner that matches the current environment, and use the local planning algorithm corresponding to the output type of the local planner to perform path planning to obtain the planned path for the robot to move to the first target area.

[0061] Among them, the optional local planning algorithms include traditional local planning algorithms such as DWA, TEB, and MPC, which are suitable for long-distance navigation movements; and local planning algorithms based on deep reinforcement learning, which are more suitable for navigation in dynamic and complex scenarios. In this way, by switching the local planning algorithm through the deep reinforcement learning model, the robot can select the optimal local planning strategy in the current environment, improving the efficiency, safety, and robustness of the scheduling.

[0062] The design of the reward value of the first reinforcement learning network can include at least one of the following: the first collision reward value, the first reward value for approaching the regional scheduling point or sub-regional scheduling point, the second reward value for reaching the regional scheduling point or sub-regional scheduling point, the third reward value for following the global planning movement, and the fourth reward value for maintaining a safe distance from obstacles.

[0063] S1235: Control the robot to move according to the planned path, and determine whether the robot has reached the current regional scheduling point. If it has reached, execute step S124 to control the robot to clean the current cleaning area; if not, return to execute step S1232 until it is determined that the robot has reached the current regional scheduling point.

[0064] S124: Control the robot to clean the current cleaning area.

[0065] In response to the robot reaching the area scheduling point of the current cleaning area, cleaning reference data is generated using the current pose of the robot and the environmental data of the current cleaning area. The cleaning reference data is input into the second reinforcement learning network for cleaning analysis to obtain a motion control instruction, and the motion control instruction is used to control the robot to perform cleaning in the current cleaning area.

[0066] In one embodiment, in response to the robot reaching the area scheduling point of the current cleaning area, the global planning map in the current cleaning environment of the robot, the acquisition data of the sensors configured on the robot, and the current pose of the robot.

[0067] According to the global planning map and the current pose of the robot, a regional static obstacle map is generated; using the environmental data collected from the current cleaning area by the sensors configured on the robot and the current pose of the robot, a regional dynamic obstacle map is generated; based on the cleaning map information and the current pose of the robot, a regional cleaning map is generated. Among them, the regional cleaning map is a map generated by the robot during the cleaning process according to the environmental information and the requirements of the cleaning task. If it is the first cleaning, the regional cleaning map can be obtained based on the environmental map constructed by the sensors and the SLAM algorithm. If it is not the first cleaning, the regional cleaning map can be obtained based on the regional cleaning map already stored in the system by the robot.

[0068] Feature extraction is respectively performed on the regional static obstacle map, the regional dynamic obstacle map, and the regional cleaning map to obtain the corresponding regional static obstacle features, regional dynamic obstacle features, and regional cleaning features; the regional static obstacle features, regional dynamic obstacle features, and regional cleaning features are fused to obtain a second fusion feature as the cleaning reference data.

[0069] The cleaning reference data is input into the second reinforcement learning network for cleaning analysis to obtain a motion control instruction, and the motion control instruction is used to control the robot to perform cleaning in the current cleaning area. In this way, the optimal motion instruction is output through the deep reinforcement learning model, achieving full coverage and edge cleaning, and improving the cleaning efficiency and coverage rate.

[0070] Among them, the second reinforcement learning network can be an Actor-Critic deep reinforcement learning network, which can be obtained through large-scale training in a simulation environment and transfer training in an actual environment; Exemplarily, the reward value of the second reinforcement learning network includes at least one of the following: a second collision reward value, a motion control behavior reward value, a cleaning reward value, and a wall-following reward value. Among them, the motion control behavior reward value is related to the motion actions of the robot during the cleaning process, and the motion actions include forward, stop, left turn, and right turn, etc.; the cleaning reward value is related to the cleaning state of the grids in the cleaning map; the wall-following reward value is related to the obstacle distance value during the edge-following process; the obstacle distance value during the edge-following process is the distance between the robot and the target-side obstacle when the distance between the robot and the physical wall edge is within a preset distance range.

[0071] S125: In response to the completion of cleaning of the current cleaning area, update the cleaning state of the current cleaning area to the cleaned state.

[0072] After the current cleaning area is cleaned, update the cleaning state of the current cleaning area to the cleaned state to indicate that the cleaning is completed. Then, re-execute steps S121 to S124 until all the to-be-cleaned areas are cleaned, and update the current cleaning area set information to the historical cleaning area information.

[0073] In one embodiment, in response to the completion of cleaning of all the to-be-cleaned areas, update the cleaned area set information after the current cleaning task update to the historical cleaning area information list, where the historical cleaning area information list is used to update the historical cleaning area set information in the system.

[0074] In one embodiment, in response to the completion of cleaning of all the to-be-cleaned areas, determine whether the number of the current historical cleaning area information list exceeds a preset threshold, and the size of the preset threshold can be set according to the hardware conditions of the robot or actual needs. If the number of the current historical cleaning area information list does not exceed the preset threshold, update the current cleaning area set information to the historical cleaning area information list. If the number of the current historical cleaning area information list exceeds the preset threshold, the historical cleaning area information in the historical cleaning area information list can be fused first. After fusion, the comprehensive cleaning area information of each historical cleaning area information is obtained, and the comprehensive cleaning area information is updated to the historical cleaning area set information, and the historical cleaning area information list is cleared to release the space for storing the historical cleaning area information. Among them, the fusion method of the historical cleaning area information can refer to the area information fusion update method in step S110, that is, based on the contour information of each cleaning area, judge the overlap degree between each cleaning area. When the overlap degree is higher than the contour overlap degree threshold, it is determined as the same area, and the area information of the same area can be fused and updated.

[0075] In addition, in the process of updating the set information of the cleaning areas after the current cleaning task is updated to the historical cleaning area information list, it is possible to first merge and empty the historical cleaning area information list and add the set information of the cleaning areas after the current cleaning task is updated. It is also possible to update the set information of the current cleaning areas to the historical cleaning area information list and then merge, and then empty the historical cleaning area information list. It is also possible to set the preset threshold to the maximum list number, remove the historical cleaning area information with the farthest time distance in the historical cleaning area information list, update the set information of the current cleaning areas to the historical cleaning area information list. The specific timing of emptying the historical cleaning area information list can be set according to actual needs and is not limited here.

[0076] To better illustrate the cleaning method of the robot provided in this application, taking the cleaning environment as an indoor scene as an example, please refer to Figure 3 , Figure 3 The following specific embodiments of the cleaning method of the robot are provided for exemplary illustration:

[0077] S1: Generate set information of cleaning areas according to the issued cleaning task.

[0078] S11: The robot device end receives the cleaning task issued by the user, such as issuing tasks such as cleaning the whole house without a map, cleaning the whole house with a map, area cleaning, and selected area cleaning, and generates set information of cleaning areas R region ={C i , s i , attr i , crowd i}, 1 ≤ i ≤ N.

[0079] Among them, N represents the number of cleaning areas in R region , C i represents the contour information of the i-th cleaning area, and the contour information is composed of contour points; s i represents the cleaning status of the i-th cleaning area (divided into four statuses: cleaned, inaccessible, not cleaned, and skipped cleaning); attr i represents the area attribute information of the i-th cleaning area (divided into attribute information such as kitchen, living room, bathroom, corridor, unknown, etc.); crowd i represents the dynamic crowding degree of the i-th cleaning area (for example, calculated from the regional pedestrian flow and pet movement).

[0080] When the cleaning task is issued for the first time, set the cleaning status of each cleaning area in the current cleaning to not cleaned, the room attribute information to unknown, and the regional dynamic crowding degree to non-crowded status.

[0081] S12: Obtain the historical cleaning area set information R history ={C i , si , attr i , crowd i}}, 1 ≤ i ≤ N history 。

[0082] Among them, N history represents the number of cleaned areas after the previous cleaning is completed. When N history = 0, it means that this cleaning is the first cleaning and there is no need to perform area information fusion, and step S2 is executed; otherwise, step S13 is executed to perform area cleaning information fusion.

[0083] S13: Compare and fuse the cleaning area set information R region with each cleaning area in the historical cleaning area set information R history .

[0084] Specifically, calculate the overlap ratio ratio contour of the two area contours. When ratio contour ≥ thres contour , it is considered that these two cleaning areas are the same area, and update the reachability information and crowding degree information in the historical cleaning area set information to the area information of the corresponding current cleaning area; otherwise, no operation is performed. Among them, thres contour is the area contour overlap ratio threshold parameter, which can take a value of 95% during actual use; the area contour overlap ratio ratio contour can be comprehensively calculated from the coincidence point ratio of area contour points, area shape similarity, and intersection over union.

[0085] S2: Judge whether the cleaning is completed according to the cleaning area and the connected domain map information.

[0086] Judge whether the current cleaning task is completed according to the cleaning area set information R region , the connected domain map map connect and the current pose P cur = {x, y, th} of the robot. Among them, x, y, th are the x coordinate, y coordinate, and heading angle in the map coordinate system respectively, and the connected domain map map connect is used to represent the passability of the planned robot for any area.

[0087] S21: Check the cleaning area information of the non-cleaned state according to the connected domain map map connect and the current pose P cur of the robot.

[0088] (1) If a cleaning area in the connected domain represents reachability, and the original cleaning state in the area information of this cleaning area is unreachable, update the cleaning state of this cleaning area to uncleaned;

[0089] (2) If a cleaning area within a connected region is characterized as unreachable, and the original cleaning status in the area information of this cleaning area is the status information of not cleaned or skipped cleaning, then update the cleaning status of this cleaning area to unreachable.

[0090] S22: Traverse the cleaning status information of each area information in the set information R of the current cleaning area region among them.

[0091] (1) If the cleaning status of each area information is cleaned, it is considered that this cleaning is over, and execute step S23 to update and iterate the cleaning area information;

[0092] (2) If the cleaning status of each area information is unreachable, execute step S24;

[0093] (3) If there is an area with a cleaning status of not cleaned or skipped cleaning among the cleaning statuses of each area information, execute step S3.

[0094] S23: The cleaning is over, and the robot returns to charge. Add the set information R of the current cleaning area region to the historical cleaning area information list list history ={R i}, 1≤i≤N list . Further update the historical cleaning area set information R through the historical cleaning area information list list history . history .

[0095] Among them, N list is the number of historical cleaning area information in the historical cleaning area information list. In actual use, it is usually taken as 5. The update and iteration steps of the historical cleaning area information are as follows: when N list ≤5, directly update the set information R of the current cleaning area to the historical cleaning area information list list region ; when N history >5, compare and integrate the area information of each cleaning area information in the historical cleaning area information list to obtain the comprehensive cleaning area information, and update the comprehensive cleaning area information to R list , and the fusion update method refers to step S13. Finally, clear the historical cleaning area information list. history

[0096] S24: Explore the door area of the unreachable area in turn according to the principle of optimal distance. When moving to the door area, judge whether the area is reachable. When the area is still unreachable, set the status of this area to cleaned; otherwise, set it to the status of not cleaned.

[0097] After the exploration of the door area is completed, when there are areas to be cleaned, step S3 is executed; otherwise, if there is no information about areas to be cleaned, step S23 is executed.

[0098] S3: Perform a comprehensive sorting of the areas to be cleaned to generate area scheduling points.

[0099] Sort each area to be cleaned in combination with the current position, congestion level, and area to be cleaned of the robot, and generate area scheduling point P goal 。

[0100] S31: Judge the set information R of the cleaning areas region to check if there are areas to be cleaned. When there are areas to be cleaned, step S32 is executed; otherwise, if there are only skipped cleaning areas, step S33 is executed.

[0101] S32: Calculate the comprehensive scheduling distance d from the current point P cur to each area to be cleaned. Sort the areas to be cleaned based on the comprehensive scheduling distance corresponding to each area to be cleaned, and select the area with the shortest comprehensive scheduling distance as the area to be cleaned. unclean

[0102] Select the point with the shortest planned distance from the current point P cur in this cleaning area as the area scheduling point P goal , and execute step S4. The comprehensive scheduling distance d unclean The calculation formula is: d unclean =α*d plan +β*area. Where d plan is the shortest planned distance from the current point to this cleaning area, area is the cleaning area of this area to be cleaned, and α and β are the proportional parameters for calculating the comprehensive scheduling distance, which can be adjusted according to the actual situation.

[0103] S33: Sort the skipped cleaning areas by congestion level, select the area with the lowest congestion level as the area to be cleaned, and select the point with the shortest planned comprehensive scheduling distance from the current point P cur in this area as the area scheduling point P goal , and execute step S4.

[0104] S4: Perform path planning movement based on reinforcement learning.

[0105] Perform path planning movement based on reinforcement learning according to the area scheduling point P goal , the current pose P of the robot cur , the global planning map map plan and the surrounding sensing information.

[0106] S41: According to the global planning map map planWith the current pose P of the robot cur , generate the current static obstacle map img static_obs ; Generate the current dynamic obstacle map img cur according to the surrounding sensor information (such as depth camera, RGB camera, and radar, etc.) and the current pose P of the robot dynamic_obs ; Generate the scheduling planning map img plan according to the global planning map map goal , the regional scheduling point P sub_goal and the sub-regional scheduling point P plan . Among them, the sub-regional scheduling point P sub_goal is generated by the regional scheduling point P goal and the current pose P of the robot cur .

[0107] S42: Input the img static_obs , img dynamic_obs and img plan generated in the previous step into their respective feature extraction networks to generate the corresponding features f static_obs , f dynamic_obs and f plan , and fuse the three features to generate the fused feature f conbine . Among them, the feature extraction network is a convolutional neural network, such as ResNet, MobileNet, etc., which can be specifically selected according to the hardware level, and the feature fusion operation can be splicing or addition, etc.

[0108] S43: Input the fused feature f conbine into the Actor-Critic deep reinforcement learning network, output the local planner type selection trigger signal, and select the local motion algorithm adapted to the current environment for local motion accordingly. Among them, the optional local motion algorithms include traditional local motion algorithms (DWA, TEB, MPC, etc.) and local motion algorithms based on deep reinforcement learning. Traditional local motion algorithms are more suitable for long-distance navigation motion, and deep reinforcement learning local motion algorithms are more suitable for dynamic and complex scene navigation.

[0109] Exemplarily, the reward value r total of this deep reinforcement learning network is specifically designed as shown in the following formulas (1)-(4). Among them, r collision is the collision reward value; r goal_reached is the reward value for reaching the target point; r approaching_goal is the reward value for approaching the target point, and δ sub_gaol is the difference in reaching the sub-target point within two unit time steps; r goal_plan is the reward value for following the global planning motion; r safe_dist is the reward value for the robot's safe distance from obstacles, Sdyn is the distance value of the robot from the dynamic obstacle.

[0110] r total = r collison + r goal_reached + r global_plan + r approaching_goal + r safe_dist (1)

[0111]

[0112]

[0113] S44: Determine whether the robot has reached the area scheduling point. When the robot reaches the area scheduling point, execute step S5; otherwise, execute step S41.

[0114] S5: Analyze the dynamic objects in the current cleaning area.

[0115] Analyze dynamic objects such as people flow and pets in the current cleaning area.

[0116] S51: Determine the cleaning status of the current cleaning area: (1) If the cleaning status is skipped, execute step S7; (2) If the cleaning status is not cleaned, execute step S52.

[0117] S52: The robot reaches the area scheduling point of the area to be cleaned and performs area exploration movement. During the exploration movement, video stream information is collected. Among them, the area exploration movement can be selected as the movement of multiple corner points of the area boundary, the movement of the center of the area partition, etc.

[0118] S53: Analyze the dynamic objects based on the collected video stream, and calculate the crowding degree crowd of the current cleaning area according to the movement of the dynamic objects.

[0119] S6: Determine whether the area is congested.

[0120] Judge whether the current cleaning area is congested through the crowding degree crowd of the current cleaning area. If the current cleaning area is congested, temporarily skip the cleaning of this cleaning area, set the status of this cleaning area to skipped, re-select the current cleaning area, and execute step S3; otherwise, execute step S7.

[0121] S7: Execute area cleaning based on reinforcement learning.

[0122] S71: According to the global planning map map plan and the current pose P of the robot cur generate the area static obstacle map img region_static_obs ; According to the surrounding sensor information and the current pose P of the robot cur, generate a dynamic obstacle map img for the area region_dynaimc_obs ; according to the cleaning map information map plan and the current pose P of the robot cur generate a cleaning map img for the area region_clean .

[0123] S72: Extract features of img region_static_obs , img region_dynaimc_obs , img region_clean respectively through their respective feature extraction networks, and then generate the cleaning feature f for the current area through the feature fusion module region .

[0124] S73: Input the cleaning feature f of the area region into the Actor-Critic deep reinforcement learning network, output a motion control instruction for area cleaning motion, and the area cleaning motion instruction includes forward, stop, left turn and right turn.

[0125] Exemplarily, the reward value r of this deep reinforcement learning network clean_total is specifically designed as shown in the following formulas (5)-(8), where r action is the reward for motion control behavior to supervise and improve the cleaning efficiency; r clean is the cleaning reward; r wall is the wall-following reward, which takes effect when the robot is close to the physical wall; state is the grid cleaning state in the cleaning map; right _ dist is the distance value of the obstacle on the right side of the robot.

[0126] r clean_total = r collision + r action + r safe_dist + r clean + r wall (5)

[0127]

[0128] S74: Judge whether the current cleaning area is cleaned up according to the cleaned information of the cleaning state in the area information. When the current cleaning area is cleaned up, set the cleaning state of this area to cleaned, and execute step S2; otherwise, execute step S71 and its subsequent steps.

[0129] Please refer to Figure 4 , Figure 4 is a schematic structural diagram of an embodiment of an electronic device provided by this application. The electronic device 60 includes a mutually connected memory 61 and a processor 62. The memory 61 is used to store a computer program, and when the computer program is executed by the processor 62, it is used to implement the cleaning method of the robot in the above embodiment.

[0130] For the method of the above embodiments, it may exist in the form of a computer program. Therefore, the present application proposes a computer-readable storage medium. Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. The computer-readable storage medium 80 is used to store a computer program 81, which can be executed to implement the cleaning method of the robot in the above embodiments.

[0131] The computer-readable storage medium 80 may be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0132] The above are only embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A robot cleaning method, characterized in that: The method comprises: In response to the robot receiving a cleaning task, generating cleaning area set information, wherein the cleaning area set information includes a plurality of area information, each of the area information respectively represents an area to be cleaned; Based on the congestion of the areas to be cleaned, a current cleaning area is selected from the areas to be cleaned, and the robot is controlled to clean the current cleaning area; this step is repeated until all the areas to be cleaned are cleaned; The congestion degree is obtained by analyzing the dynamic objects in the area to be cleaned.

2. The method according to claim 1, characterized in that: The area information includes the cleaning status of the area to be cleaned; The selecting a current cleaning area from each of the areas to be cleaned based on the congestion of the areas to be cleaned includes: Taking at least one of the to-be-cleaned areas as a cleanable area; Analyzing the cleaning status of each of the cleanable areas; In response to the cleaning state of the cleanable area being an uncleaned state, a candidate cleaning area is selected from the cleanable areas in the uncleaned state, and in response to the congestion degree of the candidate cleaning area meeting a first congestion condition, the candidate cleaning area is determined as a current cleaning area; and / or, in response to the cleaning state of each of the cleanable areas being a skip cleaning state, a current cleaning area whose congestion degree meets a second congestion condition is selected from each of the cleanable areas; After controlling the robot to clean the current cleaning area, the method further includes: In response to the current cleaning area being cleaned, the cleaning state of the current cleaning area is updated to a cleaned state.

3. The method according to claim 2, characterized in that The step of selecting a candidate cleaning area from the cleanable area in the uncleaned state comprises: Calculating a comprehensive dispatch distance between the robot and each of the cleanable areas in the uncleaned state, wherein the comprehensive dispatch distance is a weighted result of the shortest planned distance from the robot to the cleanable area and the cleaning area of ​​the cleanable area; Sorting the cleanable areas in the uncleaned state according to the size of the comprehensive scheduling distance, and taking the cleanable areas in the preset positions after sorting as the candidate cleaning areas; And / or, after selecting a candidate cleaning area from the cleanable area in the uncleaned state, the method further includes: In response to the congestion degree of the candidate cleaning area not meeting the first congestion condition, the cleaning status in the area information of the candidate cleaning area is updated to a skip cleaning status, and the analysis of the cleaning status of each of the cleanable areas and subsequent steps are re-executed.

4. The method according to claim 2, characterized in that: In response to the congestion degree of the candidate cleaning area meeting the first congestion condition, determining the candidate cleaning area as the current cleaning area includes: Controlling the robot to move to the candidate cleaning area; Controlling the robot to perform area exploration movement, and collecting image data of the candidate cleaning area during the area exploration movement; Analyzing the image data for dynamic objects to obtain the degree of congestion of the candidate cleaning area; In response to the analyzed congestion degree meeting the first congestion condition, determining the candidate cleaning area as the current cleaning area; and, In the case where the area information includes congestion, the congestion obtained by analysis is used to update the congestion in the area information of the candidate cleaning area; And / or, the area information further includes the congestion degree of the area to be cleaned; and the selecting, from each of the cleanable areas, a current cleaning area whose congestion degree satisfies a second congestion condition comprises: From each of the cleanable areas, select a current cleaning area whose congestion degree in the area information satisfies a second congestion condition; And / or, in response to the cleaning status of each of the cleanable areas being a skip cleaning status, after selecting a current cleaning area whose congestion degree satisfies a second congestion condition from each of the cleanable areas, and before controlling the robot to clean the current cleaning area, the method further includes: Control the robot to move to the current cleaning area.

5. The method according to claim 4, characterized in that Controlling the robot to move to the candidate cleaning area or the current cleaning area includes: Taking the candidate cleaning area or the current cleaning area as the first target area; determining a regional dispatch point in a first target area; Generate planning reference data using the regional scheduling point and the current position of the robot, and input the planning reference data into a first reinforcement learning network for path planning to obtain a planned path for the robot to move to the first target area; Control the robot to move to the first target area according to the planned path.

6. The method according to claim 5, characterized in that The regional scheduling point is the point in the first target area with the shortest planned distance from the current position of the robot; And / or, the generating planning reference data by using the regional scheduling point and the current posture of the robot includes: generating a current static obstacle map according to the global planning map and the current position of the robot; and, Generate a current dynamic obstacle map using the collected data of the sensor configured for the robot and the current position of the robot; and, Generate a sub-region scheduling point based on the regional scheduling point and the current position of the robot, and generate a scheduling planning map using the global planning map, the regional scheduling point and the sub-region scheduling point; Performing feature extraction on the current static obstacle map, the current dynamic obstacle map, and the scheduling planning map respectively to obtain corresponding static obstacle features, dynamic obstacle features, and scheduling planning features; The static obstacle feature, the dynamic obstacle feature and the scheduling planning feature are fused to obtain a first fused feature as the planning reference data; And / or, inputting the planning reference data into a first reinforcement learning network for path planning to obtain a planned path for the robot to move to the first target area includes: Processing the planning reference data using a first reinforcement learning network to output a local planner type that matches the current environment, performing path planning using a local planning algorithm corresponding to the output local planner type, and obtaining a planned path for the robot to move to the first target area; And / or, the reward value of the first reinforcement learning network includes at least one of the following: a first collision reward value, a first reward value for approaching an area scheduling point or a sub-area scheduling point, a second reward value for reaching an area scheduling point or a sub-area scheduling point, a third reward value for following the global planned motion, and a fourth reward value for being at a safe distance from an obstacle.

7. The method according to claim 2, characterized in that The step of taking at least one of the to-be-cleaned areas as a cleanable area comprises: Finding out the area to be cleaned whose current cleaning state is not the cleaned state as the second target area; For each of the second target areas, determining whether the robot can currently reach the second target area by using the connected domain map and the current position of the robot; The second target area that is currently reachable is used as the cleanable area.

8. The method according to claim 7, characterized in that After determining whether the robot can currently reach the second target area, the method further includes: In response to the second target area being currently unreachable and the cleaning state of the second target area being not an unreachable state, updating the cleaning state in the area information of the second target area to an unreachable state; In response to the second target area being currently reachable and the cleaning state of the second target area being an unreachable state, updating the cleaning state in the area information of the second target area to an uncleaned state; In response to the current cleaning status of all the second target areas being in an unreachable state, performing a regional exploration movement on each of the second target areas, and determining whether the second target area is reachable during the regional exploration movement process, and in response to determining that the determination result of the second target area is inaccessible, setting the regional status of the second target area to a cleaned state; The step of taking the currently reachable second target area as the cleanable area includes: The second target area whose current cleaning state is the uncleaned state or the skipped cleaning state is used as the cleanable area.

9. The method according to claim 1, characterized in that: Before the step of selecting a current cleaning area from each of the areas to be cleaned based on the congestion of the areas to be cleaned is performed for the first time, the method further includes: Acquire historical cleaning area information, wherein the historical cleaning area information includes area information of each first historical cleaning area cleaned by the robot last time; For each of the areas to be cleaned, the first historical cleaning area which is the same area as the area to be cleaned is found, and the area information of the area to be cleaned is updated using the area information of the found first historical cleaning area, wherein the update of the area information includes at least one of the following: update of the cleaning status and update of the congestion degree.

10. The method according to claim 9, characterized in that The method further comprises: In response to all the areas to be cleaned being cleaned, the cleaning area set information corresponding to the current cleaning task is updated to the historical cleaning area information list, and the historical cleaning area information list is used to update the historical cleaning area set information; and / or, In response to all the areas to be cleaned being cleaned, and the number of the historical cleaning area information list does not exceed a preset threshold, updating the current cleaning area set information to the historical cleaning area information list; and / or, In response to all the to-be-cleaned areas being cleaned, and the number of the historical cleaning area information lists exceeding the preset threshold, the historical cleaning area information in the historical cleaning area information lists are merged to obtain comprehensive cleaning area information; The comprehensive cleaning area information is updated to the historical cleaning area collection information, and the historical cleaning area information list is cleared.

11. The method according to claim 1, characterized in that: The controlling the robot to clean the current cleaning area includes: In response to the robot arriving at the current cleaning area, generating cleaning reference data using the current posture of the robot and environmental data of the current cleaning area; Inputting the cleaning reference data into a second reinforcement learning network for cleaning analysis to obtain motion control instructions; The motion control instruction is used to control the robot to perform cleaning in the current cleaning area.

12. The method according to claim 11, characterized in that The step of generating cleaning reference data by using the current posture of the robot and the environmental data of the current cleaning area includes: Generate a regional static obstacle map based on the global planning map and the current position of the robot; and, Generate a regional dynamic obstacle map using environmental data collected by sensors configured by the robot and the current position of the robot in the current cleaning area; and Generate an area cleaning map based on the cleaning map information and the current position of the robot; Performing feature extraction on the regional static obstacle map, the regional dynamic obstacle map, and the regional cleaning map respectively to obtain corresponding regional static obstacle features, regional dynamic obstacle features, and regional cleaning features; Fusing the regional static obstacle features, the regional dynamic obstacle features, and the regional cleaning features to obtain a second fused feature as the cleaning reference data; And / or, the reward value of the second reinforcement learning network includes at least one of the following: a second collision reward value, a motion control behavior reward value, a cleaning reward value, and a wall-following reward value, wherein the motion control behavior reward value is related to the movement actions of the robot during the cleaning process, and the movement actions include moving forward, stopping, turning left and turning right; the cleaning reward value is related to the grid cleaning state in the cleaning map; the wall-following reward value is related to the obstacle distance value of the edge-following process, and the obstacle distance value of the edge-following process is the distance between the robot and the target side obstacle when the distance between the robot and the physical wall is within a preset distance range.

13. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the processor is coupled to the memory, and the processor is configured to execute one or more steps of the robot cleaning method described in any one of claims 1 to 12 based on instructions stored in the memory.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the robot cleaning method according to any one of claims 1 to 12.