Autonomous operation method of robot

By generating globally optimal patrol paths, acquiring panoramic images and determining target world coordinates, designing systematic cleaning sub-tasks and monitoring status, the sanitation robot has achieved fully autonomous operation throughout the entire process, solving several defects in existing technologies and improving the precision and intelligence of operations.

CN122108124APending Publication Date: 2026-05-29CHENGDU HUMANOID ROBOT INNOVATION CENT CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU HUMANOID ROBOT INNOVATION CENT CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing sanitation robots suffer from a lack of global perspective in path planning, low accuracy in target recognition and positioning, unsystematic design of garbage collection operations, poor continuity of cyclical operations, low level of intelligence in status monitoring and triggering of auxiliary tasks, and imperfect task completion processes, making it difficult to meet the needs of refined sanitation operations.

Method used

By generating globally optimal patrol paths, combining panoramic image acquisition and target world coordinate determination, a systematic cleaning sub-task design is implemented. The status is monitored in real time and the dumping/charging task is triggered, and the task completion process is standardized to form a closed-loop collaborative operation.

Benefits of technology

It has enabled sanitation robots to operate autonomously throughout the entire process from path planning to task completion, improving the precision and intelligence of operations and solving problems such as lack of global optimality in path planning, low target positioning accuracy, unsystematic cleaning operations, poor continuity of cyclical operations, and imperfect completion processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122108124A_ABST
    Figure CN122108124A_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot autonomous operation method, the method includes: receiving environmental sanitation task and obtaining task information and global map, global optimal patrol path of covering target area and series key node is generated;Robot is controlled along path navigation, panoramic image is collected to identify target object and determine its world coordinate;When recognizing garbage, according to its coordinate, robot positioning and category, generate and execute cleaning subtask, and store classified garbage;After completing cleaning, return to original path and circulate operation, while monitoring power and storage box capacity;Capacity reaches threshold when navigating to the nearest garbage recycling point to dump, power is lower than threshold when navigating to the nearest charging point to charge;After task is completed, clean remaining garbage, supplement power and report task completion information.The method realizes the whole process of environmental sanitation operation autonomous operation, and improves operation refinement and intelligent level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and in particular to a method for autonomous operation of a robot. Background Technology

[0002] The construction of smart cities is driving the upgrading of sanitation operations towards intelligence and automation. Sanitation robots, due to their advantages such as continuous operation and reduced labor costs, have been gradually applied to garbage collection in various scenarios. Although existing sanitation robots have basic navigation and positioning, environmental perception, garbage identification and grasping, and power monitoring functions, there are still many technical deficiencies in actual operation, making it difficult to meet the needs of refined sanitation operations.

[0003] First, the path planning lacks a global perspective, often relying on preset fixed paths. It fails to combine task information and a global map to generate the optimal patrol path, which can easily lead to blind spots in operations and omissions of key nodes (such as garbage collection points and charging points).

[0004] Secondly, the target recognition and positioning accuracy is insufficient. It mainly collects local environmental images, lacks panoramic perception capabilities, and can only obtain the relative position of the garbage, but cannot determine its world coordinates, which affects subsequent cleanup operations.

[0005] Third, the garbage collection lacks a systematic sub-task design, fails to plan collection actions based on garbage type, location, and the robot's current position, and often fails to achieve garbage sorting and storage, thus not meeting operational requirements;

[0006] Fourth, the continuity of cyclical operations is poor, and it is difficult to accurately return to the original patrol route after the garbage is cleaned up, which easily interrupts the operation process;

[0007] Fifth, the level of intelligence in status monitoring and triggering of ancillary tasks is low, the monitoring of storage box capacity is incomplete, and many tasks such as dumping and charging are not guided to the nearest location, resulting in low work efficiency.

[0008] Sixth, the task completion process is incomplete; it does not automatically complete the disposal of remaining waste and recharging, requiring manual intervention and increasing operation and maintenance costs.

[0009] Therefore, it is particularly necessary to provide an autonomous operation method for robots to address the problems of existing sanitation robots, such as lack of globality in path planning, low accuracy in target recognition and positioning, unsystematic design of garbage cleaning operations, poor continuity of cyclical operations, low level of intelligence in status monitoring and triggering of auxiliary tasks, and imperfect task completion processing procedures. This would improve the precision and intelligence of autonomous operation of sanitation robots and enable them to operate autonomously throughout the entire sanitation process. Summary of the Invention

[0010] The purpose of this invention is to provide an autonomous operation method for sanitation robots to address the aforementioned problems. This method can solve the problems of existing sanitation robots, such as lack of globality in path planning, low accuracy in target recognition and positioning, unsystematic design of garbage cleaning operations, poor continuity of cyclical operations, low level of intelligence in status monitoring and triggering of auxiliary tasks, and imperfect task completion processing procedures.

[0011] The technical solution adopted in this invention is as follows:

[0012] An autonomous operation method for a robot includes the following steps:

[0013] Step S100, Task Reception and Path Determination: Receive sanitation tasks, obtain task information and global map corresponding to the sanitation tasks, and generate a globally optimal patrol path that covers the target area and connects key nodes based on the task information and global map.

[0014] Step S200, Path Patrol and Target Recognition: Control the robot to navigate along the globally optimal patrol path; during the navigation process, collect panoramic images of the robot's surrounding environment, identify target objects, and determine the world coordinates of the target objects in the working environment;

[0015] Step S300, Cleaning Task Generation and Execution: If the identified target object is garbage to be cleaned, a cleaning sub-task is generated and executed based on the world coordinates of the target object, the current location of the robot, and the category of the target object. The cleaning sub-task includes: planning a sub-path to navigate to the target object, performing a grasping operation, and storing the garbage according to its category into the corresponding storage box of the robot body.

[0016] Step S400, Cyclic Operation and Status Judgment: After completing the cleaning sub-task, control the robot to return to the globally optimal patrol path and continue to execute steps S200 and S300; at the same time, monitor the robot's power status and the storage box's capacity status in real time.

[0017] Step S500, Dumping Task Triggering and Execution: When the capacity status of the storage box is detected to reach a preset threshold, a dumping sub-task is generated, and the robot is controlled to navigate to the nearest garbage collection point to perform garbage dumping;

[0018] Step S600, Charging Task Triggering and Execution: When the robot's battery level is detected to be below a preset threshold, a charging sub-task is generated, and the robot is controlled to navigate to the nearest charging point to perform charging.

[0019] Step S700, Task Completion Processing: When the sanitation task is completed or the termination conditions are met, control the robot to go to the garbage collection point to dump the remaining garbage in the storage box, then go to the charging point to charge, and report the task completion information.

[0020] Furthermore, step S100 includes:

[0021] Step S101: Obtain the task information and global map corresponding to the sanitation task; the task information includes the global operation area and the target area located within the global operation area; the global map contains multiple preset key nodes;

[0022] Step S102: Based on the global operation area and the target area, the global map is rasterized to obtain a raster map; attribute information is configured for each grid in the raster map, the attribute information including at least terrain type, target area marker for indicating whether it belongs to the target area, and key node marker for marking key nodes; based on the key node markers and the global operation area, a set of valid key nodes located within the global operation area is selected;

[0023] Step S103: Identify the target area from the grid map based on the target area marker, and calculate the task demand intensity of each grid in the target area according to historical task data; divide the target area into multiple sub-areas of different importance levels according to the task demand intensity; perform coverage path planning using different coverage parameters for sub-areas of different importance levels, and generate a set of target area coverage path segments.

[0024] Step S104: Spatial clustering of the set of effective key nodes and the set of target area coverage path segments is performed to obtain multiple clusters; for each cluster, a shortest intra-cluster sub-path is planned to connect all effective key nodes and target area coverage path segments in the cluster; then, an inter-cluster path is planned to connect adjacent clusters, and all the shortest intra-cluster sub-paths are connected to form an initial path framework.

[0025] Step S105: Divide the initial path framework into multiple path segments, and label each path segment with its path type and path information; wherein, the path type is determined according to the terrain type of the grid through which the path segment passes; verify each path segment based on preset accessibility constraints, and perform local replanning on path segments that do not meet the constraints to obtain a feasible path that meets all accessibility constraints.

[0026] Step S106: Construct a comprehensive cost function to evaluate the feasible path, and with the goal of minimizing the comprehensive cost, optimize the feasible path under the preset global constraints to obtain the globally optimal patrol path.

[0027] Step S107: Determine whether the globally optimal patrol path satisfies the operation time limit constraint; if not, adjust the weight coefficient of the comprehensive cost function and re-execute the optimization solution process of step S106 until the globally optimal patrol path that satisfies the operation time limit constraint is obtained.

[0028] Furthermore, step S200 includes:

[0029] Step S201: Image acquisition, acquiring panoramic images of the environment surrounding the robot;

[0030] Step S202: Image enhancement. Perform enhancement preprocessing on the panoramic image targeting small targets to obtain a preprocessed image.

[0031] Step S203: Extraction of Region of Interest. The preprocessed image is semantically segmented to extract the preset region of interest, and the region of interest is divided into multiple sub-images.

[0032] Step S204: Collaborative recognition. Perform multi-level collaborative recognition on each sub-image to determine whether it contains a target object and to determine its category. The collaborative recognition includes preliminary detection based on a convolutional neural network, and cross-modal semantic verification by calling a multimodal large model based at least on the preliminary detection results.

[0033] Step S205: Target localization. Based on the position of the target object in the panoramic image and the robot's localization information, calculate the world coordinates of the target object in the working environment.

[0034] Furthermore, in step S300, generating and executing the cleanup subtask specifically includes:

[0035] Based on the world coordinates of the target object and the current location of the robot, plan a sub-path from the current location to the target object;

[0036] Control the robot to move along the sub-path to the target object;

[0037] Based on the category of the target object, the robot arm adapted to the robot body is invoked to perform a grasping operation;

[0038] The robotic arm is controlled to store the grasped target object into the classification bin corresponding to the target object's category in the robot's storage box.

[0039] Furthermore, in step S400, the real-time monitoring of the robot's battery status and the storage box's capacity status specifically includes:

[0040] The robot's power management system periodically acquires the remaining battery power percentage as the power status.

[0041] The current waste weight or filling rate is obtained by using a weighing sensor or vision sensor installed in the storage box, which serves as the capacity status.

[0042] Furthermore, step S500 specifically includes:

[0043] When the capacity of the storage box reaches a preset full-load threshold, a dumping task is triggered.

[0044] Based on the robot's current location and the location information of all garbage collection points in the global map, a first navigation path to the nearest garbage collection point is planned;

[0045] The robot is controlled to move along the first navigation path, and upon arriving at the waste collection point, the storage box is controlled to perform an emptying operation.

[0046] Furthermore, step S600 specifically includes:

[0047] When the robot's battery level falls below a preset low battery threshold, a charging task is triggered.

[0048] Based on the robot's current location and the location information of all charging points in the global map, a second navigation path to the nearest charging point is planned;

[0049] The robot is controlled to move along the second navigation path, and upon arrival at the charging point, it automatically docks with the charging interface and performs the charging process.

[0050] Furthermore, in step S700, the conditions for completion include: reaching the preset total operation time of the sanitation task, or not identifying any more garbage to be cleaned during the complete traversal of the globally optimal patrol path.

[0051] Furthermore, step S700 specifically includes:

[0052] Step S701, Final Disposal: Control the robot to navigate to the nearest waste collection point and empty all the waste in the storage bin;

[0053] Step S702, Final Charging: Control the robot to navigate from the waste collection point to the nearest charging point and perform a charging operation until the battery level reaches a preset saturation value;

[0054] Step S703, Information Reporting: Generate an execution report for the sanitation task and send it to the cloud. The execution report shall include at least the task identifier, total operation time, total amount of garbage cleaned, and final status.

[0055] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0056] This invention addresses the limitations of fixed paths by generating a globally optimal patrol path that covers the target area and connects key nodes through task reception and path determination; it achieves precise target identification and positioning through panoramic image acquisition and target world coordinate determination, solving the problem of local image recognition deviation; it systematically designs cleaning sub-tasks, combining waste type and location to plan the grabbing and classified storage process; it ensures that the waste returns to the original path after cleaning and monitors the power and capacity status in real time through cyclical operation and status monitoring; it triggers dumping / charging tasks through thresholds, navigating to the nearest location to perform the operation; and it completes the dumping of remaining waste, recharging, and information reporting through a standardized task completion process.

[0057] Each step forms a closed-loop collaboration, enabling sanitation robots to operate autonomously throughout the entire process from path planning, target recognition, garbage collection to task completion. This significantly improves the precision and intelligence of operations and completely solves the problems of existing technologies, such as lack of global optimality in path planning, low target positioning accuracy, unsystematic cleaning operations, poor continuity of cyclical operations, low intelligence of auxiliary tasks, and imperfect completion processes. Attached Figure Description

[0058] Figure 1 This is a flowchart of the present invention; Figure 2 This is a flowchart of step S100; Figure 3 This is a flowchart of step S200. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention. Figure 1 As shown, this invention discloses an autonomous operation method for a robot, comprising the following steps:

[0061] Step S100, Task Reception and Path Determination: Receive sanitation tasks, obtain task information and global map corresponding to the sanitation tasks, and generate a globally optimal patrol path that covers the target area and connects key nodes based on the task information and global map.

[0062] Step S200, Path Patrol and Target Recognition: Control the robot to navigate along the globally optimal patrol path; during the navigation process, collect panoramic images of the robot's surrounding environment, identify target objects, and determine the world coordinates of the target objects in the working environment;

[0063] Step S300, Cleaning Task Generation and Execution: If the identified target object is garbage to be cleaned, a cleaning sub-task is generated and executed based on the world coordinates of the target object, the current location of the robot, and the category of the target object. The cleaning sub-task includes: planning a sub-path to navigate to the target object, performing a grasping operation, and storing the garbage according to its category into the corresponding storage box of the robot body.

[0064] Step S400, Cyclic Operation and Status Judgment: After completing the cleaning sub-task, control the robot to return to the globally optimal patrol path and continue to execute steps S200 and S300; at the same time, monitor the robot's power status and the storage box's capacity status in real time.

[0065] Step S500, Dumping Task Triggering and Execution: When the capacity status of the storage box is detected to reach a preset threshold, a dumping sub-task is generated, and the robot is controlled to navigate to the nearest garbage collection point to perform garbage dumping;

[0066] Step S600, Charging Task Triggering and Execution: When the robot's battery level is detected to be below a preset threshold, a charging sub-task is generated, and the robot is controlled to navigate to the nearest charging point to perform charging.

[0067] Step S700, Task Completion Processing: When the sanitation task is completed or the termination conditions are met, control the robot to go to the garbage collection point to dump the remaining garbage in the storage box, then go to the charging point to charge, and report the task completion information.

[0068] This invention addresses the limitations of fixed paths by generating a globally optimal patrol path that covers the target area and connects key nodes through task reception and path determination; it achieves precise target identification and positioning through panoramic image acquisition and target world coordinate determination, solving the problem of local image recognition deviation; it systematically designs cleaning sub-tasks, combining waste type and location to plan the grabbing and classified storage process; it ensures that the waste returns to the original path after cleaning and monitors the power and capacity status in real time through cyclical operation and status monitoring; it triggers dumping / charging tasks through thresholds, navigating to the nearest location to perform the operation; and it completes the dumping of remaining waste, recharging, and information reporting through a standardized task completion process.

[0069] Each step forms a closed-loop collaboration, enabling sanitation robots to operate autonomously throughout the entire process from path planning, target recognition, garbage collection to task completion. This significantly improves the precision and intelligence of operations and completely solves the problems of existing technologies, such as lack of global optimality in path planning, low target positioning accuracy, unsystematic cleaning operations, poor continuity of cyclical operations, low intelligence of auxiliary tasks, and imperfect completion processes.

[0070] Furthermore, such as Figure 2 As shown, step S100 includes:

[0071] Step S101: Obtain the task information and global map corresponding to the sanitation task; the task information includes the global operation area and the target area located within the global operation area; the global map contains multiple preset key nodes; the key nodes include garbage collection points, charging points, etc.

[0072] Step S102: Based on the global operation area and the target area, the global map is rasterized to obtain a raster map; attribute information is configured for each grid in the raster map, the attribute information including at least terrain type, target area marker for indicating whether it belongs to the target area, and key node marker for marking key nodes; based on the key node markers and the global operation area, a set of valid key nodes located within the global operation area is selected;

[0073] Step S103: Identify the target area from the grid map based on the target area marker, and calculate the task demand intensity of each grid in the target area according to historical task data; divide the target area into multiple sub-areas of different importance levels according to the task demand intensity; perform coverage path planning using different coverage parameters for sub-areas of different importance levels, and generate a set of target area coverage path segments.

[0074] Step S104: Spatial clustering of the set of effective key nodes and the set of target area coverage path segments is performed to obtain multiple clusters; for each cluster, a shortest intra-cluster sub-path is planned to connect all effective key nodes and target area coverage path segments in the cluster; then, an inter-cluster path is planned to connect adjacent clusters, and all the shortest intra-cluster sub-paths are connected to form an initial path framework.

[0075] Step S105: Divide the initial path framework into multiple path segments, and label each path segment with its path type and path information; wherein, the path type is determined according to the terrain type of the grid through which the path segment passes; verify each path segment based on preset accessibility constraints, and perform local replanning on path segments that do not meet the constraints to obtain a feasible path that meets all accessibility constraints.

[0076] Step S106: Construct a comprehensive cost function to evaluate the feasible path, and with the goal of minimizing the comprehensive cost, optimize the feasible path under the preset global constraints to obtain the globally optimal patrol path.

[0077] Step S107: Determine whether the globally optimal patrol path satisfies the operation time limit constraint; if not, adjust the weight coefficient of the comprehensive cost function and re-execute the optimization solution process of step S106 until the globally optimal patrol path that satisfies the operation time limit constraint is obtained.

[0078] This invention discloses a complete process for the globally optimal patrol path, which sequentially executes the following steps: S101 Obtaining task information and a global map; S102 Rasterization processing and filtering of effective key node sets; S103 Calculating task demand intensity and dividing sub-regions and planning coverage path segments; S104 Spatial clustering and constructing an initial path framework; S105 Segmented verification and local replanning to obtain feasible paths; S106 Constructing a comprehensive cost function and optimizing the solution for the globally optimal patrol path; S107 Determining the operation time limit and adjusting weights to re-optimize until constraints are met and the path is output. The core feature is that it abandons the manual fixed path mode and achieves dynamic planning by combining task requirements, terrain constraints, key nodes, comprehensive cost, and operation time limit.

[0079] By employing dynamic planning throughout the entire process, the system calculates the intensity of task demand and divides the area into sub-regions based on historical task data (data collected by sanitation robots during previous similar patrol and cleaning tasks in the area, including garbage collection frequency, weight, quantity, and operation time). This enables differentiated coverage of areas with different needs, solving the problems of manual fixed paths failing to dynamically adapt to task requirements and unreasonable resource allocation. Through path segmentation verification and local replanning, obstacles caused by complex terrain can be avoided, addressing the issue of fixed paths failing to consider terrain constraints and potentially becoming impassable. Spatial clustering connects effective key nodes with covered path segments, preventing the omission of key nodes and resolving the issue of poor key node connection in fixed paths. Comprehensive cost optimization and weight adjustment adapt to operation time limits, quickly obtaining the globally optimal path that meets time constraints, solving the problem of time-consuming and labor-intensive manual replanning when fixed paths do not meet time limits. Ultimately, this achieves dynamic path adaptation, high efficiency, feasibility, and global optimization, improving patrol operation efficiency, task completion quality, and resource utilization.

[0080] Furthermore, such as Figure 3 As shown, step S200 includes:

[0081] Step S201: Image acquisition, acquiring panoramic images of the environment surrounding the robot;

[0082] Step S202: Image enhancement. Perform enhancement preprocessing on the panoramic image targeting small targets to obtain a preprocessed image.

[0083] Step S203: Extraction of Region of Interest. The preprocessed image is semantically segmented to extract the preset region of interest, and the region of interest is divided into multiple sub-images.

[0084] Step S204: Collaborative recognition. Perform multi-level collaborative recognition on each sub-image to determine whether it contains a target object and to determine its category. The collaborative recognition includes preliminary detection based on a convolutional neural network, and cross-modal semantic verification by calling a multimodal large model based at least on the preliminary detection results.

[0085] Step S205: Target localization. Based on the position of the target object in the panoramic image and the robot's localization information, calculate the world coordinates of the target object in the working environment.

[0086] This invention executes five core steps sequentially: image acquisition, image enhancement, region of interest extraction, collaborative recognition, and target localization. First, it acquires panoramic images of the robot's surrounding environment and performs enhancement preprocessing to optimize image quality for small targets. Then, it extracts preset regions of interest through semantic segmentation and divides the sub-image focusing recognition range. Next, it uses a multi-level collaborative recognition method of "preliminary detection by convolutional neural network + cross-modal semantic verification by multimodal large model" to determine whether the target object is included and to determine its category. Finally, it fuses the position of the target object in the panoramic image with the robot's positioning information to complete the calculation of the world coordinates of the target object in the working environment, forming a complete technical link of "image acquisition-optimization-recognition-localization".

[0087] This invention enables robots to accurately identify and spatially locate target objects in the work environment throughout the entire process. It enhances the recognition of small target features through image enhancement, reduces irrelevant background interference by extracting the area of ​​interest, and balances detection efficiency and judgment accuracy through multi-level collaborative recognition. Target localization realizes the conversion of image pixel positions to actual world coordinates in the work environment, providing accurate target object category and spatial location information for the robot's subsequent work decisions.

[0088] The target object is the waste to be cleaned up. The preset area of ​​interest includes the ground area and / or the area around the waste container.

[0089] Furthermore, in step S300, generating and executing the cleanup subtask specifically includes:

[0090] Based on the world coordinates of the target object and the current location of the robot, plan a sub-path from the current location to the target object;

[0091] Control the robot to move along the sub-path to the target object;

[0092] Based on the category of the target object, the robot arm adapted to the robot body is invoked to perform a grasping operation;

[0093] The robotic arm is controlled to store the grasped target object into the classification bin corresponding to the target object's category in the robot's storage box.

[0094] Based on the world coordinates of the target object and the robot's current location, a sub-path is planned to ensure that the robot accurately navigates to the location of the waste; according to the waste category, the appropriate robotic arm is called to perform the grasping operation, improving the grasping success rate; and the waste is stored in the corresponding sorting bin to achieve waste sorting and storage.

[0095] The above process improves the efficiency and accuracy of waste disposal and sorting, ensures the standardization of disposal operations, and solves the problems of unplanned disposal actions, poor grabbing adaptability, and unsorted storage in existing technologies, thus meeting the requirements of operational standards.

[0096] Furthermore, in step S400, the real-time monitoring of the robot's battery status and the storage box's capacity status specifically includes:

[0097] The robot's power management system periodically acquires the remaining battery power percentage as the power status.

[0098] The current waste weight or filling rate is obtained by using a weighing sensor or vision sensor installed in the storage box, which serves as the capacity status.

[0099] The power management system periodically collects the battery percentage to obtain the robot's battery status in real time; the weight of the waste or the filling rate is accurately obtained as the capacity status through the weighing sensor or vision sensor in the storage box.

[0100] The above process enables precise monitoring of power and capacity status, with no delay in data feedback, avoiding situations where patrols continue even when the storage tank is full or operations are interrupted due to power depletion, thus ensuring the continuity of cyclical operations.

[0101] Furthermore, step S500 specifically includes:

[0102] When the capacity of the storage box reaches a preset full-load threshold, a dumping task is triggered.

[0103] Based on the robot's current location and the location information of all garbage collection points in the global map, a first navigation path to the nearest garbage collection point is planned;

[0104] The robot is controlled to move along the first navigation path, and upon arriving at the waste collection point, the storage box is controlled to perform an emptying operation.

[0105] The dumping task is triggered by a full capacity threshold to ensure timely triggering; the first navigation path to the nearest point is planned based on the robot's current location and the location of the garbage collection point on the global map to shorten the round trip time; the storage box is controlled to automatically perform the dumping operation without human intervention.

[0106] The above process improves the efficiency of waste dumping operations, reduces non-operational time, and solves the problems of untimely triggering of dumping tasks, suboptimal navigation paths, and the need for manual intervention in existing technologies.

[0107] Furthermore, step S600 specifically includes:

[0108] When the robot's battery level falls below a preset low battery threshold, a charging task is triggered.

[0109] Based on the robot's current location and the location information of all charging points in the global map, a second navigation path to the nearest charging point is planned;

[0110] The robot is controlled to move along the second navigation path, and upon arrival at the charging point, it automatically docks with the charging interface and performs the charging process.

[0111] The robot triggers charging tasks by low battery thresholds to prevent it from interrupting operations due to insufficient power; it plans a second navigation path to the nearest charging point based on the robot's current location and the location of charging points on the global map, reducing energy consumption and time consumption; and it automatically connects to the charging interface and completes the charging process without human assistance.

[0112] The above process ensures the stability of the robot's battery life, reduces maintenance costs, and solves the problems of delayed charging task triggering, redundant navigation paths, and the need for manual charging in existing technologies.

[0113] Furthermore, in step S700, the conditions for completion include: reaching the preset total operation time of the sanitation task, or not identifying any more garbage to be cleaned during the complete traversal of the globally optimal patrol path.

[0114] The system sets two termination conditions: total job duration and no garbage collection during complete traversal. This covers both timed and quantitative cleanup scenarios, avoiding the limitations of a single condition.

[0115] The above process enables flexible determination of task completion, avoids ineffective patrols or excessive work, and improves the rationality of task completion and operational efficiency.

[0116] Furthermore, step S700 specifically includes:

[0117] Step S701, Final Disposal: Control the robot to navigate to the nearest waste collection point and empty all the waste in the storage bin;

[0118] Step S702, Final Charging: Control the robot to navigate from the waste collection point to the nearest charging point and perform a charging operation until the battery level reaches a preset saturation value;

[0119] Step S703, Information Reporting: Generate an execution report for the sanitation task and send it to the cloud. The execution report shall include at least the task identifier, total operation time, total amount of garbage cleaned, and final status.

[0120] The process of clearing the remaining waste from the storage bin by emptying it and then recharging it to the preset saturation value is standardized. The execution report is generated and submitted to provide complete data such as task identification, total operation time, and total amount of waste cleared.

[0121] The above process ensures the continuity of the robot's next operation, avoiding problems such as residual garbage occupying storage space and insufficient power affecting the next start-up; at the same time, it provides data support for optimizing subsequent operation plans.

[0122] Furthermore, in step S102, the attribute information configured for each grid also includes terrain parameters; the terrain parameters include the height difference between adjacent grids and a preset upper limit for travel speed based on the terrain type.

[0123] By supplementing the grid with terrain parameters such as adjacent height differences and upper limits of travel speed, accurate terrain data support can be provided for the path segmentation verification (such as height difference constraints and speed constraints) and terrain adaptation penalty calculation in subsequent steps S105. This makes the subsequent path traversability verification more targeted and the terrain adaptation evaluation more accurate. It solves the problem that fixed paths in the prior art do not fully consider complex terrain constraints, improves the scientificity and feasibility of path planning, and avoids the path becoming impassable due to missing terrain data.

[0124] Furthermore, in step S102, based on the key node markers and the global work area, a set of valid key nodes located within the global work area is selected. Specifically, key nodes located outside the global work area and those marked as prohibited areas in the grid map are removed, and the remaining key nodes constitute the set of valid key nodes.

[0125] By removing key nodes outside the work area and prohibited areas, it ensures that all nodes in the effective key node set are within the reachable range, avoiding the connection of inaccessible key nodes during subsequent path planning, which could lead to the path failing to execute or missing effective key nodes. This solves the problem of missing key nodes or accessing inaccessible nodes in existing technologies with fixed paths, and improves the accuracy and practicality of subsequent path planning.

[0126] Furthermore, in step S103, the task demand intensity of each grid within the target area is calculated based on historical task data. Specifically, based on historical task data, the historical task volume of each grid is calculated (the cumulative weight of garbage cleaned, the number of garbage items, or the number of garbage cleaned in past tasks of the grid). The ratio of the historical task volume to the grid area is determined as the task demand intensity of the corresponding grid.

[0127] By quantifying the task demand intensity of each grid by the ratio of historical task volume to grid area, the abstract "task demand" can be transformed into specific calculable parameters. This provides a precise quantitative basis for dividing the target area into sub-regions according to demand intensity, avoiding resource waste or insufficient coverage caused by the inability to distinguish task demands in different areas and the adoption of a uniform coverage mode. It solves the problem of poor adaptability of fixed path tasks in existing technologies, lays the foundation for differentiated coverage path planning, and improves operational efficiency and the rationality of resource allocation.

[0128] Furthermore, the task demand intensity ρ is calculated using the following formula: ρ=Q / A, where ρ represents the task demand intensity of the grid, Q represents the historical task volume of the grid, and A represents the area of ​​the grid.

[0129] Furthermore, in step S103, the target area is divided into multiple sub-regions of different importance levels according to the task demand intensity. Specifically, two task demand intensity thresholds are preset, and the target area is divided into three levels: high demand sub-region, medium demand sub-region, and low demand sub-region according to the task demand intensity from high to low. Each level of sub-region is composed of continuous grids that meet the threshold conditions.

[0130] By dividing the target area into three levels of sub-regions according to the intensity of task requirements, hierarchical management of the target area can be achieved. This provides a clear classification basis for adopting different coverage parameters for different levels of sub-regions, enabling high-demand areas to receive key coverage and low-demand areas to have their patrol frequency reasonably reduced. This solves the problems of uneven coverage of fixed paths and unreasonable resource allocation in existing technologies, thereby improving the quality of task completion and operational efficiency.

[0131] Furthermore, the two preset task demand intensity thresholds are ρ0 and ρ1, respectively, and ρ0 > ρ1;

[0132] When the task demand intensity ρ of a grid is greater than or equal to ρ0, the grid belongs to a high-demand sub-region.

[0133] When ρ1≤ρ<ρ0, the grid belongs to the medium demand sub-region;

[0134] When ρ < ρ1, the grid belongs to the low-demand sub-region.

[0135] By clarifying the relationship between the two thresholds and the criteria for determining each level of sub-region, the sub-region division criteria become clearer and more executable. This avoids confusion in the application of coverage parameters caused by ambiguous division criteria, ensuring accurate division of high, medium, and low demand sub-regions. This allows for the orderly implementation of subsequent differentiated coverage path planning and further improves the adaptability of the path to task requirements.

[0136] Furthermore, in step S103, different coverage parameters are used for coverage path planning for sub-regions of different importance levels. Specifically, the coverage parameters include coverage radius and traversal count. The coverage radius of high-demand sub-regions is the smallest and the traversal count is the largest. The coverage radius of medium-demand sub-regions and low-demand sub-regions increase sequentially, and the traversal count decreases sequentially. The coverage path planning adopts a partition scanning method combined with reciprocating path traversal for each sub-region. In the generated target area coverage path segment set, each path segment uniquely corresponds to one sub-region.

[0137] Because high-demand sub-regions use smaller coverage radii and more traversals, while medium- and low-demand sub-regions adjust their coverage parameters sequentially, this ensures that high-demand areas receive sufficient coverage and that low-demand areas do not waste patrol resources, thus solving the problem of uneven coverage along fixed paths. At the same time, the partitioned scanning method combined with reciprocating path traversal ensures that all target grids in each sub-region are covered, and that each path segment corresponds to a sub-region, providing a regular path foundation for subsequent spatial clustering and improving the comprehensiveness of regional coverage and the standardization of path planning.

[0138] Furthermore, the coverage radius of the high-demand sub-region is 3 meters and the number of traversals is no less than 2; the coverage radius of the medium-demand sub-region is 4 meters and the number of traversals is 1; and the coverage radius of the low-demand sub-region is 5 meters and the number of traversals is 1.

[0139] By specifying the exact coverage radius and traversal number for each level of sub-region, the coverage parameters become more concrete and implementable, avoiding poor coverage results caused by ambiguous parameters. This ensures that high-demand areas (such as densely populated garbage areas in sanitation scenarios) are adequately covered, while low-demand areas do not require excessive patrolling, further improving the accuracy of differentiated coverage. At the same time, it standardizes the operational criteria for coverage path planning, facilitating practical implementation.

[0140] Furthermore, the partition scanning method divides the target area evenly according to a preset number of grid rows or columns, and the reciprocating path traversal refers to planning a path in each sub-region in an alternating manner of "from left to right and from top to bottom" to ensure that all grids marked as the target area in the sub-region are covered.

[0141] By clarifying the specific operation methods of the partitioned scanning method and the reciprocating path, the problem of missed or repeated coverage caused by unclear traversal methods can be avoided, ensuring that all target grids in each sub-region can be fully and uniformly covered. At the same time, the alternating reciprocating path can reduce path redundancy, shorten the coverage path length within the sub-region, reduce operation energy consumption and time, and improve the efficiency and regularity of coverage path planning.

[0142] Furthermore, in step S104, the set of effective key nodes and the set of target area coverage path segments are spatially clustered to obtain multiple clusters. Specifically, the K-means clustering algorithm is used to cluster effective key nodes and target area coverage path segments that are spatially adjacent. The constraints that must be met during clustering include: the spatial distance between any effective key node in the cluster and the endpoint of the target area coverage path segment is not greater than a preset distance threshold d0, and each cluster contains at least one of the effective key nodes.

[0143] By using the K-means algorithm to cluster spatially adjacent effective key nodes and covered path segments, and ensuring that each cluster contains at least one effective key node, scattered key nodes and covered path segments can be aggregated according to their spatial location. This avoids excessive detours and backtracking during subsequent path connection, thus shortening the total path length. At the same time, intra-cluster distance constraints ensure the correlation between path segments and key nodes within a cluster, laying the foundation for planning the shortest intra-cluster sub-paths. This solves the problem in the background technology that fixed paths cannot effectively connect key nodes, improving the rationality and efficiency of path connection.

[0144] Furthermore, in step S104, for each cluster, a shortest intra-cluster sub-path is planned that connects all valid key nodes and target area coverage path segments within the cluster. Specifically, the endpoints of all valid key nodes and target area coverage path segments within the cluster are considered as points to be visited, and a shortest path is planned that traverses all points to be visited.

[0145] By using the effective key nodes and endpoints of the covered path segments within the cluster as the points to be visited, and planning the shortest path to traverse all the points to be visited, the cluster path length can be shortened as much as possible while connecting all core elements (key nodes and covered path segments) within the cluster, reducing path redundancy, and lowering operation energy consumption and patrol time. At the same time, it ensures that all effective key nodes within the cluster are visited, avoiding missed visits, and solves the problems of missed visits to key nodes and path redundancy in fixed paths, providing support for building an efficient initial path framework.

[0146] Furthermore, planning a shortest path that traverses all points to be visited specifically includes: constructing a matrix of the shortest travel distances between each pair of points to be visited, and solving for the shortest path using a dynamic programming algorithm.

[0147] By constructing the shortest distance matrix between the points to be visited and solving it using dynamic programming, it can be ensured that the obtained intra-cluster sub-paths are the true shortest paths, avoiding length redundancy caused by unreasonable path planning. At the same time, dynamic programming is efficient and accurate, and can quickly obtain the optimal solution, further improving the efficiency and accuracy of path planning, ensuring the optimality of intra-cluster sub-paths, and providing a guarantee for the efficiency of the subsequent initial path framework.

[0148] Furthermore, in step S104, the inter-cluster paths connecting adjacent clusters are planned, and all the shortest intra-cluster sub-paths are connected in series to form an initial path framework. Specifically, according to the spatial order of the clusters, the shortest travel path between adjacent clusters is calculated sequentially using a heuristic search algorithm (A* algorithm), and the shortest intra-cluster sub-paths of each cluster are connected through the shortest travel path to form the initial path framework.

[0149] Since the A* algorithm is used to calculate the shortest path between adjacent clusters according to their spatial location, it ensures that the inter-cluster path is the shortest, avoiding detours and backtracking when connecting clusters, and further shortening the total length of the initial path framework. At the same time, the A* algorithm is highly efficient in solving the shortest path and is adapted to grid map scenarios. It can quickly complete the inter-cluster path planning, connecting all the shortest sub-paths within the clusters to form the initial path framework, providing an efficient and regular basic path for subsequent path segmentation verification and optimization.

[0150] Furthermore, in step S105, the initial path frame is segmented to obtain multiple path segments, and each path segment is labeled with its path type and path information. Specifically, the initial path frame is divided into multiple path segments according to a continuous grid sequence of the same terrain type; wherein, the path type is the terrain type of the grids traversed by the path segment (flat road, gravel road, steps, etc.); the path information includes the length of the path segment, the starting coordinates, the ending coordinates, the maximum height difference between adjacent grids within the segment, and the upper limit of the travel speed preset based on the terrain type of the path segment.

[0151] By segmenting paths according to a continuous raster sequence of the same terrain type, and clearly defining and separately labeling the criteria for determining path types and the specific components of path information, the problem of ambiguous distinction between path types and path information is completely solved, avoiding confusion that could lead to errors in subsequent data extraction. Simultaneously, the direct correspondence between path types and terrain types provides a direct and clear basis for terrain type verification in subsequent accessibility constraint validation. Complete and targeted path information provides accurate and comprehensive parameter support for subsequent verification of height difference constraints, speed constraints, and terrain adaptation penalty calculations, ensuring that subsequent steps can quickly and accurately extract the required data. This improves the efficiency and accuracy of path segmentation verification and comprehensive cost calculation, laying the foundation for the efficient generation of feasible paths that meet all accessibility constraints.

[0152] Furthermore, in step S105, each path segment is verified based on preset traversability constraints, and local replanning is performed on path segments that do not meet the constraints to obtain a feasible path that meets all traversability constraints. Specifically, for each path segment, it is checked whether it meets the preset traversability constraints; if it does not meet the constraints, a local path is replanned using the starting and ending points of the path segment as the starting and ending points, and the original path segment is replaced by a heuristic search algorithm (A* algorithm) until all path segments meet the traversability constraints, thus obtaining the feasible path.

[0153] Because each path segment is tested for traversability separately, local replanning is only performed on road segments that do not meet the constraints, without the need to replan the overall initial path framework, which greatly improves the efficiency of path planning. At the same time, the use of the A* algorithm to replan local paths ensures that the replanned road segments are still the shortest paths and avoid impassable areas, ultimately obtaining a feasible path that meets all traversability constraints. This solves the problem in existing technologies where fixed paths do not consider terrain constraints and may be impassable, thus improving the reliability of path traversability.

[0154] Furthermore, the preset accessibility constraints are specifically as follows:

[0155] The terrain type of the grid cells traversed by the path segment cannot be a prohibited type;

[0156] The absolute value of the height difference between any two adjacent grid cells within a path segment does not exceed the preset maximum height difference Δh_max;

[0157] The planned travel speed v_plan for the path segment shall not exceed the maximum travel speed v_max preset for that path segment based on its terrain type.

[0158] By clarifying the three specific accessibility constraints, the accessibility verification standards for path segments become clearer and more executable, avoiding inaccurate verification caused by ambiguous constraints. At the same time, terrain type constraints can prevent the path from passing through prohibited areas, while height difference constraints and speed constraints can ensure that the path is adapted to the robot's accessibility, avoiding the robot's inability to pass due to complex terrain (such as excessive height difference) or unreasonable speed, further improving the reliability and safety of feasible paths, and solving the pain point of fixed paths not considering terrain constraints.

[0159] Furthermore, in step S106, constructing a comprehensive cost function to evaluate the feasible path specifically involves constructing a comprehensive cost function obtained by weighted summation of a coverage penalty term reflecting the coverage rate of the target area, a length term reflecting the total length of the path, and a terrain adaptation penalty term reflecting the terrain accessibility of the path.

[0160] Since the comprehensive cost function includes coverage penalty, path length penalty, and terrain adaptation penalty, it can comprehensively evaluate the overall performance of feasible paths, avoiding the suboptimal path caused by evaluating a single factor (such as considering only path length). At the same time, the coverage penalty can guide the path to prioritize ensuring coverage, and the terrain adaptation penalty can guide the path to prioritize selecting road segments with good terrain adaptability. This solves the problem in the background technology that fixed paths cannot take into account multiple factors and cannot achieve global optimization, providing a comprehensive and scientific evaluation basis for subsequent optimization to find the globally optimal path.

[0161] Furthermore, the expression for the comprehensive cost function F is: F = α·C + β·L + γ·T, where F represents the comprehensive cost of the path, α, β, and γ are preset weight coefficients, C represents the coverage penalty term, L represents the length term, and T represents the terrain adaptation penalty term; the weight coefficients α, β, and γ satisfy α + β + γ = 1.

[0162] By clarifying the specific expression of the comprehensive cost function, the calculation of comprehensive cost can be more standardized and accurate, avoiding deviations in optimization direction caused by ambiguity in cost calculation. At the same time, the influence of the three factors of coverage effect, path length, and terrain adaptability on comprehensive cost can be flexibly adjusted through the weight coefficients α, β, and γ, making the comprehensive cost function adaptable to different operational needs (such as increasing β to prioritize time limits), improving the flexibility and adaptability of the comprehensive cost function, and providing a clear calculation basis for subsequent weight adjustments.

[0163] Furthermore, the coverage penalty term C is calculated using the following formula:

[0164] C = (1 – S_c / S_t) × C_m

[0165] Where C represents the coverage penalty term, S_c represents the area of ​​the target region actually covered by the path, the calculation of which depends on the coverage radius defined in step S3; S_t represents the total area of ​​the target region; and C_m represents the preset maximum coverage penalty value.

[0166] Since the coverage penalty term is quantified through the above formula, the coverage effect of the path can be directly related to the overall cost. The worse the coverage effect (the smaller S_c / S_t), the larger the coverage penalty term C, and the higher the overall cost F. This quantification method can guide the optimization algorithm to prioritize the selection of paths with good coverage effect during the iteration process, which solves the problem of insufficient coverage of fixed paths in the existing technology, ensures that the globally optimal patrol path can achieve full coverage of the target area, and improves the quality of task completion.

[0167] Furthermore, the terrain adaptation penalty term T is calculated using the following formula:

[0168] T=[Σ(ω_i×L_i / L)]×T_m

[0169] Where T represents the terrain adaptation penalty, Σ represents the summation operation, ω_i represents the terrain type weight corresponding to the i-th path segment, L_i represents the length of the i-th path segment, L represents the total length of the path, and T_m represents the preset maximum terrain penalty value.

[0170] Since the terrain adaptation penalty term is quantified by the above formula, the terrain adaptation difficulty of the path can be associated with the overall cost. The more complex the terrain (the larger ω_i) and the higher the proportion of road segment length, the larger the terrain adaptation penalty term T and the higher the overall cost F. This quantification method can guide the optimization algorithm to prioritize the selection of paths with good terrain adaptability (such as a high proportion of flat roads), and avoid paths that pass through complex terrain (such as steps or gravel roads) that are impassable or have excessive energy consumption. This solves the problem that fixed paths in the existing technology do not consider terrain constraints, and improves the terrain adaptability and passage efficiency of the path.

[0171] Furthermore, in step S106, with the goal of minimizing the overall cost, the feasible paths are optimized under the preset global constraints to obtain the globally optimal patrol path. Specifically, the feasible paths are iteratively optimized using an optimization algorithm. Each iteration generates a new path population, calculates the overall cost of each path in the population, and checks whether it meets the global constraints. Finally, the path that satisfies all constraints and has the lowest overall cost is output as the globally optimal patrol path.

[0172] By iteratively optimizing the algorithm and calculating the overall cost and checking global constraints in each iteration, it is ensured that the final output path not only has the lowest overall cost but also meets the core requirements of the patrol mission (global constraints). At the same time, iterative optimization can avoid getting trapped in local optima and ensure that the global optimal solution is obtained. This solves the problem that fixed paths cannot achieve global optima in the background technology, improves the overall performance of the path (taking into account coverage, length, and terrain adaptation), and ensures that the path can efficiently complete the patrol mission.

[0173] Furthermore, the optimization algorithm is a genetic algorithm, specifically employing integer encoding, an elite retention strategy, partial matching crossover, and random mutation operations.

[0174] By employing a genetic algorithm, combined with integer encoding (adapting to the grid number characteristics of the grid path), an elite retention strategy (preserving the best individuals in each generation to avoid losing the optimal solution), partial matching crossover (ensuring the connectivity of the path after crossover), and random mutation (avoiding getting trapped in local optima), the efficiency and accuracy of optimization can be improved. Integer encoding adapts to the path characteristics of the grid map, elite retention and random mutation can ensure that the global optimal solution is found quickly, and partial matching crossover can avoid path breakage after crossover, ensuring that the optimized global optimal patrol path is connected and feasible, further improving the efficiency and quality of path planning.

[0175] Furthermore, the preset global constraints specifically include:

[0176] Key node access constraint: Each of the effective key nodes is accessed by the globally optimal patrol path 1 to 2 times;

[0177] Task time constraint: The predicted execution time of the globally optimal patrol path shall not exceed the task time specified in the patrol task;

[0178] Region boundary constraint: All points on the global optimal patrol path are located within the global operation area.

[0179] By setting three global constraints, it can be ensured that the optimized global patrol path can meet the core requirements of the patrol task: the critical node access constraint can avoid missing or over-accessing critical nodes, solving the problem of missing critical nodes in fixed paths; the operation time limit constraint can ensure that the path can be completed within the specified time, solving the problem of fixed paths not meeting the operation time limit; the area boundary constraint can ensure that the path does not exceed the operation range, avoiding ineffective patrols; the three constraints together guarantee the practicality and compliance of the global optimal patrol path, ensuring that the path can effectively complete the patrol task.

[0180] Furthermore, in step S107, adjusting the weight coefficients of the comprehensive cost function specifically involves increasing the value of the weight coefficient β, while proportionally decreasing the weight coefficients α and γ, so that α+β+γ=1 is still satisfied after adjustment.

[0181] Increasing the weight coefficient β corresponding to the path length term enhances the impact of path length on the overall cost, allowing the optimization algorithm to prioritize shorter paths during iteration (shorter path length, lower overall cost). Simultaneously, proportionally reducing α and γ maintains the reasonableness of the weight coefficients after adjustment (satisfying α+β+γ=1), avoiding deviations in overall cost calculation. This adjustment method can specifically shorten the total path length, thereby reducing path prediction execution time and solving the problem of the globally optimal path not meeting the job time limit, achieving rapid adaptation to job time limits.

[0182] Furthermore, in step S107, after re-executing the optimization solution process of step S106, the output path must simultaneously satisfy the traversability constraint of step S105 and the global constraint condition of step S106.

[0183] Since the re-optimized path must simultaneously satisfy both accessibility constraints and global constraints, it ensures that the path obtained after adjusting the weight coefficients not only meets the time limit requirements but also guarantees that the path is accessible (without impassable sections), has no missed key nodes, and does not exceed the work area. This avoids problems such as "meeting the time limit but being impassable" or "meeting the time limit but missing key nodes," further improving the reliability and practicality of the globally optimal patrol path and ensuring that the path can complete the patrol mission efficiently and safely.

[0184] Furthermore, step S201 specifically includes:

[0185] Step S2011: Simultaneously acquire multiple sets of raw images using multiple vision sensors arranged around the robot;

[0186] Step S2012: Perform optical distortion correction and time synchronization alignment on each group of original images;

[0187] Step S2013: Stitch and fuse the corrected and aligned sets of original images to generate the panoramic image covering a 360-degree field of view around the robot.

[0188] Image acquisition comprises three steps: simultaneous acquisition by multiple vision sensors, optical distortion correction and temporal synchronization alignment, and stitching and fusion. By employing a surround-type multi-sensor acquisition method, the limitations of a single sensor's field of view are overcome. First, distortion correction and temporal alignment are performed on the original images to eliminate geometric distortion and temporal misalignment. Then, stitching and fusion are used to generate a 360-degree panoramic image. Because of the surround-type multi-sensor acquisition, blind spots of a single vision sensor are eliminated. Optical distortion correction and temporal synchronization alignment ensure the geometric and temporal consistency of multiple sets of original images, avoiding ghosting and misalignment issues during stitching and fusion. The resulting 360-degree panoramic image possesses integrity and consistency, providing distortion-free and fully covered foundational image data for subsequent image enhancement, target recognition, and other steps.

[0189] Furthermore, the optical distortion correction in S2012 includes radial distortion correction and tangential distortion correction;

[0190] The mathematical model for radial distortion correction is as follows:

[0191]

[0192]

[0193] Among them, the and The normalized image plane coordinates before correction are represented. and The coordinates represent the radial distortion corrected coordinates. express and The sum of squares (i.e. ), the , , The radial distortion coefficients are obtained in advance using Zhang's calibration method;

[0194] The mathematical model for tangential distortion correction is as follows:

[0195]

[0196]

[0197] Among them, the and This represents the final coordinates after tangential distortion correction. , The tangential distortion coefficients are obtained in advance using Zhang's calibration method;

[0198] After correction, the final pixel coordinates are obtained through mapping using the intrinsic parameter matrix. The mapping formula is as follows: In the formula This is the camera intrinsic parameter matrix. , These are the pixel coordinates of the corrected image.

[0199] Optical distortion correction is refined into radial distortion correction and tangential distortion correction. Based on the distortion coefficients obtained by Zhang's calibration method, corresponding mathematical models are used to correct radial geometric distortion caused by lens optical characteristics and tangential geometric distortion caused by lens mounting deviations, respectively. Then, the corrected normalized coordinates are mapped to actual pixel coordinates through an intrinsic parameter matrix, achieving accurate geometric correction of the original image. Because a dedicated mathematical model is used to specifically correct radial and tangential distortions, and the coefficients obtained by Zhang's calibration method ensure the accuracy of the correction parameters, and the intrinsic parameter matrix completes the accurate mapping from normalized coordinates to pixel coordinates, the geometric distortion of the original image is effectively eliminated, pixel position deviations are corrected, and the geometric accuracy of the corrected image is guaranteed. This provides a unified pixel coordinate basis for the subsequent stitching and fusion of multiple images, avoiding stitching misalignment and contour distortion problems caused by image distortion.

[0200] Furthermore, the time synchronization alignment in S2012 specifically involves sending a hardware synchronization trigger signal to all visual sensors, causing all visual sensors to be exposed simultaneously within a set acquisition period, ensuring that the image acquisition time difference is ≤1 millisecond, and acquiring the original image with timestamp alignment.

[0201] By sending a unified hardware synchronization trigger signal to all visual sensors, all sensors are forced to expose simultaneously within a set acquisition period. This strictly controls the image acquisition time difference between multiple sensors to within 1 millisecond, achieving precise timestamp alignment of multiple sets of original images. Because the hardware synchronization trigger signal ensures simultaneous exposure of all visual sensors and controls the acquisition time difference to within 1 millisecond, it eliminates time deviations caused by asynchronous acquisition, avoids image content misalignment caused by robot movement or target object displacement, and guarantees the temporal consistency of multiple sets of original images. This allows for precise matching of image content and pixel positions during subsequent stitching and fusion, improving the stitching alignment effect of panoramic images.

[0202] Furthermore, the splicing and fusion in S2013 specifically includes:

[0203] Step S20131, Feature Extraction and Matching: Extract feature points between images acquired by adjacent visual sensors and perform matching; use the FLANN matcher to calculate feature matching pairs, use the RANSAC algorithm to remove mismatched pairs, and retain valid matching pairs with an inlier ratio ≥80%;

[0204] Step S20132, Transformation Relationship Solution: Based on the successfully matched feature point pairs, calculate the homography matrix between the adjacent images;

[0205] Step S20133, Image Fusion: Based on the homography matrix, the adjacent images are transformed by perspective and aligned. The overlapping areas are then fused using a weighted fusion method to eliminate seams and generate a seamless panoramic image.

[0206] This invention breaks down the stitching and fusion process into three sub-steps: feature extraction and matching, transformation relationship solving, and image fusion. It uses the FLANN matcher to quickly achieve feature point matching, uses the RANSAC algorithm to eliminate mismatched pairs and retain a high proportion of valid matching pairs, solves the homography matrix based on the valid matching pairs to determine the perspective transformation relationship between images, and finally achieves image alignment through perspective transformation and performs weighted fusion on the overlapping areas.

[0207] Because it retains effective matching pairs with an inlier ratio of ≥80%, the homography matrix is ​​solved more accurately, thereby achieving precise perspective transformation alignment of adjacent images and avoiding image shift and ghosting after stitching. The weighted fusion method effectively eliminates the stitching seams in overlapping areas, ultimately generating a seamless 360-degree panoramic image with consistent visual effects, improving the integrity and visual consistency of the base image.

[0208] Furthermore, step S202 includes the following sub-steps executed sequentially:

[0209] Step S2021: Noise suppression. The panoramic image is filtered using an adaptive median filtering algorithm to suppress noise and preserve the edge features of small target objects.

[0210] Step S2022: Detail enhancement. A multi-scale retinal enhancement algorithm is used to process the noise-suppressed image to improve the local details and contrast of small target objects.

[0211] Step S2023: Image normalization. The image after detail enhancement is normalized to obtain the preprocessed image with consistent lighting conditions.

[0212] This invention designs the image enhancement process into three progressive sub-steps executed sequentially: noise suppression, detail enhancement, and image normalization. First, noise suppression and small target edge preservation are achieved through adaptive median filtering. Then, the details and contrast of small targets are enhanced through a multi-scale retinal enhancement algorithm. Finally, image differences caused by illumination are eliminated through normalization processing, thus completing the layered optimization of the panoramic image.

[0213] By employing a progressive, hierarchical optimization strategy, adaptive median filtering preserves the edge features of small target objects while denoising. The multi-scale retinal enhancement algorithm specifically improves the detail recognition of small target objects, and image normalization eliminates brightness deviations under different lighting conditions. After three steps of processing, a preprocessed image with consistent lighting conditions and clear features is obtained, providing a high-quality image foundation for subsequent region of interest extraction and target recognition, and improving the detection sensitivity and robustness of subsequent algorithms.

[0214] Furthermore, the adaptive median filtering algorithm in step S21 is executed according to the following steps: defining the initial size and maximum size of the filtering window (the window size range is...). Initialize window size The algorithm calculates the median, minimum, and maximum values ​​of pixels within a window, and determines whether to increase the window size or output the median based on the relationship between the median and the minimum and maximum values, in order to preserve the edge details of small target objects while suppressing noise.

[0215] The adaptive median filtering algorithm defines the size range of the filtering window and dynamically adjusts it from the initial size. Based on the relationship between the median and maximum values ​​of pixels within the window, it decides whether to directly output the center pixel value or increase the window size, thereby achieving dynamic adaptation of the filtering window to the pixel features of different regions of the image.

[0216] Because the filtering window can be dynamically adjusted according to the image pixel features, rather than using a fixed window size, it effectively suppresses interference noise such as salt and pepper noise in panoramic images, while avoiding the blurring of small target object edge features caused by fixed window filtering. It fully preserves the edge details of small target objects, providing a clear target object outline basis for subsequent detail enhancement steps and ensuring the feature integrity of small target objects.

[0217] Furthermore, the multi-scale retinal enhancement algorithm in step S2022 calculates the enhanced reflection component using the following formula. :

[0218]

[0219] Among them, the The pixel intensity of the input image is represented by the Represents the natural logarithm operation, the The total number of Gaussian kernel scales selected is represented by Σ, where Σ represents the summation operation. For scale indexing, This represents a two-dimensional convolution operation, the... Indicates the first A Gaussian kernel function, defined as follows: The The standard deviation of the Gaussian kernel is given by... This represents the natural exponential operation; after calculating the reflection component, it performs linear stretching to map the pixel value to the grayscale range of 0~255.

[0220] The multi-scale retinal enhancement algorithm estimates the luminance component by performing two-dimensional convolution operations on the noise-suppressed image using multi-scale Gaussian kernels, separates the reflection component and luminance component of the image using natural logarithm operation, averages the reflection component at multiple scales to obtain the enhanced reflection component, and finally maps the pixel values ​​to the standard grayscale range through linear stretching to enhance the detailed features of small target objects.

[0221] By using a multi-scale Gaussian kernel to achieve precise separation of the reflection component and the brightness component, multi-scale fusion improves the representation ability of local details of small target objects, and linear stretching further enhances the visual recognition of detail features, making small target objects with indistinct features clearer in the image, and significantly improving the detection sensitivity of small target objects in subsequent target recognition steps.

[0222] Furthermore, the standardization process in step S2023 is Z-Score normalization, which is achieved through the following formula: in For the image with enhanced details, The mean of the image. The standard deviation of the image is . The preprocessed image is obtained after normalization.

[0223] The Z-Score normalization algorithm is employed. By calculating the pixel mean μ and standard deviation σ of the enhanced image, a normalization formula is used to convert each pixel value in the image into a standardized value. This eliminates the overall brightness deviation and pixel value fluctuations under different lighting conditions, achieving standardized grayscale processing. Because pixel value standardization is achieved through the mean and standard deviation, pixel value differences caused by uneven lighting and changes in ambient brightness in panoramic images are effectively eliminated. This ensures consistent lighting conditions in the preprocessed image, avoiding interference from lighting factors on subsequent semantic segmentation and object recognition algorithms, and improving the adaptability and robustness of subsequent algorithms under different lighting scenarios.

[0224] Furthermore, step S203 specifically includes:

[0225] Step S2031: Semantic segmentation. The preprocessed image is processed using a semantic segmentation neural network model to output an initial binary mask image, in which pixels belonging to the preset region of interest are activated.

[0226] Step S2032, Mask optimization: Perform morphological dilation and erosion operations on the initial binary mask image, and filter out connected regions with an area smaller than a threshold to obtain the optimized final mask image.

[0227] Step S2033: Region cropping, performing a dot product operation between the preprocessed image and the final mask image to extract image blocks containing only the region of interest;

[0228] Step S2034: Divide the image block into multiple sub-images of the same size using the sliding window method.

[0229] The extraction of regions of interest (ROIs) is broken down into four sub-steps: semantic segmentation, mask optimization, region truncation, and partitioning. First, semantic segmentation generates an initial binary mask image to preliminarily extract ROIs. Then, morphological operations and connected component filtering optimize the mask quality. Next, dot product operations eliminate non-ROIs. Finally, a sliding window method is used to divide the ROI image block into standardized sub-images, achieving accurate extraction and standardization of ROIs. Because semantic segmentation and mask optimization accurately extract the predefined ROIs and eliminate irrelevant background areas, the computational load for subsequent target recognition is reduced. The sliding window method divides the ROI into sub-images of the same size, adapting to the input requirements of subsequent convolutional neural network models, avoiding recognition errors caused by inconsistent image sizes, and improving the efficiency and accuracy of collaborative recognition.

[0230] Furthermore, the semantic segmentation neural network model in step S2031 is a convolutional neural network with a U-Net++ architecture. Its encoder contains multiple convolutional blocks for downsampling, and the decoder contains multiple transposed convolutional blocks for upsampling and concatenating them with the corresponding layer features of the encoder. The output layer generates the initial binary mask image through a 1×1 convolutional layer and a Sigmoid activation function.

[0231] A convolutional neural network based on the U-Net++ architecture is used as the semantic segmentation model. Its encoder extracts deep features from the preprocessed image through downsampling of convolutional blocks, while the decoder achieves multi-scale feature fusion by upsampling transposed convolutional blocks and concatenating them with features from the corresponding layer of the encoder. The output layer generates a binary mask image through a 1×1 convolutional layer and a sigmoid activation function, accurately distinguishing pixels in regions of interest from those in non-interest regions. The dense feature concatenation advantage of the U-Net++ architecture enhances the extraction and fusion capabilities of multi-scale features. The 1×1 convolutional layer and sigmoid activation function achieve accurate binary segmentation of regions of interest from those in non-interest regions. The generated initial binary mask image accurately activates pixels in the preset regions of interest, providing a high-quality initial mask foundation for subsequent mask optimization steps and reducing the correction cost of mask optimization.

[0232] Furthermore, in S2032, the morphological dilation operation uses a 3×3 rectangular structuring element to dilate the initial binary mask image, and the erosion operation uses a 3×3 rectangular structuring element to erode the dilated image. The area threshold is set to 100 pixels.

[0233] The expansion operation formula is as follows ,in The morphological dilation operator (SE) represents performing a dilation operation on the initial binary mask image M using the structuring element SE; the erosion operation formula is... ,in The morphological erosion operator represents applying the structuring element SE to the dilated mask image. Perform the erosion operation; where M is the initial binary mask image and SE is a 3×3 rectangular structuring element.

[0234] Mask optimization uses a 3×3 rectangular structuring element and performs dilation operations. Fill in the tiny holes in the initial binary mask image, and then perform erosion operation. By eliminating burrs at the mask edges and filtering out connected regions smaller than 100 pixels, false small areas are removed, thus achieving morphological optimization of the mask image. The combination of "dilation-erosion" morphological operations specifically addresses the holes and burrs in the initial mask. The 100-pixel area threshold effectively eliminates false small regions, resulting in a final mask image with regular contours and accurate regions. This ensures that subsequent region extraction steps can extract complete and accurate regions of interest, avoiding omissions or mis-extractions of regions of interest due to mask defects.

[0235] Furthermore, the dot product operation in step S2033 is to multiply the preprocessed image... With the final mask image Pixel-by-pixel multiplication yields a focused image containing only the region of interest. ,Right now .

[0236] Utilizing the characteristic of the final mask image that "pixel values ​​of 0 in non-interested areas and 1 in interested areas", the preprocessed image is... With the final mask image A pixel-by-pixel multiplication operation is performed to zero out the pixel values ​​of non-interested regions, retaining only the pixel values ​​of the interest regions. Because the multiplication operation utilizes the pixel characteristics of the mask image to accurately remove non-interested regions, only the focused image containing the preset interest region is extracted. This significantly reduces the interference of irrelevant background pixels on subsequent processing, focuses the distribution area of ​​the target object, and improves the processing efficiency of subsequent sub-image segmentation and target recognition.

[0237] Furthermore, in step S2034, the sliding window method uses a window size of 64×64 pixels, with a step size of 32 pixels, to focus the image. The image is divided into multiple sub-images.

[0238] Using a fixed window size of 64×64 pixels, with a sliding step of 32 pixels, the image is focused. The image is divided into multiple sub-images of the same size by sliding sequentially upwards, thus standardizing the input units for target recognition. By dividing the focused image into sub-images of uniform size, it accurately matches the input size requirements of subsequent convolutional neural network models, avoiding recognition errors caused by inconsistent image sizes. This allows the collaborative recognition step to perform detection using standardized image units, improving the standardization and detection efficiency of collaborative recognition.

[0239] Furthermore, step S204 specifically includes:

[0240] S2041. First-level detection: The sub-image is input into a first convolutional neural network model, which outputs the location bounding boxes of one or more candidate target object regions and their preliminary categories and confidence scores, wherein the location bounding boxes, preliminary categories and confidence scores together constitute the preliminary detection results.

[0241] S2042, Second-level classification: For each candidate target object region in the preliminary detection result, the corresponding image region is cropped from the sub-image according to its position bounding box, and the image region is input into the second convolutional neural network model for fine-grained classification to obtain a vector containing the probability values ​​of each category. The category corresponding to the highest probability value is taken as the fine classification result.

[0242] S2043, Level 3 verification: Based on the comparison between the confidence level of the fine classification result and the preset threshold, determine whether to trigger verification; if triggered, construct a text prompt that integrates the image region, the fine classification result, the preliminary category and confidence level in the preliminary detection result, the robot's current state and environmental context information, input it into the multimodal big model for analysis and judgment, and output the final recognition category of the target object from the big model.

[0243] The collaborative recognition process is broken down into three progressive sub-steps: first-level detection, second-level classification, and third-level verification. In S2041, the location bounding box, preliminary category, and confidence level together constitute the preliminary detection result. In S2042, the image region is accurately cropped based on the location bounding box of the preliminary detection result, and fine-grained classification is performed. In S2043, the confidence level of the fine-grained classification result is used to determine whether to trigger verification. If triggered, the image region, fine-grained classification result, core information from the preliminary detection result, robot state, and environmental context are fused to construct a text prompt. The prompt is then input into a multimodal large model to complete the final judgment, thus achieving multi-level collaboration of "preliminary detection - fine-grained classification - cross-modal verification". Because the three-level steps are progressive and make full use of the preliminary detection results, S2041 provides clear processing objects and basic information for subsequent steps, S2042 realizes fine-grained category determination of candidate target object regions, and S2043 supplements low-confidence results with multi-dimensional information and performs cross-modal verification through a multi-modal large model, forming a logical closed loop of collaborative recognition, which effectively improves the accuracy of target object category determination and reduces the problems of false recognition and missed recognition.

[0244] Furthermore, the environmental context information includes at least one of the following: map information of the work area and task history information.

[0245] At least one of the following—map information of the work area and task history information—is incorporated as environmental context information into the text prompts for the third-level verification. This provides the pre-trained image-text multimodal processing language model with background information about the robot's working environment, enabling the model to analyze and determine the target object category based on the actual situation of the work area. Because the environmental context information supplements the map and task history of the work area, the cross-modal semantic verification of the multimodal model can be analyzed in conjunction with the robot's actual working scenario, avoiding one-sided judgments based solely on image information. This improves the rationality and accuracy of the third-level verification results and further ensures the reliability of the final target object category identification.

[0246] Furthermore, the first convolutional neural network model in step S2041 is a YOLOv8 model, whose detection head adopts an Anchor-Free structure to directly predict the center coordinates, width and height, confidence level and class probability of the target object, and eliminate duplicate detection boxes through non-maximum suppression.

[0247] The formula for calculating the candidate region bounding box is as follows: , , , ,

[0248] In the formula , The coordinates of the target center are , Target width and height.

[0249] The YOLOv8 model is used as the first-level detection convolutional neural network model. Its anchor-free structure eliminates the need for pre-defined anchor boxes, directly predicting the center coordinates, width, height, confidence score, and class probability of the target object. The specific location of the candidate region bounding box is calculated using a dedicated formula, and then a non-maximum suppression algorithm is used to eliminate duplicate detection boxes, resulting in a unique candidate target object region. Because the anchor-free structure of the YOLOv8 model improves the detection capability and speed for small objects, the bounding box calculation formula achieves precise quantification of the candidate target object region location, and the non-maximum suppression algorithm effectively eliminates duplicate detection boxes, it can quickly and accurately obtain the location, preliminary class, and confidence score of the candidate target object region. This provides an accurate and unique candidate region foundation for subsequent second-level fine classification, improving the efficiency of fine classification.

[0250] Furthermore, the second convolutional neural network model in step S2042 is a ResNet-50 model, which extracts deep features through multiple residual blocks containing residual connection structures, and outputs the vector containing the probability values ​​of each category through global average pooling and fully connected layers.

[0251] The formula for calculating the category probability distribution is: In the formula For the first The probability value of the category, The logits value is the output of the fully connected layer. This represents the total number of target object categories.

[0252] A ResNet-50 model is used as the convolutional neural network model for the second-level classification. Deep features of candidate target object regions are extracted through residual blocks with residual connections, avoiding the gradient vanishing problem in deep networks. Global average pooling and fully connected layers then convert the deep features into class probability vectors. The class probability distribution formula is used to accurately quantify the probability values ​​of each class, determining the fine-grained classification result. Because the residual connection structure solves the gradient vanishing problem in deep networks, it ensures the effective extraction of deep features. Global average pooling and fully connected layers achieve accurate conversion from features to class probabilities. The class probability distribution formula quantifies the confidence level of each class, improving the feature representation ability and accuracy of class determination for fine-grained classification, providing reliable fine-grained classification results for the third-level validation.

[0253] Furthermore, the forward propagation of the residual connection structure is achieved through the following formula:

[0254] F out =ReLU(W2 ReLU(W1 F in +b1)+b2)+F in

[0255] Wherein, the F in Represents the input feature map, the F out This represents the output feature map, where W1 and W2 represent the weight matrices of the convolutional layer. This represents a convolution operation, where b1 and b2 represent the bias vectors of the convolutional layer, ReLU(·) represents the rectified linear activation function, and +F at the end of the formula... in This represents a shortcut connection for identity mapping.

[0256] The residual connection structure extracts the input feature map F through two convolution operations and the ReLU activation function. in The residual features are then compared with the original input feature map F. in By performing identity mapping and adding the results, we obtain the output feature map F. outThis approach enables the effective transfer of deep features, avoiding the vanishing gradient problem during the training of deep convolutional neural networks. The quick connection via identity mapping achieves the fusion of original input features and residual features, effectively solving the vanishing gradient problem in the ResNet-50 model during deep feature extraction. This ensures the model's ability to extract deep features from candidate target region images, resulting in more complete and distinctive deep features. It further optimizes the accuracy of fine-grained classification results, providing a reliable foundation for subsequent level-3 validation.

[0257] Furthermore, the text prompt constructed in step S2043 includes a description of the image region, multiple top-ranked categories of the fine classification results, preliminary categories and confidence levels in the preliminary detection results, the robot's location information in the global map, and task scenario information; the multimodal large model is a pre-trained image-text multimodal processing language large model.

[0258] The constructed text prompts need to integrate image region descriptions, the top categories of fine classification results, the preliminary categories and confidence levels in the preliminary detection results, the robot's global map location information, and task scenario information. Furthermore, the multimodal large model used for verification is a pre-trained image-text multimodal processing language large model, providing comprehensive and multidimensional judgment criteria for cross-modal semantic verification.

[0259] Furthermore, in step S205, the horizontal position coordinates of the target object in the three-dimensional world coordinate system are obtained by back-projection calculation using the two-dimensional pixel coordinates of the target object in the panoramic image, combined with the internal parameters of the vision sensor and the pose transformation matrix of the robot.

[0260] In the target localization step, the two-dimensional pixel coordinates of the target object in the panoramic image are first determined. Then, the influence of imaging system parameters on the pixel coordinates is eliminated by combining the internal parameters of the vision sensor. The robot's pose transformation matrix is ​​used to achieve coordinate transformation from the image pixel coordinate system to the robot's carrier coordinate system, and then to the three-dimensional world coordinate system. The horizontal position coordinates of the target object are obtained through back projection calculation. Because the internal parameters of the vision sensor and the robot's pose transformation matrix are integrated, a multi-dimensional coordinate transformation from two-dimensional pixel coordinates to three-dimensional world coordinates is achieved. The back projection calculation completes the accurate mapping from the pixel position to the actual working environment position, determining the horizontal position of the target object in the actual working environment, and providing accurate spatial position information of the target object for the robot's subsequent operation decisions.

[0261] Furthermore, the specific formula for back projection calculation in step S205 is as follows:

[0262]

[0263] Among them, the Represents a homogeneous world coordinate column vector, the and The horizontal coordinates of the target object in the world coordinate system, the The column vector representing the homogeneous form of the panoramic image pixel coordinates, the and The pixel coordinates of the target object in the panoramic image, The matrix representing the inverse of the intrinsic parameter matrix of the vision sensor, the This represents the homogeneous transformation matrix from the robot's carrier coordinate system to the world coordinate system.

[0264] The panoramic image pixel coordinates and world coordinates of the target object are converted into homogeneous form. First, the homogeneous pixel coordinates are compared with the inverse matrix of the visual sensor intrinsic parameter matrix. Multiply to eliminate the influence of the vision sensor's intrinsic parameters on the imaging, and then combine with the homogeneous transformation matrix from the robot's coordinate system to the world coordinate system. Multiplication, through matrix multiplication operations, achieves spatial transformation of coordinates, ultimately extracting the horizontal coordinates from the homogeneous world coordinates. , Because the coordinate transformation is quantitatively calculated through homogeneous matrix operations, the inverse matrix of the intrinsic parameter matrix accurately eliminates parameter interference from the imaging system, and the homogeneous transformation matrix achieves precise coordinate system transformation, ensuring the mathematical accuracy of the coordinate transformation and significantly improving the calculation accuracy of the world horizontal coordinates of the target object in the working environment, making the robot's positioning of the target object more accurate.

[0265] Furthermore, the robot body is a wheel-legged hybrid mobile robot, which includes a wheel-legged chassis for movement and obstacle crossing, at least one robotic arm for performing grasping operations, a sorting storage bin for holding different types of waste, and a waste compression mechanism for the storage bin.

[0266] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for autonomous operation of a robot, characterized in that, Includes the following steps: Step S100, Task Reception and Path Determination: Receive sanitation tasks, obtain task information and global map corresponding to the sanitation tasks, and generate a globally optimal patrol path that covers the target area and connects key nodes based on the task information and global map. Step S200, Path Patrol and Target Recognition: Control the robot to navigate along the globally optimal patrol path; during the navigation process, collect panoramic images of the robot's surrounding environment, identify target objects, and determine the world coordinates of the target objects in the working environment; Step S300, Cleaning Task Generation and Execution: If the identified target object is garbage to be cleaned, a cleaning sub-task is generated and executed based on the world coordinates of the target object, the current location of the robot, and the category of the target object; The cleaning sub-task includes: planning a sub-path to navigate to the target object, performing a grasping operation, and storing the garbage according to its category into the corresponding storage bins of the robot body; Step S400, Cyclic Operation and Status Judgment: After completing the cleaning sub-task, control the robot to return to the globally optimal patrol path and continue to execute steps S200 and S300; at the same time, monitor the robot's power status and the storage box's capacity status in real time. Step S500, Dumping Task Triggering and Execution: When the capacity status of the storage box is detected to reach a preset threshold, a dumping sub-task is generated, and the robot is controlled to navigate to the nearest garbage collection point to perform garbage dumping; Step S600, Charging Task Triggering and Execution: When the robot's battery level is detected to be below a preset threshold, a charging sub-task is generated, and the robot is controlled to navigate to the nearest charging point to perform charging. Step S700, Task Completion Processing: When the sanitation task is completed or the termination conditions are met, control the robot to go to the garbage collection point to dump the remaining garbage in the storage box, then go to the charging point to charge, and report the task completion information.

2. The method according to claim 1, characterized in that... Step S100 includes: Step S101, obtaining task information and a global map corresponding to the sanitation task; the task information includes a global operation area and a target area located within the global operation area; the global map contains multiple preset key nodes; Step S102: Based on the global operation area and the target area, the global map is rasterized to obtain a raster map; attribute information is configured for each grid in the raster map, the attribute information including at least terrain type, target area marker for indicating whether it belongs to the target area, and key node marker for marking key nodes; based on the key node markers and the global operation area, a set of valid key nodes located within the global operation area is selected; Step S103: Identify the target area from the grid map based on the target area marker, and calculate the task demand intensity of each grid in the target area according to historical task data; divide the target area into multiple sub-areas of different importance levels according to the task demand intensity; perform coverage path planning using different coverage parameters for sub-areas of different importance levels, and generate a set of target area coverage path segments. Step S104: Spatial clustering of the set of effective key nodes and the set of target area coverage path segments is performed to obtain multiple clusters; for each cluster, a shortest intra-cluster sub-path is planned to connect all effective key nodes and target area coverage path segments in the cluster; then, an inter-cluster path is planned to connect adjacent clusters, and all the shortest intra-cluster sub-paths are connected to form an initial path framework. Step S105: Divide the initial path framework into multiple path segments, and label each path segment with its path type and path information; wherein, the path type is determined according to the terrain type of the grid through which the path segment passes; verify each path segment based on preset accessibility constraints, and perform local replanning on path segments that do not meet the constraints to obtain a feasible path that meets all accessibility constraints. Step S106: Construct a comprehensive cost function to evaluate the feasible path, and with the goal of minimizing the comprehensive cost, optimize the feasible path under the preset global constraints to obtain the globally optimal patrol path. Step S107: Determine whether the globally optimal patrol path satisfies the operation time limit constraint; if not, adjust the weight coefficient of the comprehensive cost function and re-execute the optimization solution process of step S106 until the globally optimal patrol path that satisfies the operation time limit constraint is obtained.

3. The method according to claim 1, characterized in that, Step S200 includes: Step S201: Image acquisition, acquiring panoramic images of the environment surrounding the robot; Step S202: Image enhancement. Perform enhancement preprocessing on the panoramic image targeting small targets to obtain a preprocessed image. Step S203: Extraction of Region of Interest. The preprocessed image is semantically segmented to extract the preset region of interest, and the region of interest is divided into multiple sub-images. Step S204: Collaborative recognition. Perform multi-level collaborative recognition on each sub-image to determine whether it contains a target object and to determine its category. The collaborative recognition includes preliminary detection based on a convolutional neural network, and cross-modal semantic verification by calling a multimodal large model based at least on the preliminary detection results. Step S205: Target localization. Based on the position of the target object in the panoramic image and the robot's localization information, calculate the world coordinates of the target object in the working environment.

4. The method according to claim 1, characterized in that, In step S300, generating and executing the cleanup subtask specifically includes: Based on the world coordinates of the target object and the current location of the robot, plan a sub-path from the current location to the target object; Control the robot to move along the sub-path to the target object; Based on the category of the target object, the robot arm adapted to the robot body is invoked to perform a grasping operation; The robotic arm is controlled to store the grasped target object into the classification bin corresponding to the target object's category in the robot's storage box.

5. The method according to claim 1, characterized in that, In step S400, the real-time monitoring of the robot's battery status and the storage box's capacity status specifically includes: The robot's power management system periodically acquires the remaining battery power percentage as the power status. The current waste weight or filling rate is obtained by using a weighing sensor or vision sensor installed in the storage box, which serves as the capacity status.

6. The method according to claim 1, characterized in that, Step S500 specifically includes: When the capacity of the storage box reaches a preset full-load threshold, a dumping task is triggered. Based on the robot's current location and the location information of all garbage collection points in the global map, a first navigation path to the nearest garbage collection point is planned; The robot is controlled to move along the first navigation path, and upon arriving at the waste collection point, the storage box is controlled to perform an emptying operation.

7. The method according to claim 1, characterized in that, Step S600 specifically includes: When the robot's battery level falls below a preset low battery threshold, a charging task is triggered. Based on the robot's current location and the location information of all charging points in the global map, a second navigation path to the nearest charging point is planned; The robot is controlled to move along the second navigation path, and upon arrival at the charging point, it automatically docks with the charging interface and performs the charging process.

8. The method according to claim 1, characterized in that, In step S700, the termination conditions include: reaching the preset total operation time of the sanitation task, or not identifying any more garbage to be cleaned during the complete traversal of the globally optimal patrol path.

9. The method according to claim 1, characterized in that, Step S700 specifically includes: Step S701, Final Disposal: Control the robot to navigate to the nearest waste collection point and empty all the waste in the storage bin; Step S702, Final Charging: Control the robot to navigate from the waste collection point to the nearest charging point and perform a charging operation until the battery level reaches a preset saturation value; Step S703, Information Reporting: Generate an execution report for the sanitation task and send it to the cloud. The execution report shall include at least the task identifier, total operation time, total amount of garbage cleaned, and final status.