Target object detection processing system and method
By combining a multi-view vision module and a safety judgment module, high-precision object detection and safety assurance of cleaning robots in complex environments are achieved, solving the safety risks of large and medium-sized cleaning robots in high-traffic environments and ensuring the efficient and safe execution of cleaning tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO WENDAO ROBOT TECHNOLOGY CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-01
AI Technical Summary
Existing large and medium-sized cleaning robots lack sufficient object detection and safety performance in complex environments with high traffic, leading to misjudgments and safety risks.
Image data is acquired using a multi-view vision module, and combined with a distance measurement module and a security determination module to classify, identify, and determine the security attributes of target objects. The object detection capability and security performance are improved through semantic map construction and behavior prediction mechanisms.
It improves the object detection accuracy and safety of cleaning robots in complex environments, avoids the risk of collisions and falls, and ensures efficient and safe execution of cleaning tasks.
Smart Images

Figure CN121962870A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more specifically, to a target object detection and processing system and method. Background Technology
[0002] With the development of AI technology, the level of robot intelligence is constantly improving, and cleaning robots are rapidly developing and being applied. Cleaning robots are suitable for a variety of indoor and outdoor environments, such as commercial office buildings, government office buildings, office parks, hospitals, universities, stadiums and exhibition halls, traffic roads, and industrial workshops. They can efficiently complete intelligent cleaning tasks on different ground materials (such as paved roads, cement roads, washboard roads, and red brick roads).
[0003] Compared to small household cleaning robots, large and medium-sized cleaning robots need to cope with complex environments with large areas and high traffic, thus requiring higher levels of object detection and safety performance.
[0004] Cleaning robots need to autonomously clean tens of thousands of square meters of shopping mall space, and also need to support multi-floor navigation and automatic obstacle avoidance. Cleaning robots typically rely on laser positioning and recognition. Defects in environmental perception can lead to certain safety risks. For example, in strong light or reflective environments, they may misjudge the edge of an escalator as a flat surface, causing the robot to fall from a height, potentially damaging the equipment or injuring people below. Transparent / reflective surfaces such as glass doors and mirrors may be misjudged as open spaces, causing the robot to collide with or fall down the stairs. There may also be a delay in recognizing moving objects (such as pedestrians), which could cause collision injuries.
[0005] Therefore, how to effectively improve the higher-level object detection and safety performance of cleaning robots and realize the intelligent implementation of cleaning robots is an urgent problem to be solved. Summary of the Invention
[0006] The main objective of this invention is to disclose a target object detection and processing system and method, so as to at least solve the problem in the related art of how to effectively improve the higher-level object detection performance and safety assurance performance of large and medium-sized cleaning robots in complex environments with large areas and high traffic.
[0007] According to one aspect of the present invention, a target object detection and processing system is provided.
[0008] The target object detection and processing system according to the present invention includes: a target detection module, configured to acquire image data collected by a multi-view vision module, and obtain a target detection result based on the image data, wherein the target detection result includes: category information and location information corresponding to the one or more target objects; a distance calculation module, configured to acquire depth value information based on the image data, determine the coordinate information of the one or more target objects in the camera coordinate system based on the target detection result and the depth value information, and convert the coordinate information in the camera coordinate system into three-dimensional coordinate information in the world coordinate system; and a security determination module, configured to determine the security attribute information of the current target object when the category information corresponding to the current target object satisfies a first predetermined category, in conjunction with the currently constructed semantic map, based on the location information corresponding to the current target object and the time-series changes in location information.
[0009] According to another aspect of the present invention, a target object detection and processing method is provided.
[0010] The target object detection processing method according to the present invention includes: acquiring image data collected by a multi-view vision module, obtaining target detection results based on the image data, wherein the target detection results include: category information and location information corresponding to the one or more target objects; acquiring depth value information based on the image data, determining the coordinate information of the one or more target objects in the camera coordinate system based on the target detection results and the depth value information, and converting the coordinate information in the camera coordinate system into three-dimensional coordinate information in the world coordinate system; when the category information corresponding to the current target object satisfies a first predetermined category, determining the security attribute information of the current target object based on the location information corresponding to the current target object and the change of location information over time, in conjunction with the currently constructed semantic map.
[0011] According to the present invention, the target detection module can classify and identify the location of target objects, the distance calculation module can determine the spatial three-dimensional coordinates of the target object based on the target detection results of the target detection module, and the safety judgment module can determine the safety attributes of the current target object based on the target's location information and the changes in the location information over time, through semantic-level environmental understanding combined with semantic map construction and behavior prediction mechanisms. This can further improve the higher-level object detection performance and safety assurance performance of the cleaning robot. Attached Figure Description
[0012] Figure 1 This is a structural block diagram of a target object detection and processing system according to an embodiment of the present invention;
[0013] Figure 2 This is a structural block diagram of a target object detection and processing system according to a preferred embodiment of the present invention;
[0014] Figure 3 This is a flowchart of the business execution process of the security determination module according to a preferred embodiment of the present invention;
[0015] Figure 4 This is a flowchart of a target object detection and processing method according to an embodiment of the present invention. Detailed Implementation
[0016] The specific implementation of the present invention will now be described in detail with reference to the accompanying drawings.
[0017] According to an embodiment of the present invention, a target object detection and processing system is provided.
[0018] Figure 1 This is a structural block diagram of a target object detection and processing system according to an embodiment of the present invention. Figure 1 As shown, the target object detection and processing system includes:
[0019] The target detection module 10 is used to acquire image data collected by the multi-view vision module and obtain target detection results based on the image data. The target detection results include: category information and location information corresponding to one or more target objects.
[0020] The distance calculation module 12 is used to obtain depth information based on the above image data, determine the coordinate information of one or more target objects in the camera coordinate system according to the above target detection results and the above depth information, and convert the coordinate information in the camera coordinate system into three-dimensional coordinate information in the world coordinate system.
[0021] The security determination module 14 is used to determine the security attribute information of the current target object when the category information corresponding to the current target object meets the first predetermined category, in combination with the currently constructed semantic map, based on the location information corresponding to the current target object and the changes in the location information over time.
[0022] like Figure 1In the target object detection and processing system shown, the target detection module 10 obtains target detection results based on image data collected by the multi-view vision module. These results include the category information (pedestrians, escalators, garbage, etc.) and location information of one or more target objects, enabling target object classification and location recognition. The distance calculation module 12 obtains the spatial three-dimensional coordinates of the target object based on the detection results, supporting subsequent path planning and safety strategy judgment to avoid collisions and rationally plan cleaning areas. The safety judgment module 14, combining semantic map construction and behavior prediction mechanisms, determines the safety attributes of the current target based on its location information and the changes in its location information over time, further enhancing the cleaning robot's higher-level object detection capabilities and safety performance. This ensures that the cleaning robot can safely and efficiently perform tasks in various operating scenarios while guaranteeing the safety of the surrounding environment.
[0023] In the preferred implementation, the target detection module 120 is responsible for detecting target objects, classifying them, and identifying their locations. Specifically, the target detection module 120 can use deep learning algorithms to detect target objects such as pedestrians, escalators, and garbage, obtain the image location of the target objects, and pass it to the distance calculation module.
[0024] The object detection module takes image data acquired by the binocular vision module as input and outputs bounding boxes containing multiple targets. It is primarily used to determine the category and location information of multiple objects in an image. The module performs object detection based on a deep learning algorithm using neural networks, achieving end-to-end object detection result output.
[0025] Preferably, the distance measurement module 12 may further include the following units: a depth value calculation unit 120, used to calculate the depth value information based on the disparity information of the image data acquired by the multi-view vision module; a coordinate information determination unit 122, used to determine the coordinate information of the one or more target objects in the camera coordinate system using the camera intrinsic and extrinsic parameters in the multi-view vision module; and a coordinate information transformation unit 124, used to transform the coordinate information of the one or more target objects in the camera coordinate system to the world coordinate system through the extrinsic parameter matrix, so as to obtain the three-dimensional coordinate information of the one or more target objects in the world coordinate system.
[0026] In the preferred implementation process, the distance calculation module 122 performs coordinate system transformation based on the image position information detected by the target detection module to obtain the actual position of the target object relative to the cleaning robot.
[0027] Specifically, the distance calculation module 122 is based on a multi-view vision module (e.g., a binocular vision module). Taking a binocular vision module as an example, it uses the binocular vision module and camera calibration parameters to accurately locate the target object in three-dimensional space. The distance calculation module calculates the disparity of the target object in the left and right images based on the binocular images to obtain the target object's depth information. Combining the target detection results obtained by the target detection module with the aforementioned depth information calculated from the binocular images, the distance calculation module calculates the target's position information in the camera coordinate system and converts this position information into three-dimensional spatial coordinates in the world coordinate system, which is used to support subsequent path planning and safety strategy judgment.
[0028] The process of solving the three-dimensional coordinate position in space mainly includes the following steps:
[0029] Step 1: Estimate depth values based on disparity information from binocular images;
[0030] Step 2: Use camera intrinsic and extrinsic parameters to convert pixel coordinates into 3D coordinates in the camera coordinate system;
[0031] Step 3: Transform the results into the world coordinate system using the extrinsic parameter matrix to obtain the three-dimensional coordinates of the target object in the world coordinate system.
[0032] Preferably, such as Figure 2 As shown, the target object detection system may further include an information assistance module 16; wherein the information assistance module 16 is connected to the target detection module 10, the distance calculation module 12, and the security determination module 14 respectively, and is used to perform log information assistance operations, error information reporting SDK operations, and configuration parameter information processing operations.
[0033] The aforementioned information assistance module provides real-time feedback on the operating status and timely error information throughout the entire operation of the deep learning system.
[0034] Specifically, the information assistance module 16 is mainly used for: log information assistance, error information reporting SDK, and configuration parameter information processing, thereby assisting other modules in their work and completing real-time error information uploading.
[0035] Log information assistance: The main process logs include information such as successful initialization, test output, runtime, and version information, which facilitates problem location and debugging.
[0036] Error message reporting: Error messages are collected in logs according to the error level, and server reporting is triggered using trigger flags; for example, error levels are divided into three fault levels: low, medium and high for alarm purposes. For example, low image quality and image frame loss are low fault levels, while model errors, model loss, and abnormal binocular module parameters are medium fault levels.
[0037] Configuration parameter information: Read the configuration file to configure the hazardous attribute area parameters and dynamic attribute determination parameters. For example, the hazardous attribute area parameters include: hazardous attribute area distance parameter D, and the angular range A covered by the hazardous attribute area. Dynamic attribute parameters include: the radius R of the predetermined area and the time series length.
[0038] The safety attribute information judged by the aforementioned safety determination module may further include: whether the current target object is within the danger attribute area, and the status attribute information of the target object within the danger attribute area.
[0039] Preferably, such as Figure 2 As shown, the aforementioned security determination module 14 may further include:
[0040] The first determination unit 140 is used to determine whether the current target object is located within a dangerous attribute area according to predetermined determination conditions. The predetermined determination conditions include: whether the distance between the target object and the cleaning robot is less than the distance of the dangerous attribute area, and whether the angle between the target object and the cleaning robot is within the angle range covered by the dangerous attribute area.
[0041] The second determination unit 142 is used to determine whether the target object has not crossed the predetermined area for multiple consecutive frames when it is determined that the current target object is located in the dangerous attribute area. If not, the state attribute information of the target object is determined to be the first state attribute, and the current safety attribute information of the target object is returned in real time. In this case, part or all of the predetermined area is within the dangerous attribute area.
[0042] The third determination unit 144 is used to continue monitoring the current target object when the current target object's state attribute information is determined to be the first state attribute, and when the current target object is stable within the predetermined area for multiple consecutive frames, the state attribute information of the target object is determined to be updated to the second state attribute, and the current security attribute information of the target object is returned in real time.
[0043] Preferably, such as Figure 2 As shown, the aforementioned security determination module may further include:
[0044] The fourth determination unit 146 is used to determine and return the predetermined area as the first attribute area in real time when it is determined that the target object is located in the dangerous attribute area and the target object has not crossed the predetermined area for multiple consecutive frames.
[0045] The fifth determination unit 148 is used to continue monitoring the current target object when the predetermined area is determined to be the first attribute area. If the current target object moves out of the predetermined area in a frame of multiple consecutive frames, the predetermined area is determined to be converted from the first attribute area to the second attribute area, and the predetermined area is returned as the second attribute area in real time.
[0046] The sixth determination unit 150 is used to continue monitoring the predetermined area when the predetermined area is determined to be the first attribute area. If there is no target object in the predetermined area within multiple consecutive frames, the predetermined area is converted from the first attribute area to the third attribute area, and the predetermined area is returned as the third attribute area in real time.
[0047] In the preferred implementation process, when the target object detected by the target detection module meets the first predetermined category (e.g., movable object, such as pedestrian, pet, etc.), the safety strategy module needs to perform a safety assessment on the target object's position, determine the safety attribute information of the current target object, and return the current safety attribute information of the target object in real time. This can prevent the cleaning robot from colliding and rationally plan the cleaning area, ensuring that the robot can perform tasks safely and efficiently in various operating scenarios, while ensuring the safety of the surrounding environment.
[0048] The safety assessment module, combining semantic map construction and behavior prediction mechanisms, is responsible for determining the safety attributes of the target object based on its location information and its time-series location changes. These safety attributes include: 1. Whether the target object is within a hazardous area; 2. The state attributes of the target object within the hazardous area (e.g., dynamic or static). Safety attributes are configurable parameters, as detailed in the subsequent description of the information assistance module.
[0049] The aforementioned hazardous attribute area can be preset. For example, when the category information corresponding to the current target object meets the first predetermined category, a sector area with radius D and angle A is set as the hazardous attribute area. That is, the distance parameter of the hazardous attribute area is D, and the angular range covered by the hazardous attribute area is A.
[0050] Of course, you can also set up a safety attribute area. For example, you can set up a region within a predetermined range surrounding the boundary of the danger attribute area as a safety attribute area.
[0051] The aforementioned predetermined region can be established in various ways. For example, a circular predetermined region can be set with the center point of the lower edge of the target object as the center and R as the radius of activity. Of course, setting predetermined regions of other shapes based on the outline boundary of the target object is also within the scope of protection of this application. If the target object is within the activity range for multiple consecutive frames (i.e., a preset time series length, for example, 5 consecutive frames), then the region is determined to be the first attribute region (or it can be defined as a static region). If the target object crosses the boundary, then the state attribute information of the target object is determined to be the first state attribute (or the state attribute of the target object can be defined as a dynamic state attribute). If the dynamic region increases, then the state of the dynamic region needs to be continuously monitored.
[0052] It should be noted that the above-mentioned dangerous attribute area parameters D and A are configurable, the dynamic attribute parameter R is configurable, and the time series length (i.e., the duration of multiple consecutive frames) is configurable, all of which can be pre-configured by the above-mentioned information assistance module 16.
[0053] The following combination Figure 3 The preferred embodiments described above are further described below.
[0054] Figure 3 This is a flowchart illustrating the business execution process of the security determination module according to a preferred embodiment of the present invention. Figure 3 As shown, the security determination module performs business operations, mainly including the following steps:
[0055] Step S301: Determine whether the target object (e.g., a pedestrian, pet, etc.) is located within the hazardous area. The determination criteria are whether the distance between the target object and the cleaning robot is less than the hazardous area distance D, and whether the angle between the target object and the cleaning robot is within the angular range A covered by the hazardous area. If the distance between the target object and the cleaning robot is less than the hazardous area distance, and the angle between them is within the angular range covered by the hazardous area, then proceed to step S303. Otherwise, proceed to step S302.
[0056] Step S302: Determine that the target object is located in a non-dangerous attribute area (i.e., a safe attribute area), and return the information that the target object is located in a safe attribute area in real time.
[0057] Step S303: Establish a predetermined area with the center point of the lower edge of the target object as the center and R as the activity radius. Determine whether the target object has not exceeded the predetermined area for multiple consecutive frames; if so, proceed to step S306; otherwise, proceed to step S304.
[0058] It should be noted that, since the target object is located within a hazardous area, part or all of the predetermined area defined for the target object is located within the aforementioned hazardous area.
[0059] Step S304: When the target object moves out of the predetermined area in one of the consecutive frames (e.g., 5 consecutive frames), the state attribute of the target object is determined to be the first state attribute (also called the dynamic state attribute), and the state attribute of the target object is returned in real time as the first state attribute, and step S305 is executed.
[0060] Step S305: When the state attribute of the target object is the first state attribute (also known as the dynamic state attribute), continue to monitor the target object. When the target object is stable within its corresponding judgment area for multiple consecutive frames, determine that the state attribute of the target object is the second state attribute (also known as the static state attribute), and return the state attribute of the target object as the second state attribute in real time.
[0061] Therefore, the state attributes of a target object are not static, but can be dynamically adjusted based on the monitoring of the target object. For example, they can be transformed from dynamic state attributes to static state attributes, thus dynamically managing and updating the safety attribute information of the target object.
[0062] Step S306: If the target object does not go beyond the predetermined area for multiple consecutive frames (e.g., 5 consecutive frames), then the predetermined area is determined to be the first attribute area (also called the static area), and steps S307 and S308 are executed.
[0063] Step S307: When the predetermined area is the first attribute area (also known as the static attribute area), continue to monitor the target object. If the target object moves out of the predetermined area in one of the consecutive frames, then the predetermined area is determined to be changed from the first attribute area to the second attribute area (also known as the dynamic attribute area), and the predetermined area is returned as the second attribute area in real time.
[0064] Step S308: When the predetermined area is the first attribute area (also known as the static attribute area), continue to monitor the predetermined area. When there is no target object in the predetermined area for multiple consecutive frames, the determination area is changed from the first attribute area (also known as the static attribute area) to the third attribute area (also known as the non-dangerous attribute area or the safe attribute area), and the predetermined area is returned as the third attribute area in real time.
[0065] Therefore, the attributes of a predetermined area are not static, but can be dynamically adjusted based on the monitoring of the predetermined area. For example, the attributes can be changed from static to dynamic, or from static to safe. At this time, areas within the predetermined area that are partially or entirely located within dangerous attribute areas will no longer be considered dangerous attribute areas. Thus, dynamic management and updating of dangerous attribute areas are also realized.
[0066] According to an embodiment of the present invention, a target object detection and processing method is also provided.
[0067] Figure 4 This is a flowchart of a target object detection and processing method according to an embodiment of the present invention. Figure 4 As shown, the target object detection processing method includes:
[0068] Step S401: Acquire image data collected by the multi-view vision module, and obtain target detection results based on the image data. The target detection results include: category information and location information corresponding to one or more target objects.
[0069] Step S403: Based on the above image data, obtain depth information, determine the coordinate information of one or more target objects in the camera coordinate system according to the above target detection results and the above depth information, and convert the coordinate information in the camera coordinate system into three-dimensional coordinate information in the world coordinate system;
[0070] Step S405: When the category information corresponding to the current target object meets the first predetermined category, combine the currently constructed semantic map and determine the security attribute information of the current target object based on the location information corresponding to the current target object and the changes in the location information over time.
[0071] Adopting such Figure 4 The target object detection and processing method shown obtains target detection results based on image data collected by a multi-view vision module. These results include the category information (pedestrians, escalators, garbage, etc.) and location information of one or more target objects, enabling target object classification and location recognition. The method then obtains the spatial three-dimensional coordinates of the target objects based on the detection results, supporting subsequent path planning and safety strategy judgments to avoid collisions and rationally plan cleaning areas. Combined with semantic map construction and behavior prediction mechanisms, the method determines the safety attributes of the current target based on its location information and changes in location information over time, further enhancing the cleaning robot's higher-level object detection capabilities and safety performance. This ensures that the cleaning robot can safely and efficiently perform tasks in various operating scenarios while guaranteeing the safety of the surrounding environment.
[0072] The aforementioned safety attribute information includes, but is not limited to: whether the current target object is within the danger attribute area, and the status attribute information of the target object within the danger attribute area.
[0073] Preferably, in step S405, the above-mentioned determination of the security attribute information of the current target object based on the location information corresponding to the current target object and the changes in the location information over time, combined with the currently constructed semantic map, may further include the following processing:
[0074] According to predetermined judgment conditions, it is determined whether the current target object is located within the dangerous attribute area. The predetermined judgment conditions include: whether the distance between the target object and the cleaning robot is less than the distance of the dangerous attribute area, and whether the angle between the target object and the cleaning robot is within the angle range covered by the dangerous attribute area.
[0075] If it is determined that the current target object is located within the dangerous attribute area, it is determined whether the target object has not crossed the predetermined area for multiple consecutive frames. If not, the state attribute information of the target object is determined to be the first state attribute, and the current safety attribute information of the target object is returned in real time.
[0076] When the current target object's state attribute information is determined to be the first state attribute, the current target object is monitored. When the current target object is stable within the predetermined area for multiple consecutive frames, the state attribute information of the target object is updated to the second state attribute, and the current security attribute information of the target object is returned in real time.
[0077] Preferably, step S405 may further include the following processing:
[0078] If it is determined that the current target object is located within the dangerous attribute area, and it is determined that the target object has not crossed the predetermined area for multiple consecutive frames, the predetermined area is determined and returned in real time as the first attribute area.
[0079] When the predetermined area is determined to be the first attribute area, the current target object is monitored. If the current target object moves out of the predetermined area in one of the consecutive frames, the predetermined area is determined to be changed from the first attribute area to the second attribute area, and the predetermined area is returned as the second attribute area in real time.
[0080] When the predetermined region is determined to be the first attribute region, the predetermined region is monitored. If no target object exists in the predetermined region within multiple consecutive frames, the predetermined region is converted from the first attribute region to the third attribute region, and the predetermined region is returned as the third attribute region in real time.
[0081] The preferred implementation described above will be further described below, taking a pedestrian (belonging to one of the first predetermined categories) as an example.
[0082] Example 1: Pedestrian Detection
[0083] Images are acquired using a binocular vision module based on a cleaning robot. A target detection module (such as a deep neural network) is then used to detect target objects in the acquired images. The detection results include bounding boxes in the 2D image, primarily used to determine the category and location information of multiple target objects in the image. High-precision identification and tracking of pedestrian targets detected in the images are then performed.
[0084] The distance calculation module obtains the 3D position information of the pedestrian target based on the pedestrian target position information detected by the target detection module. Specifically, the distance calculation module estimates the depth value of the pedestrian target based on the disparity information of the binocular images, converts the pixel coordinates of the pedestrian target into 3D coordinates in the camera coordinate system using camera intrinsic and extrinsic parameters, and then converts the result to the world coordinate system through the extrinsic parameter matrix to obtain the 3D coordinates of the pedestrian target in the world coordinate system. The distance calculation module transmits this information to the safety strategy module in real time. Since the category information corresponding to the current target object (i.e., pedestrian) meets the first predetermined category, the safety strategy module needs to perform a safety assessment of the pedestrian target, determine the safety attribute information of the current pedestrian target, and return the status information of the pedestrian target. Then, based on the information returned by the safety strategy module, tasks such as dynamic obstacle avoidance and area risk assessment are performed to improve the collaborative safety between humans and machines.
[0085] Preferably, before obtaining the target detection result based on the image data acquired by the multi-view vision module, the following processing may be included: acquiring image data of target objects corresponding to the second predetermined category in real time using the multi-view vision module; performing semantic segmentation operation based on the acquired image data to identify multiple structural feature information of the target objects corresponding to the second predetermined category; in the semantic map established based on Simultaneous Localization and Mapping (SLAM), marking one or more detected areas as impassable areas based on the multiple structural feature information, dynamically updating the boundaries of the impassable areas, and setting a predetermined range of safe buffer zones around the boundaries of the areas.
[0086] When the category information corresponding to the current target object meets the above-mentioned second predetermined category, the following processing may also be included: when it is determined that the robot has entered the above-mentioned safety buffer zone, the deceleration mode is activated and the path is automatically replanned to avoid the robot entering the above-mentioned impassable area; when it is determined that the cleaning robot has entered the above-mentioned impassable area, the emergency stop function is activated, and the tilt angle change is detected by data fusion of the inertial measurement unit (IMU) to help judge the risk of fall.
[0087] The preferred implementation described above is further described below using an escalator (belonging to one of the second predetermined categories) as an example. This method is aimed at escalator areas in multi-level spatial scenarios, and adopts a structured recognition scheme based on image semantic parsing to construct a fall prevention strategy for high-risk areas at the edge of the escalator in a semantic map.
[0088] Specifically, a visual model is needed to identify the structural features of the escalator, map the detection results to the currently constructed environmental map, and mark the area as an impassable semantic region. When the cleaning robot's navigation path approaches the edge of the escalator, the system automatically replans the path to ensure equipment safety.
[0089] For example, images of the escalator area can be acquired in real time using a binocular vision module (such as an RGB-D camera). Semantic segmentation can be performed using a model (e.g., an improved YOLOv8 model) to identify escalator steps, edges, and railing structures. In the semantic map built based on SLAM, the detected escalator edges are marked as impassable areas using a rasterization algorithm, and the boundaries of the aforementioned risk areas are dynamically updated. For example, specific raster markers can be set for impassable areas, and a safety buffer zone of a predetermined range (e.g., 0.3 meters) can be set around the boundary of the area.
[0090] During the cleaning robot's operation, images are acquired using its binocular vision module. A target detection module then analyzes these images, identifying escalator targets based on their structural features. The distance calculation module calculates the escalator's 3D location information based on the detected target's location. Combining this with markers of impassable areas in the semantic map, when the cleaning robot approaches the escalator, a deceleration mode (e.g., a 30% speed reduction) can be activated when it enters a safety buffer zone. This automatically replans the path, preventing the robot from entering impassable areas and ensuring equipment safety. If the robot is determined to have entered an impassable area, an emergency stop is activated, and IMU data fusion is used to detect tilt changes to aid in assessing the risk of a fall.
[0091] The above functions, in conjunction with the semantic map system, enable the cleaning robot to have scene understanding and safety protection capabilities in complex buildings, effectively preventing physical risk events such as falls.
[0092] Preferably, when the category information corresponding to the current target object meets the above-mentioned third predetermined category, the following processing may also be included: generating a cleaning area annotation map containing the location, category, and size of the current target object in the semantic map established based on SLAM, and marking the cleaning priority in the cleaning area annotation map; performing path planning operation according to the target detection results, and dynamically adjusting the cleaning path and operation strategy.
[0093] The preferred implementation described above will be further described below using the example of the target object being garbage (belonging to one of the third predetermined categories).
[0094] The waste recognition function detects common litter on the ground, such as paper scraps, plastic bags, dirt, and cigarette butts, improving the intelligence and precision of the cleaning robot in actual cleaning tasks. In the deep learning system, images are acquired by the cleaning robot's binocular vision module. A target detection module (such as a deep learning model based on a neural network) performs high-confidence recognition and multi-category classification of litter targets in the acquired images. The distance calculation module obtains the three-dimensional spatial location information of the litter targets based on the location information detected by the target detection module. A semantic map based on SLAM is generated, showing the area to be cleaned, including the location, category, and size of the litter. Litter objects can be marked with high, medium, and low cleaning priorities. The above target detection is transmitted in real time to the path planning module for dynamic adjustment of the cleaning path and operation strategy, thereby avoiding omissions and repetitions and improving cleaning efficiency.
[0095] It should be noted that the specific details of the above-mentioned target object detection and processing methods can be found in the relevant references. Figures 1 to 3 The relevant descriptions and effects in the illustrated embodiments are for understanding purposes only and will not be repeated here.
[0096] In summary, the embodiments provided by this invention offer a target object detection and processing system and method. Based on visual perception and semantic map construction technology, it achieves various intelligent functions, including pedestrian recognition, escalator recognition, and garbage recognition, through semantic-level environmental understanding. Furthermore, by combining semantic map construction and behavior prediction mechanisms, it determines the safety attributes of the current target based on its location information and changes in location information over time, further enhancing the cleaning robot's higher-level object detection capabilities and safety performance. This ensures that the cleaning robot can safely and efficiently perform tasks in various operating scenarios while guaranteeing the safety of the surrounding environment.
[0097] The above-disclosed embodiments are merely a few specific examples of the present invention. However, the present invention is not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A target object detection and processing system, characterized in that, include: The target detection module is used to acquire image data collected by the multi-view vision module and obtain target detection results based on the image data. The target detection results include: category information and location information corresponding to the one or more target objects. The distance calculation module is used to obtain depth information based on the image data, determine the coordinate information of one or more target objects in the camera coordinate system according to the target detection result and the depth information, and convert the coordinate information in the camera coordinate system into three-dimensional coordinate information in the world coordinate system. The security determination module is used to determine the security attribute information of the current target object when the category information corresponding to the current target object meets the first predetermined category, in combination with the currently constructed semantic map, based on the location information corresponding to the current target object and the changes in the location information over time.
2. The target object detection and processing system according to claim 1, characterized in that, The distance calculation module further includes: A depth value calculation unit is used to calculate the depth value information based on the disparity information of the image data acquired by the multi-view vision module. The coordinate information determination unit is used to determine the coordinate information of one or more target objects in the camera coordinate system using the camera intrinsic and extrinsic parameter information in the multi-view vision module; The coordinate information transformation unit is used to transform the coordinate information of one or more target objects in the camera coordinate system to the world coordinate system through the extrinsic parameter matrix, so as to obtain the three-dimensional coordinate information of the one or more target objects in the world coordinate system.
3. The target object detection and processing system according to claim 1, characterized in that, The target object detection system also includes: an information assistance module; The information assistance module is connected to the target detection module, distance calculation module, and security determination module, respectively, and is used to perform log information assistance operations, error information reporting software development kit (SDK) operations, and configuration parameter information processing operations.
4. The target object detection and processing system according to claim 1, characterized in that, The safety attribute information includes: whether the current target object is within the danger attribute area, and the status attribute information of the target object within the danger attribute area.
5. The target object detection and processing system according to claim 4, characterized in that, The security determination module further includes: The first determination unit is used to determine whether the current target object is located within a dangerous attribute area according to predetermined determination conditions. The predetermined determination conditions include: whether the distance between the target object and the cleaning robot is less than the distance of the dangerous attribute area, and whether the angle between the target object and the cleaning robot is within the angle range covered by the dangerous attribute area. The second determination unit is used to determine whether the target object has not crossed the predetermined area for multiple consecutive frames when it is determined that the current target object is located in the dangerous attribute area. If not, the state attribute information of the target object is determined to be the first state attribute, and the current safety attribute information of the target object is returned in real time. The predetermined area is part or all of the dangerous attribute area. The third determination unit is used to continue monitoring the current target object when the current target object's state attribute information is determined to be the first state attribute, and when the current target object is stable within the predetermined area for multiple consecutive frames, the unit determines that the target object's state attribute information is updated to the second state attribute and returns the target object's current security attribute information in real time.
6. The target object detection and processing system according to claim 5, characterized in that, The security determination module also includes: The fourth determination unit is used to determine and return the predetermined area as the first attribute area in real time when it is determined that the current target object is located in the dangerous attribute area and the target object has not crossed the predetermined area for multiple consecutive frames. The fifth determination unit is used to continue monitoring the current target object when the predetermined area is determined to be the first attribute area. If the current target object moves out of the predetermined area in a certain frame of multiple consecutive frames, the predetermined area is determined to be converted from the first attribute area to the second attribute area, and the predetermined area is returned as the second attribute area in real time. The sixth determination unit is used to continue monitoring the predetermined region when the predetermined region is determined to be the first attribute region. If no target object exists in the predetermined region within multiple consecutive frames, the predetermined region is converted from the first attribute region to the third attribute region, and the predetermined region is returned as the third attribute region in real time.
7. A method for detecting and processing target objects, characterized in that, include: Image data acquired by a multi-view vision module is obtained, and target detection results are obtained based on the image data. The target detection results include: category information and location information corresponding to the one or more target objects. Depth information is obtained based on the image data, and the coordinate information of one or more target objects in the camera coordinate system is determined according to the target detection result and the depth information. The coordinate information in the camera coordinate system is then converted into three-dimensional coordinate information in the world coordinate system. When the category information corresponding to the current target object meets the first predetermined category, the security attribute information of the current target object is determined by combining the currently constructed semantic map and the location information corresponding to the current target object and the changes in the location information over time.
8. The target object detection and processing method according to claim 7, characterized in that, The process of determining the security attribute information of the current target object based on the location information and the time-series changes in location information, combined with the currently constructed semantic map, includes: According to predetermined judgment conditions, it is determined whether the current target object is located within a dangerous attribute area. The predetermined judgment conditions include: whether the distance between the target object and the cleaning robot is less than the distance of the dangerous attribute area, and whether the angle between the target object and the cleaning robot is within the angle range covered by the dangerous attribute area. If it is determined that the current target object is located within the dangerous attribute area, it is determined whether the target object has not crossed the predetermined area for multiple consecutive frames. If not, the state attribute information of the target object is determined to be the first state attribute, and the current safety attribute information of the target object is returned in real time. When the current target object's state attribute information is determined to be the first state attribute, the current target object is monitored. When the current target object is stable within the predetermined area for multiple consecutive frames, the state attribute information of the target object is updated to the second state attribute, and the current security attribute information of the target object is returned in real time.
9. The target object detection and processing method according to claim 7, characterized in that, Before acquiring image data collected by a multi-view vision module and obtaining target detection results based on the image data, the process further includes: The multi-view vision module is used to collect image data of target objects corresponding to the second predetermined category in real time, and semantic segmentation operation is performed based on the collected image data to identify multiple structural feature information of the target objects corresponding to the second predetermined category. In the semantic map built based on Simultaneous Localization and Mapping (SLAM), one or more detected areas are marked as impassable areas according to the multiple structural feature information, and the boundaries of the impassable areas are dynamically updated. A safe buffer zone of a predetermined range is set around the boundaries of these areas. When the category information corresponding to the current target object satisfies the second predetermined category, it also includes: When it is determined that the robot has entered the safety buffer zone, the deceleration mode is activated and the path is automatically replanned to prevent the robot from entering the impassable area. When the cleaning robot enters the impassable area, the emergency stop function is activated, and the tilt angle change is detected by inertial measurement unit (IMU) data fusion to help determine the risk of falling.
10. The target object detection and processing method according to claim 7, characterized in that, When the category information corresponding to the current target object satisfies the third predetermined category, it also includes: In the semantic map built based on SLAM, a cleaning area annotation map containing the location, category, and size of the current target object is generated, and the cleaning priority is marked in the cleaning area annotation map; Based on the target detection results, a path planning operation is performed to dynamically adjust the cleaning path and operation strategy.