Visual information application method and apparatus, and pool cleaning robot

Through visual information acquisition and image recognition model, combined with inertial measurement unit and encoder, intelligent path planning and cleaning of pool cleaning robots are realized, solving the problem of poor pool cleaning effect, especially the cleaning effect of pool walls and steps.

WO2025153085A1PCT designated stage expired Publication Date: 2025-07-24WYBOTICS CO LTD

Patent Information

Application Number
PCT/CN2025/073160
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-26
Filing Date
2025-01-18
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The existing pool cleaning robots have poor cleaning results when cleaning the pool, especially the cleaning effect of the pool wall and step areas is insufficient, and there is a lack of intelligent path planning.

Method used

The visual information acquisition module and image recognition model are used to identify the targets and obstacles to be cleaned through object detection and semantic segmentation, and path planning is carried out in combination with inertial measurement units and encoders to realize intelligent cleaning of the pool cleaning robot.

Benefits of technology

The cleaning effect of the pool cleaning robot is improved, especially the cleaning ability of the pool wall and steps, and intelligent path planning and dynamic obstacle avoidance are achieved, improving cleaning efficiency and effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073160_24072025_PF_FP_ABST
    Figure CN2025073160_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A visual information application method and apparatus, and a pool cleaning robot. The method comprises: 201, acquiring visual information collected by a pool cleaning robot; 202, controlling the pool cleaning robot on the basis of the visual information; or 203, processing the visual information on the basis of a preset mode.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for applying visual information and pool cleaning robot

[0001] This application claims priority to Chinese patent application No. 2024100781283, filed on January 18, 2024, entitled “Control method, device and underwater cleaning robot for underwater cleaning robot”, and Chinese patent application No. 2024103447824, filed on March 25, 2024, entitled “Target tracking method, device and underwater cleaning robot for underwater cleaning robot”, and Chinese patent application No. 2024109673477, filed on July 18, 2024, entitled “Mapping method, device and underwater cleaning robot for underwater cleaning robot”, and Chinese patent application No. 2024117023095, filed on November 26, 2024, entitled “Control method, device and pool cleaning robot for pool cleaning robot”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of robotics, and in particular to a method, apparatus, device, and pool cleaning robot for utilizing visual information. Background Art

[0003] With the development of computer technology, robotics technology has also developed rapidly. For example, users use sweeping robots to clean the floors of houses, use window cleaning robots to clean the windows of houses, and use pool cleaning robots to clean pools. Summary of the Invention

[0004] The present application provides a method and device for applying visual information, as well as a pool cleaning robot. The technical solution is as follows:

[0005] In one aspect, a method for utilizing visual information is provided, the method comprising:

[0006] Obtain visual information collected by the pool cleaning robot;

[0007] The pool cleaning robot is controlled based on the visual information, or the visual information is processed in a preset manner.

[0008] In one aspect, a device for utilizing visual information is provided, the device comprising:

[0009] A visual information acquisition module is used to acquire visual information collected by the pool cleaning robot;

[0010] An application module is used to control the pool cleaning robot based on the visual information, or to process the visual information in a preset manner.

[0011] On the one hand, a pool cleaning robot is provided, which includes a robot controller, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the method of using visual information.

[0012] On the one hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the method for using visual information.

[0013] On the one hand, a computer program product or computer program is provided, which includes a program code, which is stored in a computer-readable storage medium. The processor of the robot controller reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the robot controller executes the above-mentioned method of applying visual information. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] FIG1 is a schematic diagram of a pool cleaning robot in a pool provided by an embodiment of the present application;

[0016] FIG2 is a flow chart of a method for applying visual information provided by an embodiment of the present application;

[0017] FIG3 is a flow chart of another method for applying visual information provided by an embodiment of the present application;

[0018] FIG4 is a schematic diagram of a step provided in an embodiment of the present application;

[0019] FIG5 is a schematic diagram of another step provided in an embodiment of the present application;

[0020] FIG6 is a schematic diagram of a cleaning path provided in an embodiment of the present application;

[0021] FIG7 is a schematic diagram of another cleaning path provided in an embodiment of the present application;

[0022] FIG8 is a flowchart of another method for using visual information provided in an embodiment of the present application;

[0023] FIG9 is a schematic diagram of an outline and a circumscribed rectangle provided in an embodiment of the present application;

[0024] FIG10 is a schematic diagram of a first cleaning path provided in an embodiment of the present application;

[0025] FIG11 is a schematic diagram of a second cleaning path provided in an embodiment of the present application;

[0026] FIG12 is a schematic diagram of a third cleaning path provided in an embodiment of the present application;

[0027] FIG13 is a flowchart of another method for using visual information provided in an embodiment of the present application;

[0028] FIG14 is a schematic diagram of a flow chart illustrating feature matching during mapping by an underwater cleaning robot according to an embodiment of the present application;

[0029] FIG15 is a schematic diagram of a process for determining target acquisition location information provided by an embodiment of the present application;

[0030] FIG16 is a schematic diagram of an implementation process of position drift correction for an underwater cleaning robot provided in an embodiment of the present application;

[0031] FIG17 is a schematic diagram of a process for determining a target weight ratio according to an embodiment of the present application;

[0032] FIG18 is a flowchart of another method for using visual information provided in an embodiment of the present application;

[0033] FIG19 is a schematic diagram of an implementation process for obtaining real motion information of an underwater cleaning robot according to an embodiment of the present application;

[0034] FIG20 is a schematic diagram of an implementation process of trajectory association of an underwater cleaning robot provided in an embodiment of the present application;

[0035] FIG21 is a schematic diagram of a first intersection-over-union ratio between a detection area and a prediction area provided in an embodiment of the present application;

[0036] FIG22 is a schematic diagram of an implementation process of target detection by an underwater cleaning robot provided in an embodiment of the present application;

[0037] FIG23 is a schematic diagram of sampling of small and large targets by an underwater cleaning robot in the related art;

[0038] FIG24 is a flowchart of another method for using visual information provided in an embodiment of the present application;

[0039] FIG25 is a flowchart of another method for using visual information provided in an embodiment of the present application;

[0040] FIG26 is a schematic diagram of the structure of a visual information application device provided in an embodiment of the present application;

[0041] FIG27 is a schematic structural diagram of another visual information utilization device provided in an embodiment of the present application;

[0042] FIG28 is a schematic diagram of the structure of another visual information utilization device provided in an embodiment of the present application;

[0043] FIG29 is a schematic structural diagram of another visual information utilization device provided in an embodiment of the present application;

[0044] FIG30 is a schematic diagram of the structure of another visual information utilization device provided in an embodiment of the present application;

[0045] Figure 31 is a structural schematic diagram of a pool cleaning robot provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0047] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0048] First, the nouns involved in the embodiments of the present application are introduced.

[0049] Pool cleaning robot: A robot used to perform pool cleaning tasks. For example, when placed in a pool, the pool cleaning robot can clean the pool bottom. In some embodiments, the pool cleaning robot also has a wall-climbing function, capable of cleaning areas such as pool walls and steps.

[0050] Computer vision: The study of how machines can "see." Specifically, it involves using image acquisition devices and computers to replace the human eye in identifying, tracking, and measuring objects. Further image processing is performed, allowing the computer to create images more suitable for human observation or for transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting "information" from images or multidimensional data.

[0051] Object Detection: Object detection is the process of distinguishing one or more specific objects (or one or more types of objects) from other objects (or other types of objects). This includes both the identification of two very similar objects and the identification of one type of object from another type of object. In the embodiments of this application, object detection includes detecting the object to be cleaned and identifying obstacles.

[0052] Semantic segmentation: Semantic segmentation is to classify each pixel in the image, determine the category of each point, and then perform area division. In the embodiment of the present application, semantic segmentation is used to segment the bottom, wall and steps of the pool.

[0053] Visual inspection: Through machine vision systems, machines replace human eyes in measurement and judgment. It uses image processing and computer vision technologies to perform various processing on images captured by the robot to achieve the function of automatically detecting and identifying target objects.

[0054] Target environment map: A visual data structure that can be used to describe and represent various features and information in the target space environment. These features may include terrain, landforms, buildings, etc., while information may include location, shape, size, color, etc.

[0055] An inertial measurement unit (IMU) is a device that uses the principle of inertia to measure an object's motion. It typically consists of a gyroscope and an accelerometer. It senses a robot's acceleration and angular velocity in real time, thereby calculating its posture and position changes. With its high update rate and dynamic performance, IMUs are suitable for measuring high-speed motion and complex environments.

[0056] An encoder is a device that calculates the displacement of an object by measuring the number of pulse signals. It is typically used with rotating parts such as motors. It can provide accurate rotation angle and speed information for measuring mechanical motion.

[0057] A tachometer is a device used to measure water flow velocity, commonly used in underwater robotics, waterway monitoring, and other fields. It can calculate flow velocity by measuring the time or speed of water flowing through the sensor, with high accuracy and real-time performance.

[0058] Monocular visual odometry: Based on monocular vision technology, it estimates the motion of an object by continuously acquiring images and analyzing the changes between them. As a robot moves, the monocular camera continuously captures images at a high frequency. The visual odometry function estimates the robot's motion (translation and rotation) between two consecutive frames.

[0059] After introducing the nouns involved in the embodiments of the present application, the application scenarios of the embodiments of the present application are described below.

[0060] The technical solution provided by the embodiments of the present application can be applied to a scenario where a pool cleaning robot is controlled to clean the bottom of a pool. Referring to FIG1 , the pool cleaning robot 100 is capable of moving on the bottom 101 of the pool, thereby cleaning the bottom 101. As the pool cleaning robot 100 moves on the bottom 101, the water pump of the pool cleaning robot 100 is activated, drawing liquid from the pool into the robot's filter cartridge through a water inlet at the bottom of the pool cleaning robot 100. The filter cartridge filters the liquid, leaving dirt in the liquid within the filter cartridge. The filtered liquid is then discharged through the robot's drain port, thereby cleaning the bottom 101. In the embodiments of the present application, the pool cleaning robot 100 is also capable of cleaning the pool walls 102 and steps 103. The pool cleaning robot 100 utilizes the suction force generated by the water pump to adhere to the pool walls 102, thereby cleaning the pool walls 102 by moving along the walls 102. The steps 103 in the pool can be considered as a combination of multiple groups of smaller pool bottoms and pool walls, and the steps 103 can be cleaned in a corresponding manner. Of course, in addition to using the water pump and filter box to clean the pool area to be cleaned, the pool cleaning robot 100 can also use other types of cleaning units to clean the pool area to be cleaned, such as a roller brush to clean the pool area to be cleaned.

[0061] The technical solution provided in the embodiment of the present application is described below. Figure 2 is a flow chart of a method for applying visual information provided in the embodiment of the present application. Referring to Figure 2, taking the robot controller of a pool cleaning robot as an example, the method includes the following steps.

[0062] 201. The robot controller obtains visual information collected by the pool cleaning robot.

[0063] The visual information is information collected by the visual sensor of the pool cleaning robot. For example, images and point clouds are both visual information.

[0064] 202. The robot controller controls the pool cleaning robot based on the visual information.

[0065] Among them, controlling the pool cleaning machine refers to controlling the pool cleaning robot to clean garbage or objects to be cleaned.

[0066] 203. The robot controller processes the visual information according to a preset method.

[0067] Among them, the preset method is set by technical personnel according to actual conditions, and the embodiments of this application do not limit this.

[0068] In related art, when using a pool cleaning robot to clean a pool, it often traverses the pool along a pre-set path. This cleaning method is relatively simple and does not provide a good cleaning effect on the pool. The following will address this issue and provide a detailed description of the technical solution provided in the embodiments of this application, with reference to some examples. Referring to Figure 3, the method includes the following steps.

[0069] 301. When a pool cleaning robot performs a cleaning task in a pool, a robot controller obtains an image of the environment surrounding the pool cleaning robot.

[0070] Among them, the cleaning task refers to the task of cleaning the pool. The cleaning task is initiated regularly by the robot controller or manually by the user. For example, the user can set the execution cycle of the cleaning task for the pool cleaning robot, and the robot controller can initiate the cleaning task regularly according to the set execution cycle. Alternatively, the user uses the application terminal of the visual information to send a cleaning task execution instruction to the pool cleaning robot to control the pool cleaning robot to perform the cleaning task. The embodiment of the present application does not limit the method of initiating the cleaning task. The robot controller is built into the pool cleaning robot and is used to control the pool cleaning robot. When the pool cleaning robot is located in the pool, the pool cleaning robot can move and perform cleaning actions on the bottom, walls and steps of the pool. The pool cleaning robot includes an image acquisition device, which is used to capture the environmental image around the pool cleaning robot. The environmental image is used to reflect the environmental conditions around the pool cleaning robot. In some embodiments, the image acquisition device is an RGB (Red Green Blue) camera. The direction indicated by the central axis of the environmental image is the forward direction of the pool cleaning robot. That is, when a target is on the central axis of the environmental image, the pool cleaning robot can reach the location of the target by moving forward.

[0071] In some embodiments, when the pool cleaning robot performs a cleaning task, the robot controller obtains an image of the environment surrounding the pool cleaning robot through an image acquisition device.

[0072] The number of image acquisition devices may be multiple or one, and this is not limited in the present embodiment. In the case of multiple image acquisition devices, the environmental image is obtained by stitching together the environmental images captured by multiple image acquisition devices. It should be noted that when the pool cleaning robot performs the cleaning task, the robot controller will continuously acquire environmental images through the image acquisition devices.

[0073] For example, if there is only one image capture device installed directly in front of the pool cleaning robot, it can capture an image of the environment directly in front of the pool cleaning robot. The central axis of the environment image indicates the direction directly in front of the pool cleaning robot. If any object is to the left of the central axis of the environment image, it means that the object is to the left and in front of the pool cleaning robot.

[0074] For example, a plurality of image acquisition devices may be provided. A target image acquisition device is installed directly in front of the pool cleaning robot, and the remaining image acquisition devices are installed at equal intervals on either side of the target image acquisition device. The robot controller acquires an environmental image from the plurality of image acquisition devices and, with the environmental image captured by the target image acquisition device as the center, splices the environmental images captured by the remaining image acquisition devices onto the environmental image captured by the target image acquisition device based on the relative positional relationship between the remaining image acquisition devices and the target image acquisition device to obtain the environmental image.

[0075] 302. The robot controller extracts features of the environment image using the image recognition model to obtain environment image features of the environment image.

[0076] The image recognition model is a multi-task model that can perform both object detection and semantic segmentation on the input environment image. Feature extraction is used to map the environment image to a higher dimension for subsequent processing.

[0077] In some embodiments, the robot controller inputs the environment image into the image recognition model, and performs multiple convolutions on the environment image through the image recognition model to obtain image features of the environment image.

[0078] In this embodiment, by performing multiple convolutions on the environment image, the features of the environment image at different depths can be obtained, and the image features are obtained by using the features at different depths. The image features have a stronger expressive ability.

[0079] In some embodiments, the robot controller inputs the environment image into the image recognition model, and the image recognition model continuously performs multiple convolutions on the environment image to obtain first convolution image features corresponding to each convolution, where the first convolution image features corresponding to different convolutions have different sizes. The robot controller, through the image recognition model, fuses the first convolution image features corresponding to each convolution to obtain image features of the environment image.

[0080] The following describes a method for fusing the first convolution image features corresponding to each convolution in the above example to obtain the image features.

[0081] In some embodiments, the robot controller deconvolves each first convolution image feature using the image recognition model to obtain a second convolution image feature corresponding to each first convolution image feature. The robot controller fuses each first convolution image feature with the corresponding second convolution image feature using the image recognition model to obtain a plurality of third convolution image features. The robot controller fuses the plurality of third convolution image features using the image recognition model to obtain image features of the environment image.

[0082] Another implementation of the above step 302 is described below.

[0083] In some embodiments, the robot controller inputs the environment image into the image recognition model, and performs multiple full connections on the environment image through the image recognition model to obtain image features of the environment image.

[0084] In this implementation, by performing multiple full connections on the environment image, deep image features can be extracted, and the image features have strong expressive power.

[0085] In some embodiments, the robot controller inputs the environment image into the image recognition model, and performs a first full-connection on the environment image through the image recognition model to obtain a first fully-connected feature of the environment image. The robot controller uses the image recognition model to pool and normalize the first fully-connected feature to obtain a first reference feature. The robot controller uses the image recognition model to perform a second full-connection on the first reference feature to obtain a second fully-connected feature of the environment image. The robot controller uses the image recognition model to pool and normalize the second fully-connected feature to obtain a second reference feature, and so on, until the reference feature obtained from the last full-connection is determined as the image feature of the environment image.

[0086] Another implementation of the above step 302 is described below.

[0087] In some embodiments, the robot controller inputs the environment image into the image recognition model, and the image recognition model encodes the environment image based on the attention mechanism to obtain image features of the environment image.

[0088] In this implementation, the attention mechanism is used to focus on important parts of the environment image, and the image features obtained after encoding can more accurately represent the environment image.

[0089] In some embodiments, the robot controller inputs the environmental image into the image recognition model, and divides the environmental image into multiple image blocks through the image recognition model. The robot controller embeds and encodes each image block through the image recognition model to obtain the embedded features of each image block. The robot controller linearly transforms the embedded features of each image block through the image recognition model to obtain the query matrix, key matrix and value matrix of each image block. The robot controller obtains the attention weight of each image block to other image blocks based on the query matrix and key matrix of every two image blocks in the multiple image blocks through the image recognition model. The robot controller uses the attention weight to weightedly fuse the value matrices of each image block through the image recognition model to obtain the image features of the environmental image.

[0090] It should be noted that the robot controller can extract features of the environment image through any of the above methods to obtain image features, and the embodiments of the present application are not limited to this.

[0091] 303. The robot controller performs border detection on the environmental image based on the environmental image features through the image recognition model to obtain the first recognition frame and the second recognition frame. The first recognition frame is used to indicate the position of the target to be cleaned on the environmental image, and the second recognition frame is used to indicate the position of the obstacle on the environmental image.

[0092] Among them, border detection belongs to target detection, and the result of border detection is the first identification frame and the second identification frame. The first identification frame surrounds the target to be cleaned, and the second identification frame surrounds the obstacle. The target to be cleaned includes various types, such as stains, cigarette butts, debris, stones, sand and broken branches, etc. The type of the target to be cleaned is set and adjusted by the technician according to the actual situation, and the embodiment of the present application does not limit this. Obstacles also include various types, such as pillars, pipes, lighting equipment and landscape stones protruding from the plane in the pool, etc. The type of obstacles is set and adjusted by the technician according to the actual situation, and the embodiment of the present application does not limit this. Ideally, the first identification frame completely surrounds the target to be cleaned, and the second identification frame completely surrounds the obstacle. However, when the first identification frame does not completely surround the target to be cleaned, or when the second identification frame does not completely surround the obstacle, the technical solution provided by the embodiment of the present application can also be executed. The number of the first identification frame and the second identification frame can be one or more. For ease of understanding, the following is explained as an example of the number of the first identification frame and the second identification frame being one.

[0093] In some embodiments, the robot controller uses the image recognition model to control a first candidate recognition frame and a second candidate recognition frame to slide across the environment image, where the first candidate recognition frame is used to identify the target to be cleaned, and the second candidate recognition frame is used to identify the obstacle. The robot controller uses the image recognition model to determine the first recognition frame and the second recognition frame based on first sub-image features corresponding to multiple first image regions in the environment image feature and second sub-image features corresponding to multiple second image regions in the environment image feature.

[0094] The first image area is the image area covered by the first candidate identification frame on the environmental image, and the second image area is the image area covered by the second candidate identification frame on the environmental image. In some embodiments, there are multiple first candidate identification frames and multiple second candidate identification frames, each of which has a different size. This allows for border detection of different granularities in the environmental image, thereby improving the accuracy of border detection.

[0095] In this embodiment, the first candidate identification frame and the second candidate identification frame are used to slide on the environmental image, and the first candidate identification frame and the second candidate identification frame are used to cover different image areas on the environmental image. The image areas are classified to determine the first identification frame and the second identification frame, with high accuracy.

[0096] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0097] In the first part, the robot controller controls the first candidate recognition frame and the second candidate recognition frame to slide on the environment image through the image recognition model.

[0098] In some embodiments, when the number of the first candidate identification frame and the second candidate identification frame is one, the robot controller controls the first candidate identification frame and the second candidate identification frame to slide on the environmental image with a preset step size through the image recognition model, thereby covering different image areas of the environmental image, and performing target detection on the environmental image with the image area as the granularity.

[0099] Among them, the preset step size is set by technical personnel according to actual conditions, and the embodiments of the present application do not limit this.

[0100] In some embodiments, when there are multiple first candidate identification frames and second candidate identification frames, the robot controller uses the image recognition model to control the corresponding first candidate identification frames and second candidate identification frames to slide on the environmental image with preset step sizes corresponding to each first candidate identification frame and each second candidate identification frame. Different first candidate identification frames have different sizes, and different second candidate identification frames have different sizes, thereby covering different image areas of the environmental image, and performing target detection on the environmental image with image areas of different sizes as the granularity.

[0101] In the second part, the robot controller determines the first recognition frame and the second recognition frame through the image recognition model based on the first sub-image features corresponding to multiple first image areas in the environmental image features and the second sub-image features corresponding to multiple second image areas in the environmental image features.

[0102] In order to explain the above steps more clearly, the methods for determining the first recognition frame and the second recognition frame will be described below respectively.

[0103] In some embodiments, for any first image region among the multiple first image regions, the robot controller uses the image recognition model to fully connect and normalize the first sub-image features corresponding to the first image region to determine whether the first image region contains the object to be cleaned. If the first image region contains the object to be cleaned, the robot controller determines the bounding box of the first image region as a first reference recognition frame. The robot controller uses the image recognition model to determine the first recognition frame based on the multiple first reference recognition frames in the environment image.

[0104] In some embodiments, for any first image region among the multiple first image regions, the robot controller uses the image recognition model to fully connect and normalize the first sub-image features corresponding to the first image region to obtain a first region classification value for the first image region. If the first region classification value is greater than or equal to the first classification value threshold, it is determined that the first image region contains the target to be cleaned. If the first region classification value is less than the first classification value threshold, it is determined that the first image region does not contain the target to be cleaned. If the first image region contains the target to be cleaned, the robot controller determines the border of the first image region as the first reference recognition frame. If the first image region does not contain the target to be cleaned, the robot controller discards the first image region. The robot controller uses the image recognition model to perform non-maximum suppression on multiple first reference recognition frames on the environment image to remove redundant first reference recognition frames and obtain the first recognition frame, wherein the method of performing non-maximum suppression on multiple first reference recognition frames described in the above example includes non-maximum suppression (NMS), and extended methods of NMS, such as Soft-NMS, Softer-NMS and Fast-NMS, etc., which are not limited in this embodiment of the present application. The first classification value threshold is set by technical personnel according to actual conditions, which is not limited in this embodiment of the present application. The first classification value threshold is used to identify the first recognition frame containing the target to be cleaned.

[0105] In some embodiments, for any second image region among the multiple second image regions, the robot controller uses the image recognition model to fully connect and normalize the second sub-image features corresponding to the second image region to determine whether the second image region contains an obstacle. If the second image region contains an obstacle, the robot controller determines the bounding box of the second image region as a second reference recognition frame. The robot controller uses the image recognition model to determine the second recognition frame based on the multiple second reference recognition frames in the environment image.

[0106] In some embodiments, for any second image region among the multiple second image regions, the robot controller, using the image recognition model, performs full connectivity and normalization on the second sub-image features corresponding to the second image region to obtain a second region classification value for the second image region. If the second region classification value is greater than or equal to a second classification value threshold, the second image region is determined to contain an obstacle. If the second region classification value is less than the second classification value threshold, the second image region is determined to contain no obstacle. If the second image region contains an obstacle, the robot controller determines the bounding box of the second image region as a second reference recognition frame. If the second image region does not contain an obstacle, the robot controller discards the second image region. The robot controller, using the image recognition model, performs non-maximum suppression on multiple second reference recognition frames in the environment image to remove redundant second reference recognition frames, thereby obtaining the second recognition frame. The non-maximum threshold method is consistent with the previous embodiment and its implementation is not further described. The second classification value threshold is set by a skilled person based on actual circumstances and is not limited in this embodiment. The second classification value threshold is used to identify second recognition frames containing obstacles.

[0107] It should be noted that the robot controller can execute the above step 303 and the following step 304 in parallel, or execute the following steps 303 and 304 in sequence, and this embodiment of the application does not limit this. In the subsequent description process, the robot controller executes the above step 303 and the following step 304 in parallel as an example.

[0108] 304. The robot controller classifies a plurality of pixels in the environment image based on the environment image features using the image recognition model to obtain a location type of the pool cleaning robot.

[0109] Among them, the location types include the bottom of the pool, the wall of the pool and the steps. The process of classifying multiple pixel points is also the process of semantic segmentation. In some embodiments, the steps can also be subdivided into step planes and step facades, thereby improving the precision of semantic segmentation. In some embodiments, the environmental image features are obtained by the robot controller performing multiple downsamplings on the environmental image. For example, the robot controller continuously performs multiple convolutions on the environmental image to obtain the environmental image features of the environmental image. Among them, the number of times the environmental image is convolved is set by the technician according to the actual situation, and the embodiments of the present application do not limit this. The convolution target during the first convolution is the environmental image, and a first convolution feature map is obtained, and the size of the first convolution feature map is smaller than the environmental image. The convolution target during the second convolution is the first convolution feature map, and a second convolution feature map is obtained, and the size of the second convolution feature map is smaller than the first convolution feature map, and so on. The convolution feature map obtained by the last convolution is also the environmental image feature.

[0110] In some embodiments, the robot controller performs multiple upsampling of the environmental image features to obtain an upsampled feature map of the environmental image. The upsampled feature map has the same size as the environmental image and includes multiple channels corresponding to multiple candidate location types. The pixel value of each channel represents the confidence that the corresponding pixel in the environmental image is in the candidate location type corresponding to the channel. Based on the upsampled feature map, the robot controller determines the location type corresponding to each pixel in the environmental image. Based on the location types corresponding to the multiple pixels, the robot controller determines the location type of the pool cleaning robot.

[0111] Among them, multiple candidate location types include pool wall, pool bottom, step plane and step facade. In addition, in other possible implementations, the multiple candidate location types can also include other location types, such as raised planes in the pool, etc., which is not limited in this embodiment of the present application.

[0112] In this embodiment, by classifying the pixels in the environment image, the position type corresponding to the pixel is determined. The position type of the pool cleaning robot is determined based on the position type corresponding to the pixel, and the accuracy of the position type is high.

[0113] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0114] In the first part, the robot controller upsamples the environment image features multiple times to obtain an upsampled feature map of the environment image.

[0115] In some embodiments, the robot controller continuously upsamples the environmental image features multiple times to obtain an upsampled feature map of the environmental image.

[0116] Among them, the number of times the environmental image is upsampled is set by the technician according to the actual situation, and the embodiments of the present application do not limit this. The upsampling target during the first upsampling is the environmental image feature, and a first upsampling feature map is obtained, and the size of the first upsampling feature map is larger than the environmental image feature. The upsampling target during the second upsampling is the first upsampling feature map, and a second upsampling feature map is obtained, and the size of the second upsampling feature map is larger than the first upsampling feature map, and so on. The upsampling feature map obtained by the last upsampling is also the environmental image feature, and the size of the environmental image feature is the same as that of the environmental image. In some embodiments, the upsampling method includes deconvolution, interpolation and other upsampling methods, which are not limited in the embodiments of the present application.

[0117] In the second part, the robot controller determines the position type corresponding to each pixel point in the environment image based on the upsampled feature map.

[0118] In some embodiments, the robot controller determines the position type corresponding to each pixel point in the environment image based on the pixel values ​​of the pixel points in multiple channels of the upsampled feature map.

[0119] In some embodiments, for any pixel among the plurality of pixels in the environment image, the robot controller determines the pixel value of the pixel in each channel of the upsampled feature map. The robot controller determines the channel corresponding to the maximum pixel value as the target channel, and the candidate location type corresponding to the target channel is the location type corresponding to the pixel. The maximum pixel value also indicates the maximum confidence. For example, if the candidate location type corresponding to the target channel is the pool bottom, then the location type corresponding to the pixel is the pool bottom.

[0120] Part three: The robot controller determines the position type of the pool cleaning robot based on the position types corresponding to the multiple pixel points.

[0121] In some embodiments, the robot controller determines the position type of the pool cleaning robot ahead in the direction of travel based on the position types corresponding to the multiple pixel points. The robot controller determines the position type of the pool cleaning robot's location based on the position type of the pool cleaning robot ahead in the direction of travel.

[0122] In this embodiment, the position type in front of the pool cleaning robot's travel direction is determined based on the position types of multiple pixels in the environmental image, thereby determining the position type of the pool cleaning robot's location, which is more efficient.

[0123] In some embodiments, the robot controller clusters the multiple pixels on the environmental image based on the location types corresponding to the multiple pixels, obtaining multiple location indication areas on the environmental image. Different location indication areas correspond to different location types, and the dividing line between adjacent different location indication areas serves as the boundary between different location types. The robot controller determines the location type ahead of the pool cleaning robot in its travel direction based on the locations of the multiple location indication areas on the environmental image. The robot controller determines the location type of the pool cleaning robot based on the location type ahead of the pool cleaning robot in its travel direction.

[0124] In order to explain the above example more clearly, the method in which the robot controller determines the position type ahead of the traveling direction of the pool cleaning robot based on the positions of the multiple position indication areas on the environment image in the above example will be explained below.

[0125] In some embodiments, the robot controller determines the position type corresponding to the position indication area located at the bottom of the environment image among the multiple position indication areas as the position type in front of the travel direction of the pool cleaning robot.

[0126] For example, when the position type corresponding to the position indication area at the bottom of the environmental image is the pool bottom, the robot controller determines the position type in front of the pool cleaning robot's travel direction as the pool bottom; when the position type corresponding to the position indication area at the bottom of the environmental image is the pool wall, the robot controller determines the position type in front of the pool cleaning robot's travel direction as the pool wall; when the position type corresponding to the position indication area at the bottom of the environmental image is a step, the robot controller determines the position type in front of the pool cleaning robot's travel direction as a step; when the position type corresponding to the position indication area at the bottom of the environmental image includes the dividing line between the pool bottom and the pool bottom, the robot controller determines the position type in front of the pool cleaning robot's travel direction as the dividing line between the pool bottom and the pool bottom.

[0127] 305. The robot controller determines a cleaning mode corresponding to the location type of the pool cleaning robot.

[0128] In some embodiments, when the pool cleaning robot is located at the pool wall, the robot controller determines the cleaning mode of the pool cleaning robot to be the roller brush cleaning mode.

[0129] The objects to be cleaned on the pool wall are typically microorganisms, algae, and scale attached to the pool wall. When cleaning the pool wall, the pressure between the roller brush and the pool wall is increased to enhance the friction between the roller brush and the pool wall, thereby enhancing the cleaning effect on the pool wall. The roller brush cleaning mode described in the above embodiment refers to a cleaning mode that switches to the roller brush and increases the pressure between the roller brush and the pool wall. Of course, based on this, the robot controller can also control other cleaning components of the pool cleaning robot to clean the pool wall, and this embodiment of the application is not limited to this.

[0130] In some embodiments, when the pool cleaning robot is located at the bottom of the pool, the robot controller determines the cleaning mode of the pool cleaning robot to be the water pumping mode.

[0131] Among them, the targets to be cleaned on the bottom of the pool are usually pebbles, fallen leaves, sand and other sediments. During the cleaning process, a larger water flow is required to allow as many targets to be cleaned as possible to enter the filter box through the water flow and remain in the filter box, thereby achieving efficient cleaning of the pool bottom. The water pump extraction mode in the above-mentioned embodiment refers to increasing the power of the water pump to increase the water flow, thereby sucking the targets to be cleaned into the filter box. Accordingly, on this basis, the robot controller can also control other cleaning components of the pool cleaning robot to clean the pool wall, and the embodiments of the present application are not limited to this.

[0132] In some embodiments, when the pool cleaning robot is located at the boundary between the bottom and the wall of the pool, the robot controller determines the cleaning mode of the pool cleaning robot to be the edge cleaning mode.

[0133] Among them, the edge cleaning mode refers to cleaning close to the dividing line. This is because sand and gravel are easily accumulated on the dividing line between the pool bottom and the pool wall. The edge cleaning mode can improve the cleaning effect of the dividing line between the pool bottom and the pool wall, thereby improving the cleaning effect of the pool.

[0134] In some embodiments, when the pool cleaning robot is located on the steps of the pool, the robot controller determines the cleaning mode of the pool cleaning robot to be a basic cleaning mode or a deep cleaning mode. In the basic cleaning mode, the step plane is cleaned but the step facade is not cleaned. In the deep cleaning mode, the step facade and the step plane are cleaned.

[0135] As shown in Figure 4, in basic cleaning mode, the pool cleaning robot only cleans the flat surface of the steps, not the vertical surface. In deep cleaning mode, the pool cleaning robot cleans both the flat surface and the vertical surface. Furthermore, the boundary between the flat surface and the vertical surface of the steps is similar to the boundary between the pool bottom and the pool wall. The pool cleaning robot can also use edge cleaning mode to clean the boundary between the flat surface and the vertical surface of the steps, thereby improving the cleaning effect on the steps.

[0136] 306. The robot controller performs path planning for the pool cleaning robot based on the environment image, the first recognition frame, the second recognition frame, and the cleaning mode to obtain a target movement trajectory.

[0137] Here, cleaning the target to be cleaned includes moving to the location of the target to be cleaned and performing a cleaning action.

[0138] In some embodiments, the robot controller determines a preset cleaning trajectory corresponding to the cleaning mode. The robot controller determines a reference movement trajectory based on the positions of the first identification frame and the second identification frame in the environment image, the obstacle type of the obstacle in the second identification frame, and the size of the second identification frame. The robot controller combines the preset cleaning trajectory and the reference movement trajectory to obtain the target movement trajectory.

[0139] Among them, the preset cleaning trajectory corresponding to the cleaning mode is also called the basic cleaning path corresponding to the cleaning mode, which is the cleaning path of the pool cleaning robot in this cleaning mode when there are no obstacles and objects to be cleaned.

[0140] In this implementation, a preset cleaning trajectory corresponding to a cleaning mode is determined. A reference trajectory is determined based on the positions of the first and second recognition frames in the environmental image, the types of obstacles in the second recognition frame, and the size of the second recognition frame. A target trajectory is then determined based on the preset cleaning trajectory and the reference trajectory, enabling intelligent path planning for the pool cleaning robot. Using the recognition frames for path planning, rather than further segmenting the frames, significantly reduces computational complexity while ensuring accurate paths.

[0141] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0142] In the first part, the robot controller determines the preset cleaning trajectory corresponding to the cleaning mode.

[0143] In some embodiments, the robot controller performs a query based on the cleaning mode to obtain a preset cleaning trajectory corresponding to the cleaning mode.

[0144] Among them, the correspondence between the cleaning mode and the preset cleaning trajectory is set by the technician according to the actual situation, and the embodiments of the present application do not limit this. In some embodiments, the preset cleaning trajectory corresponding to the water pump extraction mode is the "bow"-shaped path shown in Figure 6. The "bow"-shaped path includes multiple horizontal sub-paths and multiple vertical sub-paths. Any horizontal sub-path in the multiple horizontal sub-paths is perpendicular to any vertical sub-path in the multiple vertical sub-paths. Any two horizontal sub-paths in the multiple horizontal sub-paths are parallel to each other, and any two vertical sub-paths in the multiple vertical sub-paths are parallel to each other. In order to facilitate the distinction between horizontal sub-paths and vertical sub-paths, the sub-path with a longer length is determined as a vertical sub-path, and the sub-path with a shorter length is determined as a horizontal sub-path. In some embodiments, referring to FIG7 , the length of the horizontal sub-path is related to the visual cone of the image acquisition device of the pool cleaning robot, that is, when the pool cleaning robot moves on the vertical sub-path, the visual cone of the image acquisition device can cover the central axis between the vertical sub-path and the two adjacent vertical sub-paths. The horizontal distance covered by the visual cone of the image acquisition device is greater than the length of the horizontal sub-path to achieve full coverage of the pool. This can reduce the density of the cleaning path of the pool cleaning robot and improve cleaning efficiency.

[0145] In some embodiments, when the pool is cleaned without the image acquisition device, the preset cleaning trajectory can also be a "bow"-shaped path. Compared to the "bow"-shaped path in Figure 6 above, the length of the horizontal sub-path is associated with the width of the pool cleaning robot. In some embodiments, the length of the horizontal sub-path is less than the width of the pool cleaning robot, that is, the distance between two adjacent vertical sub-paths is less than the width of the pool cleaning robot, thereby ensuring that when the pool cleaning robot cleans the pool according to the preset cleaning trajectory, there is an area of ​​repeated coverage when walking on any two adjacent vertical sub-paths to ensure the cleaning effect. Of course, when using the image acquisition device to clean the pool, the preset cleaning trajectory can be switched to the form shown in Figure 6.

[0146] In the second part, the robot controller determines a reference movement trajectory based on the positions of the first recognition frame and the second recognition frame in the environment image, the obstacle type of the obstacle in the second recognition frame, and the size of the second recognition frame.

[0147] In some embodiments, the robot controller determines whether the pool cleaning robot needs to avoid the obstacle based on the obstacle type, the position of the second identification frame in the environmental image, and the size of the second identification frame. If the pool cleaning robot needs to avoid the obstacle, the robot controller determines a first movement direction and a first movement distance of the pool cleaning robot based on the relative positional relationship between the first and second identification frames and the center point of the environmental image. The robot controller generates the reference movement trajectory based on the first movement direction and the first movement distance of the pool cleaning robot.

[0148] Among them, the obstacle types include protruding pillars, pipes, lighting equipment and landscape stones in the pool, and the obstacle types are obtained by classifying the second identification frame. The position of the second identification frame in the environmental image can reflect the relative position relationship between the obstacle and the pool cleaning robot. The position of the second identification frame on the environmental image and the size of the second identification frame can reflect the size of the obstacle and the distance between the obstacle and the pool cleaning robot. The relative position relationship between the first identification frame and the center point of the environmental image includes the distance and angle between the first identification frame and the center point of the environmental image. Correspondingly, the relative position relationship between the second identification frame and the center point of the environmental image includes the distance and angle between the second identification frame and the center point of the environmental image.

[0149] Because objects appear larger when they are closer, or smaller when they are farther away, the size of the second recognition frame alone cannot be used to determine the distance between the pool cleaning robot and the obstacle, nor the size of the obstacle. The second recognition frame's position within the surrounding image is also required for comprehensive judgment. For example, the closer the vertical coordinate of the second recognition frame's center point is to the bottom edge of the surrounding image, the closer it is to the pool cleaning robot.

[0150] In order to explain the above embodiment more clearly, the above embodiment will be further described in several parts below.

[0151] A. The robot controller determines whether the pool cleaning robot needs to avoid the obstacle based on the obstacle type, the position of the second identification frame in the environment image, and the size of the second identification frame.

[0152] In some embodiments, the robot controller determines the size of the obstacle based on the position of the second identification frame in the environment image and the size of the second identification frame. The robot controller determines whether the pool cleaning robot needs to avoid the obstacle based on the obstacle type and size of the obstacle.

[0153] In some embodiments, the robot controller determines the distance between the obstacle and the pool cleaning robot based on the position of the second identification frame in the environmental image. The robot controller determines the size of the obstacle based on the distance between the obstacle and the pool cleaning robot and the size of the second identification frame. If the obstacle type is a preset type and the size of the obstacle is less than or equal to the preset size, it is determined that the pool cleaning robot does not need to avoid the obstacle; if the obstacle type is not the preset type or the size of the obstacle is greater than the preset size, it is determined that the pool cleaning robot needs to avoid the obstacle. The preset type and the preset size are set by technicians according to actual conditions and are not limited in this embodiment of the application.

[0154] B. When the pool cleaning robot needs to avoid the obstacle, the robot controller determines the first moving direction and the first moving distance of the pool cleaning robot based on the relative position relationship between the first identification frame, the second identification frame and the center point of the environment image.

[0155] In some embodiments, when the pool cleaning robot needs to avoid the obstacle, the robot controller determines the relative positional relationship between the object to be cleaned, the obstacle, and the pool cleaning robot based on the relative positional relationship between the first identification frame, the second identification frame, and the center point of the environment image. The robot controller performs path planning based on the relative positional relationship between the object to be cleaned, the obstacle, and the pool cleaning robot, and determines a first moving direction and a first moving distance for the pool cleaning robot, wherein the purpose of path planning is to avoid the obstacle and reach the location of the object to be cleaned.

[0156] C. The robot controller generates the reference movement trajectory based on the first movement direction and the first movement distance of the pool cleaning robot.

[0157] In some embodiments, the robot controller combines the first moving direction and the first moving distance of the pool cleaning robot in a sequential order to obtain the reference moving trajectory.

[0158] The above description is based on an example in which the pool cleaning robot needs to avoid the obstacle. The following description will describe a case in which the pool cleaning robot does not need to avoid the obstacle.

[0159] In some embodiments, when the pool cleaning robot does not need to avoid the obstacle, the robot controller determines a second movement direction and a second movement distance of the pool cleaning robot based on the relative positional relationship between the first recognition frame and the center point of the environment image. The robot controller generates the reference movement trajectory based on the second movement direction and the second movement distance of the pool cleaning robot.

[0160] In which case, when the pool cleaning robot does not need to avoid the obstacle, it does not need to consider the obstacle when generating the reference movement trajectory, and only needs to consider the target to be cleaned.

[0161] In the third part, the robot controller combines the preset cleaning trajectory and the reference movement trajectory to obtain the target movement trajectory.

[0162] In some embodiments, the robot controller splices the preset cleaning trajectory and the reference movement trajectory to obtain the target movement trajectory.

[0163] Another implementation of the above step 306 is described below.

[0164] In some embodiments, the robot controller determines the pool cleaning robot's obstacle climbing level, which is positively correlated with the pool cleaning robot's obstacle climbing capability. If the pool cleaning robot's obstacle climbing level is less than or equal to a preset level, the robot controller segments the obstacle in the second recognition frame to obtain a contour of the obstacle in the environmental image. The robot controller then performs path planning for the pool cleaning robot based on the environmental image, the first recognition frame, the contour of the obstacle, and the cleaning mode to obtain a target movement trajectory.

[0165] The obstacle climbing level and the preset level of the pool cleaning robot are set by technicians according to actual conditions, and are not limited in the embodiments of the present application. The obstacle climbing level of the pool cleaning robot is related to the hardware structure of the pool cleaning robot.

[0166] In this embodiment, when the pool cleaning robot has a poor obstacle climbing capability, the obstacles in the second recognition frame are further segmented to improve the accuracy of the obstacles and the accuracy of the path planning.

[0167] Based on the above embodiment, when the obstacle climbing level of the pool cleaning robot is greater than the preset level, the robot controller plans the path of the pool cleaning robot based on the environmental image, the first identification frame, the second identification frame and the cleaning mode to obtain the target movement trajectory.

[0168] That is, there is no need to further identify the second identification frame, and obstacle avoidance can be performed with the second identification frame as the granularity. In other words, the second identification frame can be avoided during the path planning process. For the specific implementation process, please refer to the description in the previous implementation method, which will not be repeated here.

[0169] 307. The robot controller controls the pool cleaning robot to move along the target movement trajectory in the cleaning mode to clean the target to be cleaned in the pool.

[0170] In some embodiments, the robot controller controls the pool cleaning robot to activate the cleaning mode, wherein the robot moves along a preset cleaning trajectory in the target movement trajectory. After reaching a designated location, the robot switches to a reference movement trajectory to clean the target. When the target is cleaned, the robot returns to the preset cleaning trajectory.

[0171] It should be noted that, during the movement of the pool cleaning robot, the above steps 301-307 will be repeatedly executed to achieve dynamic path planning and improve the cleaning effect.

[0172] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0173] According to the technical solution provided by the embodiment of the present application, when a pool cleaning robot is performing a cleaning task in a pool, an environmental image of the pool cleaning robot is obtained, and the environmental image can reflect the environmental conditions surrounding the pool cleaning robot. The environmental image is input into an image recognition model, and the image recognition model performs target detection and semantic segmentation on the environmental image to obtain a first recognition frame, a second recognition frame, and the position type of the pool cleaning robot's location on the environmental image. The first recognition frame is used to indicate the location of the target to be cleaned on the environmental image, and the second recognition frame is used to indicate the location of the obstacle on the environmental image. The image recognition model is a multi-task model. Based on the environmental image, the first recognition frame, the second recognition frame, and the position type of the pool cleaning robot's location, the pool cleaning robot is controlled to clean the target to be cleaned. That is, the target to be cleaned, the obstacle, and the position type identified by the image recognition model are used for overall planning, so that the pool cleaning robot can intelligently clean the target to be cleaned in combination with the environment in which it is located, thereby improving the cleaning effect of the pool cleaning robot on the pool.

[0174] In response to the above-mentioned problem of poor cleaning effect on the pool, another implementation method provided by the embodiment of the present application will be described in more detail with reference to some examples. Referring to FIG8 , the method includes the following steps.

[0175] 801. The robot controller determines the garbage types of multiple candidate garbage around the pool cleaning robot.

[0176] The robot controller is built into the pool cleaning robot and is used to control the pool cleaning robot. When the pool cleaning robot is located in a pool, the pool cleaning robot can move and perform cleaning actions on the pool bottom, pool walls, and steps. The pool cleaning robot includes a visual sensor that is used to capture an image of the environment surrounding the pool cleaning robot and / or a point cloud of the environment surrounding the pool cleaning robot. The image and point cloud are used to reflect the environmental conditions surrounding the pool cleaning robot. Candidate garbage is garbage located around the pool cleaning robot and identified by the robot controller. Garbage refers to objects that can be cleaned, such as pebbles, cigarette butts, paper balls, branches, leaves, and mud balls, etc., although this embodiment of the present application is not limited to this. The garbage type of the candidate garbage is used to indicate the shape of the candidate garbage. That is, candidate garbage of different shapes will be classified into different garbage types. In this embodiment of the present application, the shape of the garbage refers to the shape of the garbage when viewed from above, that is, the shape of the garbage when viewed from above. In some embodiments, garbage types include point types, line segment types, and surface types (also referred to as rectangle types).

[0177] In some embodiments, the robot controller determines an outline of each candidate trash item. Based on the outline of each candidate trash item, the robot controller determines a bounding rectangle of each candidate trash item. Based on the size information of the bounding rectangle of each candidate trash item, the robot controller determines the trash type of each candidate trash item.

[0178] The outline of a candidate trash item can be considered a representation of its shape. It closely resembles the actual shape of the candidate trash item, and is therefore typically an irregular polygon. The bounding rectangle of a candidate trash item is determined based on its outline, so it also reflects its shape. The size of the bounding rectangle is used to reflect the shape of the candidate trash item. Furthermore, compared to the outline, the bounding rectangle is more regular, which aids in classification.

[0179] Through the technical solution provided in the embodiments of the present application, the circumscribed rectangle corresponding to the outline of each candidate garbage is determined, and the garbage type of each candidate garbage is determined using the size information of the circumscribed rectangle corresponding to each candidate garbage. The difficulty of classifying each candidate garbage is low and the efficiency is high.

[0180] In order to explain the above embodiment more clearly, the above embodiment will be explained in several parts below.

[0181] In the first part, the robot controller determines the outline of each candidate garbage.

[0182] In some embodiments, the robot controller acquires an image of the environment surrounding the pool cleaning robot. The robot controller performs a perspective transformation on the image to obtain a BEV image corresponding to the image. The robot controller performs object detection on the BEV image to obtain the outlines of each candidate garbage item.

[0183] Among them, the environmental image can reflect the environmental conditions of the environment in which the pool cleaning robot is located. The environmental image is an image taken from the perspective of the pool cleaning robot. The perspective transformation of the environmental image is to convert the environmental image from the perspective of the pool cleaning robot to a projection perspective, that is, BEV (Bird's Eye View), so as to facilitate the determination of the outline of each candidate.

[0184] In order to explain the above embodiment more clearly, the following will further describe the above embodiment in several parts.

[0185] A. The robot controller obtains an image of the environment surrounding the pool cleaning robot.

[0186] In some embodiments, the robot controller controls the pool cleaning robot to rotate and obtains multiple initial environment images during the rotation process. The robot controller splices the multiple initial environment images to obtain the environment image.

[0187] Among them, one initial environment image corresponds to one rotation angle, and stitching multiple initial environment images means stitching the multiple initial environment images in the order of acquisition time to obtain an environment image reflecting the global environment around the pool cleaning robot.

[0188] In this embodiment, the pool cleaning robot is controlled to rotate and multiple initial environment images are obtained during the rotation process. The multiple initial environment images are spliced ​​together to obtain an environment image, which can more comprehensively reflect the environmental conditions around the pool cleaning robot.

[0189] In some embodiments, the robot controller controls the pool cleaning robot to rotate. During the rotation of the pool cleaning robot, the visual sensor captures images at predetermined rotation angles to obtain multiple initial environment images. The robot controller stitches the multiple initial environment images in order of acquisition time to obtain the environment image.

[0190] B. The robot controller transforms the perspective of the environment image to obtain the BEV image corresponding to the environment image.

[0191] In one possible implementation, the robot controller acquires multiple reference environment images whose generation time is adjacent to the environment image. The robot controller performs a perspective transformation on the environment image based on the multiple reference environment images and the environment image to obtain a BEV image corresponding to the environment image.

[0192] The multiple reference environment images include reference environment images whose generation time is between the environment images, and reference environment images whose generation time is after the environment image.

[0193] In this embodiment, the environmental image and multiple reference environmental images of the environmental image are used to perform perspective transformation on the environmental image to obtain a BRV image corresponding to the environmental image. The acquisition cost of the BEV image is relatively low.

[0194] In some embodiments, the robot controller acquires multiple reference environment images whose generation time is adjacent to that of the environment image. The robot controller arranges the environment image and the multiple reference environment images in chronological order of generation time to obtain an environment image set. The robot controller inputs the environment image set into a first perspective transformation model, and performs feature extraction on the environment image set through the first perspective transformation model to obtain image set features of the environment image set. The robot controller encodes the image set features based on the attention mechanism through the first perspective transformation model to obtain BEV features of the environment image set. The robot controller performs multiple rounds of iterative decoding on the BEV features through the first perspective transformation model to obtain the BEV image corresponding to the environment image.

[0195] Among them, the first perspective transformation model can adopt any type of BEV perception model based on image sequence (environmental image set) in the relevant technology, and the embodiment of the present application is not limited to this.

[0196] Another implementation method of obtaining the BEV image corresponding to the environment image is described below.

[0197] In some embodiments, the robot controller obtains depth information corresponding to the environment image and performs a perspective transformation on the environment image based on the environment image and the depth information to obtain a BEV image corresponding to the environment image.

[0198] The depth information includes the distance between pixels in the environment image.

[0199] The depth information corresponding to the environmental image is obtained through the point cloud corresponding to the environmental image, or is obtained synchronously when the environmental image is obtained (for example, the visual sensor is a depth camera), which is not limited in this embodiment of the present application.

[0200] In this embodiment, the perspective transformation of the environment image is performed by combining the environment image and the depth information, and the obtained BEV image has higher accuracy.

[0201] In some embodiments, the robot controller inputs the environmental image into a second perspective transformation model, extracts features from the environmental image using the second perspective transformation model, and obtains environmental image features for the environmental image. The robot controller uses the second temporal transformation model to concatenate the depth information of the environmental image with the environmental image features to obtain depth image features. The robot controller uses the second temporal transformation model to encode the depth image features based on an attention mechanism to obtain BEV features for the environmental image. The robot controller uses the second perspective transformation model to perform multiple rounds of iterative decoding on the BEV features to obtain a BEV image corresponding to the environmental image.

[0202] Among them, the second perspective transformation model can adopt any type of BEV perception model based on image sequence and depth information in the relevant technology, and the embodiment of the present application is not limited to this.

[0203] C. The robot controller performs target detection on the BEV image to obtain the outlines of each candidate garbage.

[0204] Among them, target detection is to find each candidate garbage from the BEV image and then obtain the outline of each candidate garbage.

[0205] In some embodiments, the robot controller performs feature extraction on the BEV image to obtain BEV image features. Based on the BEV image features, the robot controller performs garbage detection to obtain multiple candidate detection frames, each corresponding to a piece of garbage. The robot controller then performs contour recognition within the multiple candidate detection frames to obtain contours for each piece of garbage.

[0206] In some embodiments, the robot controller inputs the BEV image into a target detection model, and uses the target detection model to perform multiple convolutions, multiple full connections, or attention-based encoding on the BEV image to obtain BEV image features of the BEV image. Using the target detection model, the robot controller controls a candidate recognition frame to slide over the BEV image features, where the candidate recognition frame is used to identify the candidate garbage. Using the target detection model, the robot controller determines multiple candidate detection frames based on sub-image features corresponding to multiple candidate image regions within the BEV image features. The candidate image regions are the areas covered by the candidate recognition frames. The robot controller performs boundary recognition within the multiple candidate detection frames to obtain the outlines of each candidate garbage item.

[0207] Among them, the target detection model can be any type of target detection model in the relevant technology, and the embodiments of the present application are not limited to this.

[0208] In the second part, the robot controller determines the circumscribed rectangle of each candidate garbage based on the outline of each candidate garbage.

[0209] In some embodiments, for any one of the multiple candidate trash items, the robot controller determines four vertices corresponding to the outline of the candidate trash item, where the four vertices are the outermost vertices of the outline of the candidate trash item. The robot controller sequentially connects the four vertices to obtain a bounding rectangle for the candidate trash item.

[0210] For example, see FIG9 , which shows the outline (dashed line) and circumscribed rectangle (solid line) of the candidate garbage. d Indicates the length of the circumscribed rectangle, X d Indicates the width of the bounding rectangle.

[0211] In the third part, the robot controller determines the garbage type of each candidate garbage based on the size information of the circumscribed rectangle of each candidate garbage.

[0212] In some embodiments, the size information includes length and width. For any of the multiple candidate garbage items, if both the length and width of the bounding rectangle of the candidate garbage item are less than or equal to a preset threshold, the robot controller determines the target garbage type of the target garbage item as a point type. The preset threshold is associated with the single cleaning width of the pool cleaning robot. If the length of the bounding rectangle of the candidate garbage item is less than or equal to the preset threshold and the width is greater than the preset threshold, or if the width of the bounding rectangle of the candidate garbage item is less than or equal to the preset threshold and the length is greater than the preset threshold, the robot controller determines the target garbage type of the target garbage item as a line segment type. If both the length and width of the bounding rectangle of the candidate garbage item are greater than the preset threshold, the robot controller determines the target garbage type of the target garbage item as a surface type.

[0213] Among them, the single cleaning width of the pool cleaning robot refers to the maximum effective cleaning width of the pool cleaning robot during its movement, that is, the garbage within the single cleaning width of the pool cleaning robot can be effectively removed. Generally speaking, the maximum effective cleaning width is slightly smaller than the width of the sewage suction port of the pool cleaning robot. The single cleaning width is determined by the structure of the pool cleaning robot, and the embodiment of the present application does not limit this. The preset threshold is smaller than the single cleaning width. Generally speaking, the preset threshold is slightly smaller than the single cleaning width. The preset threshold is set by technical personnel according to actual conditions, and the embodiment of the present application does not limit this. The length and width of the circumscribed rectangle of the candidate garbage are both smaller than or equal to the preset threshold, indicating that the pool cleaning robot can clean the candidate garbage by moving to the location of the candidate garbage. Therefore, the candidate garbage can be regarded as a "point", and the candidate garbage is defined as a point type. If the length of the candidate garbage's circumscribed rectangle is less than or equal to the preset threshold, and the width is greater than the preset threshold, or if the width of the candidate garbage's circumscribed rectangle is less than or equal to the preset threshold, this indicates that the pool cleaning robot can clean the candidate garbage with a single movement along the length or width of the candidate garbage. Therefore, the candidate garbage can be considered a "line segment," and the candidate garbage is defined as a line segment type. If both the length and width of the candidate garbage's circumscribed rectangle are greater than the preset threshold, this indicates that the pool cleaning robot needs to make multiple movements along the length or width of the candidate garbage to clean it. Therefore, the candidate garbage can be considered a "surface," and the candidate garbage is defined as a surface type.

[0214] In this implementation, the candidate garbage is classified using the relationship between the length and width of the circumscribed rectangle of the candidate garbage and the preset threshold, and the classification efficiency is high.

[0215] Another implementation of the third part is described below.

[0216] In some embodiments, the size information includes area. For any of the multiple candidate trash items, if the area of ​​the bounding rectangle of the candidate trash item is less than or equal to a first area threshold, the robot controller determines the target trash type of the target trash item as a point type. If the area of ​​the bounding rectangle of the candidate trash item is greater than the first area threshold and less than or equal to a second area threshold, the robot controller determines the target trash type of the target trash item as a line segment type. If the area of ​​the bounding rectangle of the candidate trash item is greater than the second area threshold, the robot controller determines the target trash type of the target trash item as a surface type.

[0217] Among them, the first area threshold and the second area threshold are set by technical personnel according to actual conditions, and the embodiments of the present application do not limit this.

[0218] In this implementation, the candidate garbage is classified using the relationship between the area of ​​the circumscribed rectangle of the candidate garbage and the first area threshold and the second area threshold, and the classification efficiency is high.

[0219] 802. The robot controller determines the cleaning cost between the pool cleaning robot and each candidate garbage based on the position of the pool cleaning robot, the garbage type, position, and garbage area of ​​each candidate garbage.

[0220] The garbage area of ​​the candidate garbage can refer to the area enclosed by the outline of the candidate garbage or the area of ​​the circumscribed rectangle of the candidate garbage, and this embodiment of the present application does not limit this. The cleaning cost is used to represent the cost required for the pool cleaning robot to clean the corresponding candidate garbage. In the embodiment of the present application, the cleaning cost includes a distance cost and an area cost. The distance cost represents the cost corresponding to the distance the pool cleaning robot needs to move to the location of the candidate garbage, and the area cost represents the cost corresponding to the cleaning area when the pool cleaning robot cleans the candidate garbage. In the embodiment of the present application, the cleaning cost is used to select the target garbage to be cleaned.

[0221] In some embodiments, the robot controller determines the distance between the pool cleaning robot and each candidate garbage based on the location of the pool cleaning robot and the garbage type and location of each candidate garbage. The robot controller determines the cleaning cost between the pool cleaning robot and each candidate garbage based on the distance between the pool cleaning robot and each candidate garbage and the garbage area of ​​each candidate garbage.

[0222] In order to explain the above embodiment more clearly, a method of determining the garbage area of ​​candidate garbage will be described first.

[0223] In some embodiments, the robot controller determines the area of ​​the circumscribed rectangle of each candidate garbage as the garbage area of ​​each candidate garbage. The robot controller determines the area enclosed by the outline of each candidate garbage as the garbage area of ​​each candidate garbage.

[0224] In order to explain the above embodiment more clearly, the following will be divided into several parts to explain the above embodiment.

[0225] In the first part, the robot controller determines the distance between the pool cleaning robot and each candidate garbage based on the position of the pool cleaning robot and the garbage type and position of each candidate garbage.

[0226] In some embodiments, for any of the multiple candidate trash items, if the candidate trash item is a point type, the robot controller determines the center position of the center point of the candidate trash item based on the location of the candidate trash item. The robot controller determines the distance between the pool cleaning robot and the center position as the distance between the pool cleaning robot and the candidate trash item. If the candidate trash item is a line segment type, the robot controller determines the midpoint position of the midpoints of the two short sides of the circumscribed rectangle of the candidate trash item based on the location of the candidate trash item. The robot controller determines the distance between the pool cleaning robot and the candidate trash item based on the location of the pool cleaning robot and the midpoint position of the two short sides. If the candidate trash item is a surface type, the robot controller determines the vertex positions of the four vertices of the circumscribed rectangle of the candidate trash item based on the location of the candidate trash item. The robot controller determines the distance between the pool cleaning robot and the candidate trash item based on the location of the pool cleaning robot and the vertex positions of the four vertices.

[0227] The location of the candidate garbage includes the locations of at least three vertices of the circumscribed rectangle of the candidate garbage. The short sides of the circumscribed rectangle refer to the two shorter sides of the circumscribed rectangle, and correspondingly, the long sides of the circumscribed rectangle refer to the two longer sides of the circumscribed rectangle.

[0228] In order to explain the above embodiment more clearly, the following will be divided into several parts to explain the above embodiment.

[0229] A. When the candidate garbage type is a point type, the robot controller determines the center position of the center point of the candidate garbage based on the position of the candidate garbage.

[0230] In some embodiments, when the candidate garbage type is a point type, the robot controller determines the position of the geometric center of the circumscribed rectangle of the candidate garbage based on the positions of at least three vertices of the circumscribed rectangle of the candidate garbage, and the robot controller determines the position of the geometric center as the center position of the center point of the candidate garbage.

[0231] B. The robot controller determines the distance between the position of the pool cleaning robot and the center position as the distance between the pool cleaning robot and the candidate garbage.

[0232] In some embodiments, the robot controller determines the distance between the position of the pool cleaning robot and the center position based on the first coordinate corresponding to the position of the pool cleaning robot and the second coordinate corresponding to the center position, thereby obtaining the distance between the pool cleaning robot and the candidate garbage.

[0233] C. When the candidate garbage type is a line segment type, the robot controller determines the midpoint position of the two short side midpoints of the circumscribed rectangle of the candidate garbage based on the position of the candidate garbage.

[0234] In some embodiments, when the candidate garbage type is a line segment type, the robot controller determines the midpoint position of the midpoints of the two short sides of the circumscribed rectangle of the candidate garbage based on the positions of at least three vertices of the circumscribed rectangle of the candidate garbage.

[0235] D. The robot controller determines the distance between the pool cleaning robot and the candidate garbage based on the position of the pool cleaning robot and the midpoint position of the midpoints of the two short sides.

[0236] In some embodiments, the two short side midpoints include a first short side midpoint and a second short side midpoint, and the robot controller determines a first short side reference distance between the position of the pool cleaning robot and the midpoint of the first short side midpoint. The robot controller determines a second short side reference distance between the position of the pool cleaning robot and the midpoint of the second short side midpoint. The robot controller determines the shorter of the first short side reference distance and the second short side reference distance as the distance between the pool cleaning robot and the candidate garbage.

[0237] E. When the candidate garbage type is a surface type, the robot controller determines the vertex positions of the four vertices of the circumscribed rectangle of the candidate garbage based on the position of the candidate garbage.

[0238] In some embodiments, when the candidate garbage type is a face type, the robot controller determines the vertex positions of four vertices of the circumscribed rectangle of the candidate garbage based on the positions of at least three vertices of the circumscribed rectangle of the candidate garbage.

[0239] F. The robot controller determines the distance between the pool cleaning robot and the candidate garbage based on the position of the pool cleaning robot and the vertex positions of the four vertices.

[0240] In some embodiments, the four vertices include a first vertex, a second vertex, a third vertex, and a fourth vertex, and the robot controller determines a first vertex reference distance between the position of the pool cleaning robot and the vertex position of the first vertex. The robot controller determines a second vertex reference distance between the position of the pool cleaning robot and the vertex position of the second vertex. The robot controller determines a third vertex reference distance between the position of the pool cleaning robot and the vertex position of the third vertex. The robot controller determines a fourth vertex reference distance between the position of the pool cleaning robot and the vertex position of the fourth vertex. The robot controller determines the shorter of the first vertex reference distance, the second vertex reference distance, the third vertex reference distance, and the fourth vertex reference distance as the distance between the pool cleaning robot and the candidate garbage.

[0241] In the second part, the robot controller determines the cleaning cost between the pool cleaning robot and each candidate garbage based on the distance between the pool cleaning robot and each candidate garbage and the garbage area of ​​each candidate garbage.

[0242] In some embodiments, the robot controller multiplies the distance cost weight by the distance between the pool cleaning robot and each candidate trash item to obtain a distance cost between the pool cleaning robot and each candidate trash item. The robot controller divides the area cost weight by the trash area of ​​each candidate trash item to obtain an area cost for each candidate trash item. The robot controller adds the distance cost and area cost corresponding to each candidate trash item to obtain a cleaning cost between the pool cleaning robot and each candidate trash item.

[0243] Among them, the distance cost weight and the area cost weight are set by technical personnel according to actual conditions, and the embodiments of this application do not limit this.

[0244] In some embodiments, the above embodiment can be expressed by the following formula (1): cost = ω d *dis+ω a / area (1)

[0245] Among them, cost is the cleaning cost, ω d is the distance cost weight, dis is the distance between the pool cleaning robot and each candidate garbage, ω a is the area cost weight, area is the garbage area of ​​each candidate garbage, ω d *dis is the distance cost, ω a / area is the area cost.

[0246] 803. The robot controller determines the target garbage from the multiple candidate garbage based on the cleaning costs between the pool cleaning robot and each candidate garbage. The target garbage is the candidate garbage with the minimum cleaning cost.

[0247] The target garbage is the garbage selected from multiple candidate garbage to be cleaned.

[0248] In some embodiments, the robot controller determines the candidate garbage with the lowest cleaning cost among the multiple candidate garbages as the target garbage.

[0249] 804. The robot controller obtains the position of the pool cleaning robot and the position of target garbage to be cleaned around the pool cleaning robot.

[0250] The pool cleaning robot also includes a positioning component for determining the pool cleaning robot's position, thereby locating the pool cleaning robot. The robot controller can combine the positioning component with information collected by the visual sensor to determine the location of trash around the pool cleaning robot, thereby locating the trash. In step 802 above, the pool cleaning robot's position and the location of the candidate trash items were used to determine the cleaning cost. If the target trash item belongs to multiple candidate items, then in step 804, the result of step 802 can be directly used.

[0251] In order to more clearly illustrate the technical solution provided in the embodiment of the present application, the method of determining the position of the pool cleaning robot and the position of the candidate garbage is described below.

[0252] In some embodiments, the robot controller determines the position of the pool cleaning robot in the cleaning area through a positioning component. The robot controller determines the position of each candidate garbage in the cleaning area through a visual sensor and the position of the pool cleaning robot in the cleaning area.

[0253] The cleaning area refers to the area where the pool cleaning robot is performing cleaning tasks. Generally speaking, the cleaning area is the area where the pool to be cleaned is located. Before the pool cleaning robot performs cleaning tasks in the cleaning area, the robot controller will map the cleaning area to obtain a regional map corresponding to the cleaning area. Accordingly, the position of the pool cleaning robot in the cleaning area refers to the position of the pool cleaning robot in the regional map, and the position of the candidate garbage in the cleaning area refers to the position of the candidate garbage in the regional map. In the above embodiment, the visual sensor is used to determine the relative position between the candidate garbage and the pool cleaning robot. This relative position, combined with the position of the pool cleaning robot, can determine the position of the candidate garbage in the cleaning area. The position of the candidate garbage includes the positions of at least three vertices of the circumscribed rectangle of the candidate garbage.

[0254] 805. The robot controller generates a target cleaning path for the pool cleaning robot based on the position of the pool cleaning robot, the position of the target garbage, and the target garbage type of the target garbage. The target cleaning path is the cleaning path used when cleaning the target garbage, and the target garbage type is used to represent the shape of the target garbage.

[0255] The target waste type represents the shape of the target waste. That is, in this embodiment of the present application, the waste type is determined based on the shape of the waste. Accordingly, the target cleaning path generated using the target waste type is also associated with the shape of the waste. In this embodiment of the present application, the target cleaning path includes two cleaning paths. The first cleaning path is the path the pool cleaning robot takes to the location of the target waste, and the second cleaning path is the path used to clean the target waste. The target waste type significantly influences the second cleaning path. The location of the target waste includes the positions of at least three vertices of the circumscribed rectangle of the target waste.

[0256] In some embodiments, the target cleaning path is a first cleaning path, a second cleaning path, or a third cleaning path. When the target garbage type is a point type, the robot controller determines the center position of the center point of the target garbage based on the position of the target garbage. The robot controller performs path planning between the position of the pool cleaning robot and the center position to obtain the first cleaning path. When the target garbage type is a line segment type, the robot controller determines the line segment position of the line segment corresponding to the target garbage based on the position of the target garbage. The robot controller generates the second cleaning path based on the position of the pool cleaning robot and the line segment position. When the target garbage type is a surface type, the robot controller determines the area where the target garbage is located based on the position of the target garbage. The robot controller generates the third cleaning path based on the position of the pool cleaning robot and the area where the target garbage is located.

[0257] The purpose of path planning is to obtain the shortest cleaning path while avoiding obstacles. The robot controller can adopt any path planning method in the relevant technology to obtain the cleaning path, and the embodiments of the present application are not limited to this.

[0258] In order to explain the above embodiment more clearly, the following will be divided into several parts to explain the above embodiment.

[0259] A. When the target garbage type is a point type, the robot controller determines the center position of the center point of the target garbage based on the position of the target garbage.

[0260] In some embodiments, when the target garbage type is a point type, the robot controller determines the position of the geometric center of the circumscribed rectangle of the target garbage based on the positions of at least three vertices of the circumscribed rectangle of the target garbage, and the robot controller determines the position of the geometric center as the center position of the center point of the target garbage.

[0261] B. The robot controller performs path planning between the position of the pool cleaning robot and the center position to obtain the first cleaning path.

[0262] In some embodiments, the robot controller performs path planning between the position of the pool cleaning robot and the center position to obtain the first reference cleaning path. The robot controller extends the first reference cleaning path by a preset length in the original direction to obtain the first cleaning path.

[0263] The first cleaning path is a cleaning path that passes through the center point of the target garbage.

[0264] In some embodiments, referring to FIG. 10 , a first cleaning path 500 starts from a pool cleaning robot 501 and passes through a center point of a point-type target garbage 502 .

[0265] C. When the target garbage type is a line segment type, the robot controller determines the line segment position of the line segment corresponding to the target garbage based on the position of the target garbage.

[0266] The line segment position includes the midpoint of the two short sides of the circumscribed rectangle of the target garbage.

[0267] In some embodiments, when the target garbage type is a line segment type, the robot controller determines the midpoint position of the midpoints of the two short sides of the circumscribed rectangle of the target garbage based on the positions of at least three vertices of the circumscribed rectangle of the target garbage.

[0268] D. The robot controller generates the second cleaning path based on the position of the pool cleaning robot and the position of the line segment.

[0269] In some embodiments, the robot controller determines the starting point of the line segment corresponding to the target waste from the line segment position. The robot controller performs path planning between the pool cleaning robot's position and the starting point to obtain a first initial cleaning path. The robot controller adds a second initial cleaning path of the line segment corresponding to the target waste to the first initial cleaning path to obtain a second cleaning path, wherein the second initial cleaning path passes through the starting point and the end point of the line segment corresponding to the target waste.

[0270] The starting point refers to the midpoint of the two short side midpoints that is closer to the pool cleaning robot, and the closer short side midpoint is also called the starting point. Correspondingly, the end point refers to the midpoint of the two short side midpoints that is farther from the pool cleaning robot, and the farther short side midpoint is also called the end point.

[0271] In some embodiments, referring to FIG. 11 , the second cleaning path 600 is a cleaning path starting from the pool cleaning robot 601 and passing through the location of the line segment type target garbage 602 .

[0272] E. When the target garbage type is a surface type, the robot controller determines the area where the target garbage is located based on the position of the target garbage.

[0273] In some embodiments, when the target garbage type is a surface type, the robot controller determines the area where the target garbage is located based on the positions of at least three vertices of a circumscribed rectangle of the target garbage.

[0274] F. The robot controller generates the third cleaning path based on the position of the pool cleaning robot and the area where the target garbage is located.

[0275] In some embodiments, the robot controller determines the vertex position of any vertex of the target waste within the area where the target waste is located. The robot controller performs path planning between the pool cleaning robot's position and the vertex position to obtain a third initial cleaning path. The robot controller uses the vertex position as a starting point and performs path planning based on the area where the target waste is located to obtain a fourth initial cleaning path, which covers the area where the target waste is located. The robot controller combines the third initial cleaning path with the fourth initial cleaning path to obtain the third cleaning path.

[0276] The vertices of the target garbage refer to the vertices of the circumscribed rectangle of the target garbage.

[0277] In some embodiments, referring to FIG. 12 , a third cleaning path 700 is a cleaning path starting from a pool cleaning robot 701 and covering an area where a target garbage 702 of a surface type is located.

[0278] 806. The robot controller controls the pool cleaning robot to clean the target garbage based on the target cleaning path.

[0279] Controlling the pool cleaning robot to clean the target garbage based on the target cleaning path refers to controlling the pool cleaning robot to move according to the target cleaning path, thereby achieving the cleaning of the target garbage.

[0280] In some embodiments, the robot controller controls the pool cleaning robot to move according to the target cleaning path, turns on the water pump of the pool cleaning robot during the movement, and uses the water pump to suck the target garbage from the suction port into the filter box, thereby cleaning the target garbage.

[0281] Optionally, after step 806, the following steps can also be performed.

[0282] In some embodiments, after the target cleaning path is completed, the robot controller determines whether there is still candidate garbage to be cleaned. If there is still candidate garbage to be cleaned, the robot controller redefines the target garbage from the candidate garbage to be cleaned and redefines the target cleaning path. The robot controller controls the pool cleaning robot to clean the redetermined target garbage based on the redetermined target cleaning path.

[0283] Based on the above embodiment, if there is no candidate garbage to be cleaned, the robot controller performs path planning based on the location of the pool cleaning robot and the location of the base station to obtain a target recharging path. The robot controller controls the pool cleaning robot to dock with the base station along the target recharging path.

[0284] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0285] Using the technical solutions provided in the embodiments of this application, the position of the pool cleaning robot and the location of the target waste to be cleaned are obtained. Using the pool cleaning robot's position, the location of the target waste, and the target waste type, a target cleaning path is generated for the target waste. This target cleaning path closely matches the location and type of the target waste. By controlling the pool cleaning robot to clean the target waste based on the target cleaning path, the target waste can be more thoroughly removed, thereby achieving a better cleaning effect.

[0286] With the development of computer technology, robot technology has also developed rapidly. For example, users use sweeping robots to clean the floor of the house, use window cleaning robots to clean the windows of the house, and use pool cleaning robots to clean the pool.

[0287] In related technologies, the positioning and mapping capabilities of pool cleaning robots are easily invalidated in the case of weak textures, that is, they cannot meet the cleaning needs in weak texture environments and it is difficult to meet the needs of various user scenarios.

[0288] In some embodiments, the technical solutions provided in the embodiments of the present application can be executed independently by the robot controller of the pool cleaning robot, or by a terminal or server connected to the pool cleaning robot via a network, and the embodiments of the present application are not limited to this. The network can be a medium that provides a communication link between the pool cleaning robot and the terminal or server, or it can be the Internet including network equipment and transmission media, but is not limited thereto. The transmission medium can be a wired link, such as but not limited to coaxial cable, optical fiber and digital subscriber line (DSL), or a wireless link, such as but not limited to wireless fidelity (WIFI), Hypertext Transfer Protocol (HTTP), Bluetooth and mobile device network.

[0289] The terminal can interact with the pool cleaning robot through the network to receive messages from the pool cleaning robot or send messages to the pool cleaning robot. The terminal can be hardware or software. When the terminal is hardware, it can be various electronic devices, including but not limited to smart watches, smart phones, tablet computers, laptop portable computers and desktop computers. When the terminal is software, it can be installed in the electronic devices listed above, which can be implemented as multiple software or software modules (for example: to provide distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.

[0290] The above-mentioned server can be a server that provides control services for various pool cleaning robots. It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module, which is not specifically limited here. The server can be, but is not limited to, a hardware server, a virtual server, a cloud server, etc.

[0291] In response to the above problems, another implementation method provided in the embodiment of the present application is described in more detail with reference to some examples. Referring to FIG13 , the method includes the following steps.

[0292] 1301. The robot controller obtains an environmental image collected by the pool cleaning robot.

[0293] When the pool cleaning robot moves in the target space, it can collect environmental images within its visual range through a first sensor. The above-mentioned first sensor can include but is not limited to monocular cameras, binocular cameras and other image acquisition devices.

[0294] 1302. The robot controller extracts features from the environment image collected by the pool cleaning robot to obtain target structure information in the target space where the pool cleaning robot is located.

[0295] The target structure information may include, but is not limited to, at least one of the following: target space component element information and target space auxiliary element information in the target space; the target space component element information may include, but is not limited to, structural boundary information of the target space corresponding component element, and the component element may include, but is not limited to, at least one of the following: bottom surface, wall surface, step plane, and step elevation; the target space auxiliary element information may include, but is not limited to, structural boundary information of the target space corresponding auxiliary element, and the auxiliary element may include, but is not limited to, at least one of the following auxiliary devices: wall lamp, floor lamp, water inlet, and water outlet. The structural boundary information may include, but is not limited to, first position information corresponding to the structural boundary of the component element or auxiliary element in the environmental image and second position information of the contour containing area corresponding to the structural boundary in the environmental image.

[0296] In some embodiments, the above-mentioned feature extraction of the environmental image collected by the pool cleaning robot to obtain the target structure information in the target space where the pool cleaning robot is located may include, but is not limited to: first extracting the structural boundary information in the environmental image collected by the pool cleaning robot, the above-mentioned structural boundary information may include, but is not limited to, the first position information corresponding to the structural boundary existing in the environmental image and the second position information of the area included in the contour corresponding to the structural boundary. For example, but not limited to, a segmentation model, a boundary extraction model, or an edge recognition model may be used to segment and extract the constituent elements of the target space area and common devices (auxiliary elements) in the target space, and identify the pixel-level structural boundary coordinates (first position information) and the specific position information of the included area (second position information) in the image coordinate system corresponding to the environmental image. Then, based on the above-mentioned structural boundary information, the target structure information in the target space is determined. For example, but not limited to, the structural boundary information corresponding to each of the multiple environmental images collected during the movement of the pool cleaning robot can be fused to obtain the target structure information in the target space.

[0297] In some embodiments, the above-mentioned feature extraction of the environmental image collected by the pool cleaning robot to obtain the target structure information in the target space where the pool cleaning robot is located may also include, but is not limited to: directly using the multi-task model to segment and identify the target space component element information and the target space auxiliary element information in the environmental image collected by the pool cleaning robot, respectively, to obtain the target structure information in the target space where the pool cleaning robot is located. The above-mentioned target space component element information may include, but is not limited to, the structural boundary information of the components such as the pool bottom, pool wall, step plane, step facade, etc. that constitute the target space, and the above-mentioned target space auxiliary element information may include, but is not limited to, the structural boundary information of auxiliary devices (auxiliary elements) such as wall lamps, floor lamps, water inlets, water outlets, etc. installed in the target space. The above-mentioned multi-task model may be, but is not limited to, trained based on multiple sample environmental images with known spatial component element information and spatial auxiliary element information.

[0298] 1303. The robot controller performs feature matching based on target structure information in the target space corresponding to two adjacent frames of environmental images to obtain a corresponding target feature matching result.

[0299] In some embodiments, after obtaining the target structure information corresponding to the environmental image, the target structure information in the target space corresponding to the two adjacent frames of the environmental image may be subjected to feature matching, such as, but not limited to, similarity calculation, to obtain target structure boundary points that match the two adjacent frames of the environmental image, i.e., structure boundary points that belong to the same position on the same structure (component element or auxiliary element) in the two adjacent frames of the environmental image, or structure boundary points in the two adjacent frames of the environmental image whose corresponding similarity is greater than a threshold. The above-mentioned similarity calculation may be based on, but is not limited to, features such as the pixel, color, position, and slope or curvature of the corresponding position of the boundary line corresponding to the structure boundary point in the environmental image.

[0300] In some embodiments, as shown in FIG. 14 , the above 1303 performs feature matching based on target structure information in the target space corresponding to two adjacent frames of environmental images to obtain a corresponding target feature matching result, and the implementation process may include but is not limited to the following steps:

[0301] 1401. The robot controller predicts predicted acquisition position information of the pool cleaning robot corresponding to the latter frame of the two adjacent frames of environmental images based on the previous frame acquisition position information of the pool cleaning robot corresponding to the previous frame of the two adjacent frames of environmental images.

[0302] Among them, the pool cleaning robot can, but is not limited to, use its built-in kinematic model to estimate its moving direction and distance during the interval between two adjacent frames of environmental images based on the acquisition time interval corresponding to the two adjacent frames of environmental images, as well as the moving speed, steering angle, etc., and then combine the estimated moving direction and distance with the previous frame acquisition position information of the pool cleaning robot corresponding to the previous frame of environmental image to determine the predicted acquisition position information of the pool cleaning robot corresponding to the next frame of environmental image.

[0303] 1402. The robot controller predicts predicted structure information in a next frame of the environment image based on the predicted acquisition position information and the target structure information in the previous frame of the environment image.

[0304] Among them, after obtaining the predicted acquisition position information of the pool cleaning robot corresponding to the next frame of the environmental image, it is also possible, but not limited to, to translate and / or rotate the target structure information in the previous frame of the environmental image based on the difference between the previous frame of the environmental image corresponding to the pool cleaning robot and the predicted acquisition position information of the next frame of the environmental image corresponding to the pool cleaning robot, so as to obtain the corresponding predicted structure information in the next frame of the environmental image, that is, the structural boundary information of the constituent elements and / or auxiliary elements in the previous frame of the environmental image corresponding to the structural boundary in the next frame of the environmental image.

[0305] 1403. The robot controller matches the target structure information in the next frame of the environment image with the predicted structure information to obtain a predicted structure matching range corresponding to the target structure information in the next frame of the environment image.

[0306] Among them, after obtaining the target structure information actually extracted and the predicted structure information predicted in the next frame of the environmental image, the target structure information corresponding to the next frame of the environmental image can be directly matched with the predicted structure information, but is not limited to, for example, but not limited to matching the target structure information with the corresponding structural boundary points in the predicted structure information (for example, but not limited to similarity calculation, etc.), and screening out the set of structural boundary points in which the target structure information matches the predicted structure information (for example, but not limited to the similarity being greater than a preset value), thereby obtaining the predicted structure matching range corresponding to the target structure information in the next frame of the environmental image.

[0307] 1404. The robot controller determines the matching target structure boundary points in the target structure information corresponding to two adjacent frames of the environment image based on the geometric features of each structure boundary point in the previous frame of the environment image and the geometric features of each structure boundary point within the predicted structure matching range.

[0308] Among them, the above-mentioned geometric features may include, but are not limited to, slopes and curvatures corresponding to the structural boundary points. After obtaining the predicted structural matching range corresponding to the next frame of the environmental image, in order to ensure the accuracy of the feature matching, it is also possible but not limited to directly calculating the difference or similarity between the geometric features of each structural boundary point in the previous frame of the environmental image and the geometric features of each structural boundary point within the predicted structural matching range corresponding to the next frame of the environmental image, and then determining the matching target structural boundary points in the target structural information corresponding to each of the two adjacent frames of the environmental image based on the above-mentioned difference or similarity. The above-mentioned matching target structural boundary points may include, but are not limited to, structural boundary points whose difference is less than the target value or whose similarity is greater than the threshold.

[0309] In an embodiment of the present application, by performing feature matching based on the target structure information in the corresponding target space of two adjacent frames of environmental images in the above manner, the matching target structure boundary points in the two adjacent frames of environmental images can be accurately found, thereby ensuring the continuity and accuracy of subsequent mapping.

[0310] Next, please continue to refer to FIG. 13 . As shown in FIG. 13 , after performing feature matching based on target structure information in the target space corresponding to two adjacent frames of environmental images in step 1302 and obtaining corresponding target feature matching results, the technical solution further includes:

[0311] 1304. The robot controller determines the target depth information of the target structure boundary point in the target space in the target structure information based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot. The target acquisition position information is comprehensively determined based on the positioning information obtained by the various sensors of the pool cleaning robot. The target structure boundary point is the structure boundary point that matches in the target structure information corresponding to each of two adjacent frames of environmental images.

[0312] The multiple sensors described above may include, but are not limited to, a first sensor, a second sensor, and an ultrasonic sensor. The first sensor is used to collect environmental images of the pool cleaning robot, and the second sensor is used to collect motion data of the pool cleaning robot. The second sensor may include, but is not limited to, at least one of the following: an inertial measurement unit, an encoder, or a flow meter. The ultrasonic sensor is used to collect ultrasonic positioning data of the pool cleaning robot. The target acquisition position information is the actual acquisition position information of the pool cleaning robot corresponding to the latter of two adjacent frames of environmental images corresponding to the target feature matching results. The target depth information may include, but is not limited to, the position coordinates of the target structure boundary points in the three-dimensional target space.

[0313] Furthermore, the target feature matching result includes target structure boundary points that match in two adjacent frames of environmental images. The above-mentioned 1304, the implementation process of determining the target depth information of the target structure boundary point in the target space based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot, may include, but is not limited to: estimating the target depth information of the target structure boundary point in the target space based on the position of the target structure boundary point in two adjacent frames of environmental images and the target acquisition position information corresponding to the pool cleaning robot, for example, but not limited to, first determining the relative position information of the target structure boundary point and the pool cleaning robot based on the difference between the positions of the target structure boundary point in two adjacent frames of environmental images, and then calculating the target depth information of the target structure boundary point in the target space based on the relative position information and the target acquisition position information corresponding to the pool cleaning robot, that is, the position coordinates of the pool cleaning robot in the target space when the next frame of environmental image is acquired.

[0314] 1305. The robot controller constructs a target environment map of the target space based on the target structure information and the target depth information corresponding to the boundary points of the target structure.

[0315] Among them, it is possible but not limited to directly using the target structure information corresponding to the collected environment image and the target depth information corresponding to the target structure boundary points during the pool cleaning robot mapping process to construct a target environment map corresponding to the target space.

[0316] In some embodiments, when a pool cleaning robot is required to map a target space, upon receiving the corresponding mapping instructions, the pool cleaning robot may, but is not limited to, move along a target mapping path corresponding to the target space, and during movement, collect environmental images within a preset visual range at a preset frequency. After completing the target mapping path, the pool cleaning robot may construct a target environmental map of the target space based on the target structure information corresponding to each of the multiple environmental images along the target mapping path, as well as the target depth information corresponding to the target structure boundary points that match between two adjacent frames of the multiple environmental images.

[0317] In some embodiments, after traversing the target space according to the target mapping path, the pool cleaning robot can construct a target environment map of the target space based on all target mapping paths during the traversal process and the corresponding target structure information and target depth information corresponding to the target structure boundary points. It can also construct a target environment map of the target space in real time according to the real-time changing target mapping path and the corresponding target structure information and target depth information corresponding to the target structure boundary points during the traversal of the target space, that is, construct a target environment map for the corresponding traversed positions in the target space while traversing the target space, etc. The embodiments of the present application are not limited to this.

[0318] In some embodiments, after constructing the target environment map of the target space, the pool cleaning robot can also use various optimization algorithms to optimize the target environment map, such as removing mismatched features, adjusting the scale of the map, etc., to improve the accuracy and quality of the target environment map.

[0319] In an embodiment of the present application, on the one hand, target acquisition position information corresponding to the pool cleaning robot is obtained by comprehensively utilizing multiple sensors to ensure the positioning accuracy of the pool cleaning robot in an underwater environment with a large number of weak textures; on the other hand, feature matching is performed based on the target structure information in the target space corresponding to two adjacent frames of environmental images, and the target depth information of the matching structural boundary points in the target space in the two adjacent frames of environmental images is estimated based on the target acquisition position information obtained by multiple sensors, so as to construct a target environment map of the target space based on the extracted target structure information and the corresponding target depth information, thereby comprehensively utilizing multiple sensors, feature extraction and feature matching, and depth estimation to realize perception and analysis of the internal structure of the swimming pool (target space), thereby improving the mapping capability and stability of the pool cleaning robot.

[0320] In some possible embodiments, as shown in FIG15 , before determining target depth information of the target structure boundary point in the target structure information in the target space based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot in step 1303, the technical solution may also include, but is not limited to: first acquiring two adjacent frames of environmental images based on the first sensor, acquiring actual motion data of the pool cleaning robot (such as, but not limited to, actual moving distance, direction, speed, etc.) based on one or more second sensors, and acquiring ultrasonic positioning data of the pool cleaning robot based on the ultrasonic sensor;

[0321] Then, based on the two adjacent frames of environmental images and the previous frame acquisition position information corresponding to the previous frame of the environmental image in the two adjacent frames of environmental images, the first acquisition position information corresponding to the next frame of the environmental image in the two adjacent frames of environmental images is determined (for example, but not limited to, the first acquisition position information corresponding to the next frame of the environmental image can be calculated based on the position change information of the same structural element in the two adjacent frames of environmental images and the previous frame acquisition position information corresponding to the previous frame of the environmental image), and the second acquisition position information corresponding to the next frame of the environmental image is determined based on the actual motion data (for example, but not limited to, the corresponding moving distance and direction between the two adjacent frames of environmental images and the previous frame acquisition position information corresponding to the previous frame of the environmental image can be calculated). and determining the third acquisition position information corresponding to the next frame of the environmental image based on the ultrasonic positioning data (for example, but not limited to, directly determining the position information indicated in the ultrasonic positioning data as the third acquisition position information corresponding to the next frame of the environmental image); finally, determining the target acquisition position information corresponding to the next frame of the environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information. For example, but not limited to, weighted fusion of the first acquisition position information, the second acquisition position information, and the third acquisition position information can be performed according to a preset weight ratio to obtain the target acquisition position information corresponding to the next frame of the environmental image.

[0322] In an embodiment of the present application, the pool cleaning robot can be positioned comprehensively through multiple sensors such as visual sensors (first sensors), motion sensors (second sensors), and ultrasonic sensors, thereby enabling perception of the current environment in a weak texture environment while ensuring real-time and accurate positioning of the pool cleaning robot.

[0323] In some embodiments, in order to further ensure the accuracy of the real-time positioning of the pool cleaning robot and improve the positioning accuracy and positioning capability of the pool cleaning robot, as shown in FIG16 , after the target acquisition position information corresponding to the next frame of the environmental image is determined based on the first acquisition position information, the second acquisition position information and the third acquisition position information, the technical solution further includes: first determining the position loss corresponding to the next frame of the environmental image based on the predicted acquisition position information, the first acquisition position information, the second acquisition position information and the third acquisition position information corresponding to the next frame of the environmental image, for example but not limited to first calculating the position differences between the predicted acquisition position information and the first acquisition position information, the second acquisition position information and the third acquisition position information, and then weightedly fusing the corresponding position differences to obtain the position loss corresponding to the next frame of the environmental image; then, based on The image loss corresponding to the next frame of the environmental image is determined based on the predicted structural information and the target structural information corresponding to the next frame of the environmental image, for example, but not limited to, the position loss corresponding to the next frame of the environmental image is comprehensively calculated based on the position deviation value and the geometric feature deviation value between the corresponding structural boundaries in the above-mentioned predicted structural information and the target structural information; finally, the position drift of multiple sensors is corrected based on the position loss and image loss corresponding to the next frame of the environmental image, for example, but not limited to, adjusting the internal parameters of multiple sensors such as temperature drift, zero votes, etc., or the positioning data collected by the sensor (such as an ultrasonic sensor) itself, so as to minimize the position loss and image loss corresponding to the above-mentioned next frame of the environmental image, thereby continuously correcting the positioning error of the pool cleaning robot during the mapping process, and improving the positioning accuracy of the next frame of the environmental image corresponding to the next frame of the environmental image.

[0324] In some embodiments, in order to eliminate the interference of abnormal ultrasonic positioning data on the precise positioning of the pool cleaning robot, before determining the target collection position information corresponding to the next frame of the environmental image based on the first collection position information, the second collection position information and the third collection position information, it is also possible but not limited to first judging whether the ultrasonic positioning data collected by the ultrasonic sensor meets the non-line-of-sight condition based on the next frame of the environmental image. The above-mentioned non-line-of-sight condition is used to characterize the situation where there is occlusion in the next frame of the environmental image and the ultrasonic positioning data is abnormally jittery. The above-mentioned abnormal jitter situation is used to characterize the situation where the difference between the ultrasonic positioning data corresponding to the next frame of the environmental image and the ultrasonic positioning data corresponding to the previous frame of the environmental image is greater than a preset ultrasonic positioning threshold, that is, the ultrasonic positioning jitter amplitude is large; If the ultrasonic positioning data corresponding to the next frame of the environmental image meets the non-line-of-sight condition, it means that the ultrasonic positioning data is abnormal positioning data and there may be a large positioning error. In this case, the ultrasonic positioning data will be discarded, and only the target acquisition position information corresponding to the next frame of the environmental image will be determined based on the first acquisition position information and the second acquisition position information. If the ultrasonic positioning data does not meet the non-line-of-sight condition, it means that the ultrasonic positioning data is normal positioning data and the positioning is relatively accurate. In this case, the above steps of determining the target acquisition position information corresponding to the next frame of the environmental image based on the first acquisition position information, the second acquisition position information and the third acquisition position information can be directly executed, thereby only using more accurate ultrasonic positioning data to assist the pool cleaning robot in meeting the more accurate positioning requirements in a weak texture environment.

[0325] In some embodiments, in order to avoid the influence of abnormal visual tracking of the pool cleaning robot on its positioning accuracy, as shown in FIG17, before determining the target acquisition position information corresponding to the next frame of the environment image based on the first acquisition position information, the second acquisition position information and the third acquisition position information, it is also possible but not limited to first

[0326] The number of target structure boundary points that meet preset constraints in the next frame of the environmental image and the total number of structure boundary points in the corresponding target structure information of the next frame of the environmental image are counted, wherein the preset constraint condition is that there are matching target structure boundary points in at least two consecutive frames of environmental images. The preset constraint condition may include, but is not limited to, good constraints (i.e., there are matching target structure boundary points in more than two consecutive frames of environmental images, i.e., the same target structure boundary point is tracked in more than two consecutive frames of environmental images) and general constraints (i.e., there are matching target structure boundary points only in two consecutive frames of environmental images, i.e., the same target structure boundary point is tracked only in two consecutive frames of environmental images). Since the higher the ratio between the number of target structure boundary points that meet the preset constraint conditions and the total number of structure boundary points, the better the tracking effect of the pool cleaning robot, and the higher the accuracy of visual positioning based on the environmental image collected by the underwater robot, the visual constraint weight corresponding to the first sensor can be determined directly based on the ratio between the number of target structure boundary points that meet the preset constraint conditions and the total number of structure boundary points. For example, but not limited to, different ratios correspond to different visual constraint weights, or, when the ratio is greater than or equal to a first threshold (for example, but not limited to 0.9, 0.8, etc.), the visual constraint weight is set to a first weight (for example, but not limited to 1, 0.9, etc.), and when the ratio is less than the first threshold and greater than a second threshold (for example, but not limited to 0.6, 0.5, etc.), the visual constraint weight is set to a second weight (for example, but not limited to 0.6, 0.5, etc.). etc.), when the ratio is less than or equal to the second threshold, the visual constraint weight is set to a third weight (for example, but not limited to 0, 0.2, etc.); at the same time, the weights corresponding to the second sensor and the ultrasonic sensor can be determined based on their respective built-in covariance matrices; finally, the target weight ratios corresponding to the first sensor, the second sensor and the ultrasonic sensor can be determined based on the visual constraint weight corresponding to the first sensor and the weights corresponding to the second sensor and the ultrasonic sensor, and the first acquisition position information obtained by the first sensor, the second acquisition position information obtained by the second sensor, and the third acquisition position information obtained by the ultrasonic sensor are weightedly fused according to the above target weight ratio to obtain more accurate target acquisition position information corresponding to the next frame of environmental image.

[0327] With the rapid development of computer vision object detection technology in recent years, the effectiveness of tracking through detection has been significantly improved, and its real-world applications are increasing. Currently, most algorithms for multi-target tracking using computer vision object detection use a Kalman filter based on the constant velocity motion model to predict the position of the currently detected object bounding box in the next frame and directly match it with the detection results in the next frame to achieve data association.

[0328] In camera motion application scenarios such as pool cleaning robots, there are many types of targets that the pool cleaning robot needs to clean, and it is difficult to ensure that the pool cleaning robot moves at a constant speed. Therefore, the bounding box prediction result of the pool cleaning robot tracking target using the Kalman filter with constant speed motion assumption is obviously not accurate enough, which can easily affect the accuracy and efficiency of target tracking.

[0329] In response to the above problems, another implementation method provided in the embodiment of the present application is described in more detail with some examples. Referring to FIG18 , the method includes the following steps.

[0330] 1801. When the pool cleaning robot is inspecting the area to be cleaned, the robot controller obtains a first target detection result corresponding to a first environment image and a second target detection result corresponding to a second environment image within a visual detection range of the pool cleaning robot.

[0331] The pool cleaning robot is equipped with a walking unit at its bottom, which can be driven to control the robot to move and inspect the area to be cleaned. The walking unit includes walking wheels or a propeller. Driving the walking unit means driving the walking wheels to rotate or driving the propeller to generate thrust, thereby driving the pool cleaning robot to move. When inspecting the area to be cleaned, the pool cleaning robot also uses an image acquisition device to obtain environmental images within the pool cleaning robot's visual detection range at a preset acquisition frequency, such as, but not limited to, a first environmental image and the next frame of the first environmental image, namely, a second environmental image. The robot then uses an object detection algorithm or an image detection algorithm to perform target detection on the first and second environmental images to detect whether there is a target to be cleaned in the first and second environmental images within the visual detection range of the underwater cleaning robot, and if there is a target to be cleaned, the target's corresponding location, area, detection area, type, and other detection information are detected, thereby obtaining a first target detection result corresponding to the first environmental image and a second target detection result corresponding to the second environmental image.

[0332] Optionally, after acquiring the first environmental image, the pool cleaning robot may first extract environmental features from the first environmental image. Such environmental features may include, but are not limited to, shape, size, color, texture, etc., and then, based on the extracted environmental features, a classifier may be used to distinguish between the target to be cleaned and the background in the first environmental image to perform target detection on the first environmental image, thereby obtaining a first target detection result. The classifier may include, but is not limited to, a support vector machine, a convolutional neural network, etc.

[0333] In some embodiments, the method for obtaining the second target detection result corresponding to the second environment image is consistent with the method for obtaining the first target detection result corresponding to the first environment image, and will not be repeated here.

[0334] Optionally, in order to ensure the accuracy of target detection, before performing target detection on the first environmental image and the second environmental image, the first environmental image and the second environmental image may be preprocessed, such as but not limited to processing operations including noise reduction, contrast enhancement, color adjustment, etc., so that more accurate target detection can be performed subsequently based on the preprocessed first environmental image and the second environmental image.

[0335] In some embodiments, the aforementioned area to be cleaned may be an area corresponding to a swimming pool or pool to be cleaned, or may be any underwater area pre-defined by the user, etc., and this is not limited in the present embodiment. The aforementioned pool cleaning robot includes at least one image acquisition device; the aforementioned visual detection range is determined based on the parameters, installation position, and rotatable angle range of the at least one image acquisition device.

[0336] It should be noted that to improve cleaning efficiency, the pool cleaning robot can be equipped with multiple image acquisition devices. These devices can work collaboratively to expand the underwater robot's visual detection range by splicing or fusing captured environmental images. Specifically, the visual detection range can include the sum of the horizontal and vertical ranges detectable by the multiple image acquisition devices. Furthermore, the visual detection range can also be updated as the operating status of the image acquisition devices changes. For example, if one of the image acquisition devices suddenly malfunctions or becomes damaged during the cleaning process, the pool cleaning robot's visual detection range will also decrease. Alternatively, the pool cleaning robot's visual detection range can be the sum of the detection ranges of the pool cleaning robot's image acquisition device and other sensing devices, such as lidar and ultrasonic sensors. By combining the environmental information acquired by these other sensing devices with the environmental images captured by the image acquisition devices, a more comprehensive environmental perception can be achieved, thereby achieving more efficient and accurate cleaning of the area to be cleaned.

[0337] 1802. If the first target detection result contains first detection information of the target to be cleaned, the robot controller generates target prediction information corresponding to the second environment image based on the first environment image and the actual motion information of the pool cleaning robot.

[0338] Among them, the above-mentioned targets to be cleaned may include but are not limited to garbage, oil stains, stones, nuts, leaves, sand piles, gravel piles, etc. If the first detection information of the target to be cleaned exists in the first target detection result, it can be considered that the target to be cleaned exists in the first environmental image, and then the target prediction information corresponding to the target to be cleaned in the second environmental image can be estimated based on the first environmental image corresponding to the target to be cleaned and the real motion information of the pool cleaning robot, so as to subsequently realize trajectory association. The above-mentioned first detection information may include but is not limited to characteristic information such as the target type, target position, target detection area and target detection area corresponding to the target to be cleaned in the first environmental image. The above-mentioned target detection area may be but is not limited to the minimum circumscribed rectangular area corresponding to the target to be cleaned in the first environmental image. The above-mentioned target position is used to characterize the relative position of the target to be cleaned in the first environmental image and the pool cleaning robot. The above-mentioned target prediction information may include but is not limited to the predicted area corresponding to the target to be cleaned in the first environmental image in the second environmental image and the predicted area and predicted position corresponding to the predicted area.

[0339] Furthermore, the implementation process of generating target prediction information corresponding to the second environmental image based on the first environmental image and the actual motion information of the pool cleaning robot in the above 1802 may include: first using a Kalman filter or a variant of the Kalman filter to generate initial prediction information corresponding to the second environmental image based on the first environmental image, that is, predicting the area position of the target to be cleaned in the first environmental image in the next frame of the environmental image, and then correcting the above initial prediction information based on the actual motion information of the pool cleaning robot to make the prediction process of the Kalman filter more in line with the actual motion condition corresponding to the pool cleaning robot, eliminating the influence of the Kalman filter assumed by the constant speed motion model on the accuracy of the target prediction information, and thus obtaining more accurate target prediction information corresponding to the second environmental image, thereby improving the accuracy of subsequent trajectory association and the target tracking efficiency of the pool cleaning robot.

[0340] In some embodiments, variants of the above-mentioned Kalman filter may include, but are not limited to, an extended Kalman filter, a four-element based Kalman filter, a state error Kalman filter, and the like.

[0341] Optionally, in addition to including an image acquisition device, the pool cleaning robot may also include, but is not limited to, a motion sensor. The motion sensor may include, but is not limited to, an inertial measurement unit, an encoder, a flow sensor, etc., for detecting the motion information of the pool cleaning robot. As shown in FIG19 , before generating target prediction information corresponding to the second environment image based on the first environment image and the actual motion information of the pool cleaning robot in step 1802 , the following may also be included, but is not limited to:

[0342] 1901. The robot controller estimates first motion information of the pool cleaning robot within a target time period through image registration based on the first environment image and the second environment image.

[0343] The target time period is the time period between the first acquisition moment corresponding to the first environmental image and the second acquisition moment corresponding to the second environmental image. After the pool cleaning robot acquires the first and second environmental images, it can first extract a first feature from the first environmental image and a second feature from the second environmental image. Then, based on the first and second features corresponding to corresponding points between the first and second environmental images, it can estimate a transformation matrix between the first and second environmental images, such as, but not limited to, translation, rotation, and scaling. Finally, based on this transformation matrix, it can estimate first motion information of the pool cleaning robot during the target time period between the first and second environmental images. The transformation matrix can include, but is not limited to, a translation component and a rotation and scaling component. The translation component is used to represent the displacement of the pool cleaning robot between the first and second acquisition moments, and the rotation and scaling components are used to represent information about changes in the pool cleaning robot's posture and field of view. The first motion information can include, but is not limited to, first posture information and first displacement information of the pool cleaning robot in the image coordinate system during the target time period.

[0344] 1902. The robot controller obtains second motion information of the pool cleaning robot within a target time period through a motion sensor.

[0345] The second motion information may include, but is not limited to, second posture information and second displacement information of the pool cleaning robot in the world coordinate system within the target time period.

[0346] Optionally, when the motion sensor included in the pool cleaning robot is an inertial measurement unit, the angular velocity and acceleration of the pool cleaning robot in the world coordinate system during the target time period can be first obtained through the inertial measurement unit, and then the second posture information of the pool cleaning robot in the world coordinate system can be obtained by integrating the angular velocity, the speed of the pool cleaning robot can be obtained by integrating the acceleration, and the second displacement information of the pool cleaning robot in the world coordinate system can be obtained by integrating the speed.

[0347] In some embodiments, the pool cleaning robot may include multiple motion sensors. To ensure the accuracy of the second motion information, and thereby improve the accuracy of the pool cleaning robot's true motion information, the second motion information may be obtained by, but is not limited to, weighted summing of multiple motion information acquired by the multiple motion sensors of the pool cleaning robot.

[0348] 1903. The robot controller fuses the first motion information with the second motion information to obtain real motion information corresponding to the pool cleaning robot.

[0349] Among them, after obtaining the first motion information and the second motion information, the second motion information can be first mapped to the image coordinate system, and then weighted fused with the first motion information to obtain more accurate real motion information corresponding to the pool cleaning robot.

[0350] In some embodiments, since the first posture information in the first motion information estimated by image registration is relatively accurate, and the second displacement information in the second motion information obtained by the motion sensor is relatively accurate, in order to obtain more accurate real motion information of the pool cleaning robot, when the first motion information and the second motion information are fused, the fusion weight of the first posture information in the first motion information can be set to be greater than the fusion weight of the second posture information in the second motion information, and the fusion weight of the first displacement information in the first motion information can be set to be less than the fusion weight of the second displacement information in the second motion information.

[0351] Optionally, in order to ensure the efficiency of obtaining the real motion information corresponding to the pool cleaning robot, the first motion information obtained by image registration estimation in the above 1901 or the second motion information obtained by the motion sensor in 1902 can also be directly used as the real motion information corresponding to the pool cleaning robot to participate in the generation of target prediction information corresponding to the second environment image, but is not limited to.

[0352] 1803. The robot controller determines the trajectory correlation between the first environment image and the second environment image based on the second target detection result and the target prediction information according to the trajectory correlation strategy corresponding to the type of the target to be cleaned.

[0353] Among them, the above-mentioned trajectory correlation is used to determine whether the trajectories of the pool cleaning robots corresponding to the first environmental image and the second environmental image respectively represent the motion cleaning trajectories of the pool cleaning robot tracking the same target to be cleaned. If there is also a target to be cleaned in the second environmental image, it means that there may be a certain target tracking correlation between the second environmental image and the first environmental image. Then, the trajectory correlation between the first environmental image and the second environmental image can be determined based on the second target detection result and the target prediction information according to the trajectory correlation strategy corresponding to the type of the target to be cleaned. The target to be cleaned belongs to different types, and the corresponding trajectory correlation strategy is also different. Therefore, a targeted trajectory correlation strategy can be adopted according to the type of the target to be cleaned, thereby improving the operating efficiency and target tracking efficiency of the pool cleaning robot, realizing targeted tracking and cleaning according to the type of the target to be cleaned, and further improving the cleaning effect and inspection efficiency of the pool cleaning robot.

[0354] As shown in FIG20 , the second target detection result includes the detection area corresponding to the target to be cleaned in the second environment image. The target prediction information includes the prediction area corresponding to the target to be cleaned in the second environment image. The above 1803, according to the trajectory association strategy corresponding to the type of the target to be cleaned, the implementation process of determining the trajectory correlation between the first environment image and the second environment image based on the second target detection result and the target prediction information may include: first calculating the first intersection-and-union ratio between the detection area and the prediction area, that is, the ratio of the intersection area to the union area of ​​the detection area and the prediction area; then, according to the type of the target to be cleaned, according to the corresponding trajectory association strategy, determining the trajectory correlation between the first environment image and the second environment image based on the first intersection-and-union ratio.

[0355] In some embodiments, as shown in Figure 21, after obtaining the detection area 520 and the prediction area 530 corresponding to the target to be cleaned in the second environment image 510, it can be calculated that the intersection area corresponding to the intersection area 540 between the detection area 520 and the prediction area 530 is 6, and the union area corresponding to the union area 550 between the detection area 520 and the prediction area 530 is 18, then the first intersection ratio between the detection area 520 and the prediction area 530 can be calculated to be 1 / 3.

[0356] Furthermore, as shown in FIG20 , if the target to be cleaned is a rigid target, in order to reduce the computational complexity of trajectory association and improve the operating efficiency and target tracking efficiency of the pool cleaning robot, the trajectory association between the first environment image and the second environment image can be determined directly based on the first intersection-and-union ratio between the detection area and the prediction area corresponding to the target to be cleaned in the second environment image. For example, but not limited to, the ratio corresponding to the above-mentioned first intersection-and-union ratio is directly determined as the association value corresponding to the above-mentioned trajectory association, or, when the above-mentioned first intersection-and-union ratio is greater than or equal to the association threshold, it is determined that there is trajectory association between the first environment image and the second environment image, and the above-mentioned second environment image The trajectory of the image corresponding to the detection area can be associated with the cleaning trajectory of the pool cleaning robot corresponding to the first environmental image, that is, the trajectory between the first environmental image and the second environmental image is the trajectory generated by the pool cleaning robot tracking the same target to be cleaned; when the first intersection-over-union ratio is less than the correlation threshold, it is determined that there is no trajectory correlation between the first environmental image and the second environmental image, and the trajectory of the detection area corresponding to the second environmental image cannot be associated with the cleaning trajectory of the pool cleaning robot corresponding to the first environmental image, that is, the trajectory between the first environmental image and the second environmental image is the trajectory generated when the pool cleaning robot tracks different targets to be cleaned. The above-mentioned rigid targets are used to represent targets to be cleaned that will not undergo non-rigid deformation, such as but not limited to nuts, stones, etc.

[0357] If the target to be cleared is a non-rigid target, to ensure the accuracy of trajectory association and avoid tracking inaccuracies caused by non-rigid deformation of the target to be cleared during tracking, a feature correlation between the target to be cleared in the first and second environmental images can be determined based on a first feature corresponding to the target to be cleared in the first environmental image and a second feature corresponding to the target to be cleared in the second environmental image. Then, based on the feature correlation and a first intersection-over-union ratio, the trajectory correlation between the first and second environmental images can be determined. For example, but not limited to, the trajectory correlation can be obtained by weighted summing the feature correlation and the first intersection-over-union ratio. The first feature is used to re-identify the identity of the target to be cleared in the first environmental image, such as, but not limited to, the appearance features of the target to be cleared in the first environmental image. The second feature is used to re-identify the identity of the target to be cleared in the second environmental image, such as, but not limited to, the appearance features of the target to be cleared in the second environmental image. The first and second features can be, but not limited to, extracted using a pre-trained object detection model. The non-rigid target is used to characterize targets to be cleared that are subject to non-rigid deformation, such as, but not limited to, leaves, garbage bags, etc.

[0358] As the network depth in the target detection model increases, using a pixel in a high-level feature map to map back to the original image will represent a relatively large area. For example, a convolutional network that has undergone 5 downsamplings with a step size of 2, a 1*1 pixel of its highest-level feature information mapped back to the original image represents a 32*32 area. For small targets, this 32*32 area may be much larger than the actual scale of the smaller target, that is, the features on the high-level feature map cannot accurately describe the features of the small target. Therefore, for small targets, regardless of whether they will undergo rigid deformation, their feature information can be used for trajectory association in the embodiments of the present application.

[0359] That is, if the target to be cleaned is a small target, the trajectory association between the first environment image and the second environment image can be determined based on its first intersection-and-union ratio and the first intersection-and-union ratio coefficient corresponding to the first intersection-and-union ratio, thereby ensuring the accuracy of the trajectory association of the small target. The above-mentioned small target is used to represent the target to be cleaned whose imaging area in the first environment image is smaller than a preset value or whose short side dimension is smaller than a preset size. The above-mentioned first intersection-and-union ratio coefficient is greater than or equal to 1.

[0360] Furthermore, in order to avoid the problem of small targets that originally had trajectory correlation being misjudged due to the shaking of the image acquisition device or the offset of the prediction area during the tracking movement of the pool cleaning robot, and to improve the underwater tracking ability of the pool cleaning robot for small targets to be cleaned, the first intersection-and-union ratio coefficient corresponding to the above-mentioned small targets can be set to be greater than 1, for example but not limited to, it can be directly determined based on the proportional relationship between the predicted area of ​​the prediction area and the detection area of ​​the detection area.

[0361] In some embodiments, the first intersection-over-combination coefficient k may be:

[0362] Among them, spred is the predicted area of ​​the prediction area; sdet is the detected area of ​​the detection area.

[0363] In some embodiments, as shown in FIG. 22 , the implementation process of 1801 of obtaining a first target detection result corresponding to a first environment image and a second target detection result corresponding to a second environment image within the visual detection range of the pool cleaning robot may include: obtaining a first environment image and a second environment image within the visual detection range of the pool cleaning robot through an image acquisition device; inputting the first environment image into a target detection model and outputting the first target detection result corresponding to the first environment image; and inputting the second environment image into the target detection model and outputting the first target detection result corresponding to the second environment image. The target detection model is trained based on a sample training set, which includes a plurality of target sample images with known sample cleaning targets and corresponding target real information.

[0364] Furthermore, the above-mentioned target real information may include, but is not limited to, the type of sample cleaning target in the target sample image, and the real position and real area of ​​the real area corresponding to the sample cleaning target. Before inputting the first environmental image into the target detection model and outputting the first target detection result corresponding to the first environmental image, a sample training set will be obtained first, and the initial detection model will be used to generate a candidate area corresponding to the sample cleaning target based on the target sample image; then, the target judgment parameter corresponding to the target sample image is determined based on the second intersection-union ratio between the real area and the candidate area, and whether the target sample image is a positive target sample image is determined based on the target judgment parameter; if so, it means that the candidate area generated by the initial detection model for the target sample image is close to its corresponding real area, and the detection accuracy is high, then the initial detection model can be trained based on each positive target sample image in the sample training set to obtain a target detection model. The above-mentioned positive target sample image is used to characterize the target sample image in the sample training set whose corresponding target judgment parameter is greater than the target judgment threshold.

[0365] Most current object detection models divide positive and negative samples by comparing the intersection-over-union (IoU) of the candidate regions generated by the network and the corresponding true regions of the target with a pre-set threshold. This method of dividing positive and negative samples is not very friendly to small targets. For example, Figure 23 shows a 25*25 image space, where each small square represents a pixel. The left image shows a relatively small target with a scale of 6*6, and the right image shows a relatively large target with a scale of 18*18. For the small target shown on the left in Figure 23, the intersection area of ​​the true region 711 and the candidate region 712 is 25, and the union area is 47. The corresponding original IoU = 25 / 47 ≈ 0.53. When the candidate region 712 is offset by two pixels to the right and downward, the intersection area of ​​the true region 711 and the candidate region 713 is 9, and the union area is 63. The IoU after offset = 9 / 63 ≈ 0.14. For the relatively large target in the right image of Figure 23, the initial IoU of the ground truth region 721 and the candidate region 722 is 289 / 359 ≈ 0.81. After both ground truth region 721 and candidate region 722 are shifted rightward and downward by two pixels, the IoU of the candidate region 723 after the shift is 225 / 423 ≈ 0.53. Therefore, when the position of the candidate region is slightly jittered, the impact on the positive and negative sample judgment of small targets is much greater than that of relatively large samples. If 0.5 is used as the target judgment threshold for positive and negative samples, in the above example, a slight shift will cause the candidate region of the small target to change from a positive sample to a negative sample, while the relatively large target will remain a positive sample. Therefore, during the training of the target detection model, if the IoU of the candidate region and the ground truth region is directly used as the judgment principle for positive and negative samples, the corresponding target sample images of small targets are more likely to be judged as negative samples, and their features will not be learned during the model training process, which greatly affects the detection performance of the target detection model for relatively small targets, and thus relatively small targets are more likely to be missed during the target detection process.

[0366] Based on this, during the target detection model training process, when the target sample image is divided into positive and negative samples, if the actual area of ​​the target sample image corresponding to the sample cleaning target is smaller than the preset area, the sample cleaning target can be considered to be a small target. The intersection-and-union (IoU) with the IoU coefficient will be used as the basis for dividing whether the target sample image corresponding to the small target is a positive target sample image, thereby enhancing the target detection model's learning of small target features and improving the pool cleaning robot's tracking effect on small targets.

[0367] If the true area of ​​the true region is smaller than the preset area, a second IoU coefficient corresponding to the second IoU can be determined based on the ratio of the true area mean to the true area of ​​the true region; then, the target determination parameter corresponding to the target sample image is determined based on the second IoU and the second IoU coefficient corresponding to the second IoU. The true area mean is used to represent the average true area of ​​the true region corresponding to the sample cleaning target for each target sample image in the sample training set, and the second IoU coefficient is greater than 1.

[0368] In some embodiments, the second IoU coefficient may be, but is not limited to:

[0369] Among them, t is the second intersection-and-union coefficient; is the mean of the true area; is the true area of ​​the true area of ​​the sample cleaning target corresponding to the i-th target sample image in the sample training set; is the second intersection-and-union between the true area corresponding to the i-th target sample image and the candidate area; is the second intersection-and-union coefficient threshold corresponding to the second intersection-and-union, for example but not limited to equal to 2 or 3.

[0370] As can be seen from the above formula, the smaller the actual area of ​​the sample cleanup target, the larger the weight coefficient of the corresponding second intersection-over-union ratio. This alleviates the susceptibility of small targets to frame jitter during the positive-negative sample classification process, making the target sample images corresponding to small targets more easily classified as positive target sample images for target detection model training. At the same time, by using the second intersection-over-union ratio coefficient threshold m to set an upper limit on the area impact factor, this limits the noise interference caused by the features of extremely small but inaccurate candidate regions on the target detection model network, further improving the detection performance of the target detection model.

[0371] Next, please refer to Figure 24, which is another flowchart provided by an embodiment of the present application. As shown in Figure 24, the method includes the following steps:

[0372] 2401. When the pool cleaning robot is inspecting the area to be cleaned, the robot controller obtains a first target detection result corresponding to a first environment image and a second target detection result corresponding to a second environment image within a visual detection range of the pool cleaning robot.

[0373] Among them, 2401 is consistent with 1801 and will not be repeated here.

[0374] 2402. If the first target detection result contains first detection information of the target to be cleaned, the robot controller generates target prediction information corresponding to the second environment image based on the first environment image and the actual motion information of the pool cleaning robot.

[0375] Among them, 2402 is the same as 1802 and will not be repeated here.

[0376] 2403. The robot controller determines the trajectory correlation between the first environment image and the second environment image based on the second target detection result and the target prediction information according to the trajectory correlation strategy corresponding to the type of the target to be cleaned.

[0377] Among them, 2403 is the same as 1803 and will not be repeated here.

[0378] Furthermore, if the first target detection result contains first detection information of the target to be cleared, the method may further include:

[0379] 2404. The robot controller controls the pool cleaning robot to move toward the target to be cleaned based on the first detection information, so as to track and clean the target to be cleaned.

[0380] The first detection information includes the relative position of the target to be cleaned and the pool cleaning robot. The relative position can be determined based on the coordinates, direction, size, and other relative position information of the target to be cleaned in the first environment image. If the first detection information of the target to be cleaned is present in the first target detection result, it can be considered that the target to be cleaned exists in the first environment image. The pool cleaning robot can then be controlled to move to the target position of the target to be cleaned based on the relative position to track and clean the target to be cleaned.

[0381] Optionally, if there are multiple targets to be cleaned in the first environment image, the cleaning priorities corresponding to the multiple targets to be cleaned can be determined based on the feature information corresponding to the multiple targets to be cleaned and the weights corresponding to the feature information. The feature information can include but is not limited to at least one of the following: the type, area, color, relative position of the target to be cleaned and the pool cleaning robot, etc., and then the pool cleaning robot is controlled to move to the target positions of the multiple targets to be cleaned in turn according to the cleaning priorities corresponding to the multiple targets to be cleaned, so as to track and clean the multiple targets to be cleaned.

[0382] Optionally, after the pool cleaning robot tracks and cleans the target, the method may further include the following steps:

[0383] 2405. If the pool cleaning robot has cleaned the target object for a target number of times and the target object is still not cleaned, the robot controller determines that the target object is a stubborn stain.

[0384] If the pool cleaning robot has cleaned the target a target number of times (e.g., but not limited to, 3 or 5 times) and the target remains uncleaned, it indicates that the pool cleaning robot is currently unable to clean the target, and the target can be determined to be a stubborn stain for the pool cleaning robot. The number of times the pool cleaning robot cleans the target can be determined, but is not limited to, based on the number of times the pool cleaning robot passes by the target location of the target.

[0385] Optionally, after determining that the object to be cleaned is a stubborn stain, the method may also include, but is not limited to, one or more of the following steps:

[0386] 2406. The robot controller optimizes the target inspection and cleaning path of the pool cleaning robot in the area to be cleaned based on the target location of the stubborn stains in the area to be cleaned.

[0387] The target inspection and cleaning path is used to represent the initial inspection and cleaning path of the pool cleaning robot when inspecting the area to be cleaned. After determining the stubborn stains in the area to be cleaned, the target inspection and cleaning path of the pool cleaning robot in the area to be cleaned can be promptly optimized based on the target location of the stubborn stains in the area to be cleaned, so that the next time the pool cleaning robot inspects the area to be cleaned, it can bypass the target location of the stubborn stain in advance or directly skip (pass by) the target location of the stubborn stain according to the optimized target inspection and cleaning path, and will not stop at the target location for cleaning, thereby avoiding the pool cleaning robot from ineffectively cleaning the stubborn stains in the area to be cleaned and improving the inspection and cleaning efficiency of the pool cleaning robot in the area to be cleaned.

[0388] 2407. The robot controller optimizes the target detection model corresponding to the pool cleaning robot based on the first environment image corresponding to the stubborn stain.

[0389] Among them, after determining that the target to be cleaned is a stubborn stain, it is also possible but not limited to updating and optimizing the target detection model corresponding to the pool cleaning robot based on the first environmental image corresponding to the stubborn stain, so that the pool cleaning robot has the ability to detect stubborn stains, further improving the inspection and cleaning effect of the pool cleaning robot.

[0390] 2408. During the mapping or positioning process of the pool cleaning robot in the area to be cleaned, the robot controller performs loop detection based on the target position of stubborn stains to eliminate the accumulated error of the pool cleaning robot's odometer.

[0391] Among them, after determining the stubborn stains existing in the area to be cleaned, the pool cleaning robot can also, but is not limited to, perform loop detection based on the target position of the stubborn stains during the mapping or positioning process of the area to be cleaned, so as to eliminate the cumulative error of the pool cleaning robot's odometer. That is, when the pool cleaning robot moves to the target position where the stubborn stain is located again, the cumulative error of the pool cleaning robot's odometer is cleared to zero, thereby ensuring the consistency and closure of the map after mapping the area to be cleaned, and improving the mapping and positioning effect of the area to be cleaned.

[0392] Next, please refer to Figure 25, which is another flow chart provided in an embodiment of the present application. As shown in Figure 25, the method includes the following steps:

[0393] 2501. When the pool cleaning robot is inspecting the area to be cleaned, the robot controller obtains a first target detection result corresponding to a first environment image and a second target detection result corresponding to a second environment image within a visual detection range of the pool cleaning robot.

[0394] Among them, 2501 is consistent with 1801 and will not be repeated here.

[0395] 2502. If the first target detection result contains first detection information of the target to be cleaned, the robot controller generates target prediction information corresponding to the second environment image based on the first environment image and the actual motion information of the pool cleaning robot.

[0396] Among them, 2502 is consistent with 1802 and will not be repeated here.

[0397] 2503. The robot controller determines the trajectory correlation between the first environment image and the second environment image based on the second target detection result and the target prediction information according to the trajectory correlation strategy corresponding to the type of the target to be cleaned.

[0398] Among them, 2503 is the same as 1803 and will not be repeated here.

[0399] 2504. The robot controller determines a target cleaning trajectory corresponding to the target to be cleaned based on the trajectory correlation.

[0400] Among them, after determining the trajectory correlation between the first environmental image and the second environmental image, if the trajectory correlation is greater than or equal to the correlation threshold, that is, the trajectories between the first environmental image and the second environmental image are correlated, then the trajectories of the corresponding pool cleaning robots between the first environmental image and the second environmental image can be connected together as the target cleaning trajectory corresponding to the target to be cleaned.

[0401] Optionally, if the trajectory correlation is less than the correlation threshold, that is, the trajectories between the first environmental image and the second environmental image are not correlated, the trajectory of the pool cleaning robot corresponding to the first environmental image can be used as the target cleaning trajectory corresponding to the target to be cleaned, that is, the target position corresponding to the target to be cleaned in the first environmental image is the end point of its corresponding target cleaning trajectory, and the trajectory of the pool cleaning robot corresponding to the second environmental image is used as the new target cleaning trajectory or inspection trajectory.

[0402] 2505. The robot controller performs path planning based on the target cleaning trajectory to obtain the cleaning path to be inspected corresponding to the pool cleaning robot.

[0403] Among them, after determining the target cleaning trajectory corresponding to the target to be cleaned currently tracked by the pool cleaning robot in the first environmental image, the subsequent path of the pool cleaning robot can also be planned based on the target cleaning trajectory to obtain the cleaning path to be inspected corresponding to the pool cleaning robot, thereby avoiding the pool cleaning robot from repeatedly inspecting the same location subsequently, and improving the inspection efficiency of the pool cleaning robot in the area to be cleaned.

[0404] Optionally, after the above 2504 determines the target cleaning trajectory corresponding to the target to be cleaned based on the trajectory correlation, the above 2505 performs path planning based on the target cleaning trajectory to obtain the cleaning path to be inspected corresponding to the pool cleaning robot. When the computing power of the robot controller of the pool cleaning robot is greater than the preset computing power, the target cleaning trajectory can be optimized by using an interpolation algorithm, that is, the trajectory position of the missing intermediate frame is predicted based on the trajectory position of the previous and next frames in the target cleaning trajectory, and then path planning is performed based on the optimized target cleaning trajectory and the inspected cleaning path of the pool cleaning robot to obtain the cleaning path to be inspected corresponding to the pool cleaning robot, thereby avoiding the problem that one or two frames in the target cleaning trajectory are blocked by other things due to different postures of the pool cleaning robot to the target to be cleaned, and there is no corresponding detection information in the environmental image, which causes the trajectory of the frame to be interrupted, thereby improving the continuity of the target cleaning trajectory, and thus ensuring the inspection efficiency and inspection effect of the pool cleaning robot in the area to be cleaned.

[0405] 2506. The robot controller controls the pool cleaning robot to continue patrolling and cleaning the area to be cleaned based on the patrolled and cleaned path.

[0406] Among them, after obtaining the cleaning path to be inspected corresponding to the pool cleaning robot, the pool cleaning robot can be further controlled to continue patrol cleaning in the area to be cleaned based on the cleaning path to be inspected, thereby ensuring that the pool cleaning robot can efficiently complete the patrol cleaning of the entire area to be cleaned.

[0407] In some embodiments, after determining the trajectory correlation between the first environment image and the second environment image based on the second target detection result and the target prediction information according to the trajectory correlation strategy corresponding to the type of the target to be cleaned, the method further includes:

[0408] If the trajectory correlation is less than the correlation threshold and the target to be cleaned belongs to a specified type (for example, but not limited to a small stone, a leaf, etc.), when the trajectory correlation between the N frames of environmental images corresponding to the specified type after the first environmental image and the first environmental image is less than the correlation threshold, that is, N frames of trajectories corresponding to the specified type appear repeatedly but cannot be associated with the previous trajectory, it can be considered that a new target to be cleaned of the same type as the specified type has appeared, such as a new stone, leaf, etc., and the cleaning trajectory corresponding to the N frames of environmental images can be initialized as the new cleaning trajectory corresponding to the pool cleaning robot; the above N is a positive integer greater than 1, for example, but not limited to 10, 8, etc. The target to be cleaned of the specified type is used to represent the target to be cleaned that the pool cleaning robot can complete cleaning in one go.

[0409] For non-specified objects (such as, but not limited to, pebble piles, sand piles, etc.), the pool cleaning robot will clean the surface of the object during the cleaning process. Due to the limitations of the pool cleaning robot's width and effective cleaning area, it may not be able to complete the cleaning of the entire pebble pile or sand pile in one go. After cleaning, a large non-specified object may become multiple small objects. For example, a large sand pile may become two or more small sand piles, or a large pebble pile may become two or more small pebble piles. For this case, if the trajectory correlation is less than the correlation threshold and the target to be cleaned belongs to a non-specified type, that is, the second environmental image after the first environmental image of the pool cleaning robot can no longer be associated with the first environmental image for trajectory, it can be considered that the pool cleaning robot has cleaned the target to be cleaned in the first environmental image, and there is no need to perform trajectory correlation calculation on the environmental image after the second environmental image and the first environmental image. When the target to be cleaned appears in the third environmental image, for example, when it is subsequently detected that the non-specified type of target to be cleaned in the first environmental image corresponds to a small target to be cleaned that has not been cleaned, it can be treated as a new target to be cleaned for target tracking and cleaning, that is, the new cleaning trajectory corresponding to the pool cleaning robot can be determined directly based on the cleaning position corresponding to the third environmental image. The above-mentioned third environmental image is the environmental image collected after the first environmental image. Targets to be cleaned that belong to non-specified types are used to characterize targets to be cleaned that the pool cleaning robot cannot complete cleaning in one go.

[0410] Figure 26 is a structural diagram of a visual information application device provided in an embodiment of the present application. Referring to Figure 26, the device includes: a visual information acquisition module 2601 and an application module 2602.

[0411] The visual information acquisition module 2601 is used to acquire visual information collected by the pool cleaning robot;

[0412] The application module 2602 is used to control the pool cleaning robot based on the visual information, or to process the visual information in a preset manner.

[0413] In some embodiments, referring to FIG. 27 , the visual information acquisition module 2601 includes an environment image acquisition unit 2701 , and the application module 2602 includes a detection and segmentation unit 2702 and a control unit 2703 .

[0414] The environmental image acquisition unit 2701 is used to acquire an environmental image around the pool cleaning robot when the pool cleaning robot performs a cleaning task in the pool.

[0415] The detection and segmentation unit 2702 is used to input the environmental image into the image recognition model, perform target detection and semantic segmentation on the environmental image through the image recognition model, and obtain the first recognition frame, the second recognition frame and the position type of the position of the pool cleaning robot on the environmental image. The first recognition frame is used to indicate the position of the target to be cleaned on the environmental image, and the second recognition frame is used to indicate the position of the obstacle on the environmental image.

[0416] The control unit 2703 is used to control the pool cleaning robot to clean the target based on the environment image, the first identification frame, the second identification frame and the location type of the pool cleaning robot.

[0417] In some embodiments, the detection and segmentation unit 2702 is configured to extract features from the environment image using the image recognition model to obtain environment image features of the environment image. Based on the environment image features, the image recognition model performs bounding box detection on the environment image to obtain the first recognition frame and the second recognition frame. Based on the environment image features, the image recognition model classifies multiple pixels in the environment image to obtain a location type of the pool cleaning robot.

[0418] In some embodiments, the detection and segmentation unit 2702 is used to control the first candidate recognition frame and the second candidate recognition frame to slide on the environmental image, where the first candidate recognition frame is used to identify the target to be cleaned, and the second candidate recognition frame is used to identify the obstacle. The first recognition frame and the second recognition frame are determined based on the first sub-image features corresponding to multiple first image areas in the environmental image features and the second sub-image features corresponding to multiple second image areas in the environmental image features. The first image area is the image area covered by the first candidate recognition frame on the environmental image, and the second image area is the image area covered by the second candidate recognition frame on the environmental image.

[0419] In some embodiments, the detection and segmentation unit 2702 is configured to perform full connection and normalization on the first sub-image features corresponding to any first image region among the multiple first image regions to determine whether the first image region contains the target to be cleaned. If the first image region contains the target to be cleaned, the border of the first image region is determined as the first reference recognition frame. The first recognition frame is determined based on the multiple first reference recognition frames on the environmental image. For any second image region among the multiple second image regions, the second sub-image features corresponding to the second image region are fully connected and normalized to determine whether the second image region contains an obstacle. If the second image region contains an obstacle, the border of the second image region is determined as the second reference recognition frame. The second recognition frame is determined based on the multiple second reference recognition frames on the environmental image.

[0420] In some embodiments, the detection and segmentation unit 2702 is used to perform multiple upsampling of the environmental image features to obtain an upsampled feature map of the environmental image. The size of the upsampled feature map is the same as that of the environmental image. The upsampled feature map includes multiple channels, and the multiple channels correspond to multiple candidate position types. The pixel value of the pixel point in each channel represents the confidence that the corresponding pixel point in the environmental image is at the candidate position corresponding to the channel; based on the upsampled feature map, the position type corresponding to each pixel point in the environmental image is determined; based on the position types corresponding to the multiple pixel points, the position type of the pool cleaning robot is determined.

[0421] In some embodiments, the detection and segmentation unit 2702 is configured to determine a location type ahead of the pool cleaning robot in a travel direction based on the location types corresponding to the plurality of pixel points, and to determine a location type of the pool cleaning robot based on the location type ahead of the pool cleaning robot in a travel direction.

[0422] In some embodiments, the control unit 2703 is configured to determine a cleaning mode corresponding to the location type of the pool cleaning robot. A path is planned for the pool cleaning robot based on the environmental image, the first identification frame, the second identification frame, and the cleaning mode to obtain a target movement trajectory. The pool cleaning robot is controlled to move along the target movement trajectory in the cleaning mode to clean the target in the pool.

[0423] In some embodiments, the control unit 2703 is configured to determine a preset cleaning trajectory corresponding to the cleaning mode. A reference movement trajectory is determined based on the positions of the first and second recognition frames in the environment image, the obstacle type of the obstacle in the second recognition frame, and the size of the second recognition frame. The preset cleaning trajectory and the reference movement trajectory are combined to obtain the target movement trajectory.

[0424] In some embodiments, the control unit 2703 is configured to determine whether the pool cleaning robot needs to avoid the obstacle based on the obstacle type, the position of the second identification frame in the environment image, and the size of the second identification frame. If the pool cleaning robot needs to avoid the obstacle, the control unit 2703 determines a first movement direction and a first movement distance of the pool cleaning robot based on the relative positional relationship between the first and second identification frames and the center point of the environment image. The reference movement trajectory is generated based on the first movement direction and the first movement distance of the pool cleaning robot.

[0425] In some embodiments, the control unit 2703 is further configured to determine a second movement direction and a second movement distance of the pool cleaning robot based on the relative positional relationship between the first recognition frame and the center point of the environment image when the pool cleaning robot does not need to avoid the obstacle, and generate the reference movement trajectory based on the second movement direction and the second movement distance of the pool cleaning robot.

[0426] In some embodiments, the control unit 2703 is further configured to determine the obstacle climbing level of the pool cleaning robot, where the obstacle climbing level is positively correlated with the obstacle climbing capability.

[0427] The control unit 2703 is configured to segment the obstacle in the second recognition frame and obtain a contour of the obstacle in the environmental image when the obstacle climbing level of the pool cleaning robot is less than or equal to a preset level. The control unit 2703 is configured to perform path planning for the pool cleaning robot based on the environmental image, the first recognition frame, the contour of the obstacle, and the cleaning mode to obtain a target movement trajectory.

[0428] In some embodiments, the control unit 2703 is configured to, when the pool cleaning robot is located at the pool wall, determine the cleaning mode of the pool cleaning robot to be a roller brush cleaning mode. When the pool cleaning robot is located at the bottom of the pool, determine the cleaning mode of the pool cleaning robot to be a water pump extraction mode. When the pool cleaning robot is located at the dividing line between the pool bottom and the pool wall, determine the cleaning mode of the pool cleaning robot to be a side cleaning mode. When the pool cleaning robot is located on the steps of the pool, determine the cleaning mode of the pool cleaning robot to be a basic cleaning mode or a deep cleaning mode, wherein the basic cleaning mode cleans the step plane but does not clean the step facade, and the deep cleaning mode cleans the step facade and the step plane.

[0429] In some embodiments, referring to FIG. 28 , the visual information acquisition module 2601 includes a position acquisition unit 2801 , and the application module 2602 includes a cleaning path generation unit 2802 and a control unit 2803 .

[0430] The position acquisition unit 2801 is used to acquire the position of the pool cleaning robot and the position of target garbage to be cleaned around the pool cleaning robot.

[0431] The cleaning path generation unit 2802 is used to generate a target cleaning path for the pool cleaning robot based on the position of the pool cleaning robot, the position of the target garbage, and the target garbage type of the target garbage. The target cleaning path is the cleaning path used when cleaning the target garbage, and the target garbage type is used to represent the shape of the target garbage.

[0432] The control unit 2803 is used to control the pool cleaning robot to clean the target garbage based on the target cleaning path.

[0433] In one possible embodiment, the target cleaning path is a first cleaning path, a second cleaning path, or a third cleaning path. The cleaning path generation unit 2802 is configured to, when the target garbage type is a point type, determine the center position of the center point of the target garbage based on the position of the target garbage. Path planning is performed between the position of the pool cleaning robot and the center position to obtain the first cleaning path. When the target garbage type is a line segment type, the line segment position of the line segment corresponding to the target garbage is determined based on the position of the target garbage. The second cleaning path is generated based on the position of the pool cleaning robot and the line segment position. When the target garbage type is a surface type, the area where the target garbage is located is determined based on the position of the target garbage. The third cleaning path is generated based on the position of the pool cleaning robot and the area where the target garbage is located.

[0434] In one possible embodiment, the cleaning path generation unit 2802 is configured to determine the starting point of the line segment corresponding to the target waste from the line segment position. Path planning is performed between the pool cleaning robot's position and the starting point to obtain a first initial cleaning path. A second initial cleaning path of the line segment corresponding to the target waste is added to the first initial cleaning path to obtain a second cleaning path, where the second initial cleaning path passes through the starting point and end point of the line segment corresponding to the target waste.

[0435] In one possible embodiment, the cleaning path generation unit 2802 is configured to determine the vertex position of any vertex of the target waste within the area where the target waste is located. Path planning is performed between the pool cleaning robot's position and the vertex position to obtain a third initial cleaning path. Path planning is performed based on the vertex position and the area where the target waste is located, to obtain a fourth initial cleaning path that covers the area where the target waste is located. The third initial cleaning path and the fourth initial cleaning path are combined to obtain the third cleaning path.

[0436] In one possible implementation, the application module 2602 further includes a target garbage determination unit configured to determine the garbage type of multiple candidate garbage items surrounding the pool cleaning robot. Based on the pool cleaning robot's location, the garbage type, location, and garbage area of ​​each candidate garbage item, a cleaning cost between the pool cleaning robot and each candidate garbage item is determined. Based on the cleaning costs between the pool cleaning robot and each candidate garbage item, a target garbage item is determined from the multiple candidate garbage items, with the target garbage item being the candidate garbage item with the lowest cleaning cost.

[0437] In one possible embodiment, the target garbage determination unit is configured to determine an outline of each candidate garbage item, determine a bounding rectangle of each candidate garbage item based on the outline of each candidate garbage item, and determine the garbage type of each candidate garbage item based on the size information of the bounding rectangle of each candidate garbage item.

[0438] In one possible embodiment, the target garbage identification unit is configured to obtain an image of the environment surrounding the pool cleaning robot, perform a perspective transformation on the image to obtain a BEV image corresponding to the image, and perform target detection on the BEV image to obtain the outlines of each candidate garbage item.

[0439] In a possible embodiment, the target garbage determination unit is configured to control the pool cleaning robot to rotate and acquire multiple initial environment images during the rotation process, and to splice the multiple initial environment images to obtain an environment image surrounding the pool cleaning robot.

[0440] In one possible embodiment, the size information includes length and width. The target waste determination unit is configured to, for any of the multiple candidate waste items, determine the target waste type of the target waste item as a point type if both the length and width of the bounding rectangle of the candidate waste item are less than or equal to a preset threshold, where the preset threshold is associated with the single cleaning width of the pool cleaning robot. If the length of the bounding rectangle of the candidate waste item is less than or equal to the preset threshold and the width is greater than the preset threshold, or if the width of the bounding rectangle of the candidate waste item is less than or equal to the preset threshold and the length is greater than the preset threshold, determine the target waste type of the target waste item as a line segment type. If both the length and width of the bounding rectangle of the candidate waste item are greater than the preset threshold, determine the target waste type of the target waste item as a surface type.

[0441] In a possible embodiment, the target garbage determination unit is further configured to determine the garbage area of ​​each candidate garbage by the area of ​​the circumscribed rectangle of each candidate garbage, or to determine the garbage area of ​​each candidate garbage by the area enclosed by the outline of each candidate garbage.

[0442] In one possible embodiment, the target garbage determination unit is configured to determine the distance between the pool cleaning robot and each candidate garbage item based on the location of the pool cleaning robot and the garbage type and location of each candidate garbage item. Based on the distance between the pool cleaning robot and each candidate garbage item and the garbage area of ​​each candidate garbage item, the cleaning cost between the pool cleaning robot and each candidate garbage item is determined.

[0443] In one possible embodiment, the target garbage determination unit is configured to, for any candidate garbage among the multiple candidate garbage items, determine the center position of the center point of the candidate garbage item based on the location of the candidate garbage item when the candidate garbage item is of point type. The distance between the pool cleaning robot and the center position is determined as the distance between the pool cleaning robot and the candidate garbage item. If the candidate garbage item is of line segment type, the midpoint position of the midpoints of the two short sides of the circumscribed rectangle of the candidate garbage item is determined based on the location of the candidate garbage item. The distance between the pool cleaning robot and the candidate garbage item is determined based on the location of the pool cleaning robot and the midpoint position of the two short sides. If the candidate garbage item is of surface type, the vertex positions of the four vertices of the circumscribed rectangle of the candidate garbage item are determined based on the location of the candidate garbage item. The distance between the pool cleaning robot and the candidate garbage item is determined based on the location of the pool cleaning robot and the vertex positions of the four vertices.

[0444] In one possible embodiment, the two short side midpoints include a first short side midpoint and a second short side midpoint, and the target garbage identification unit is configured to determine a first short side reference distance between the position of the pool cleaning robot and the midpoint of the first short side midpoint. Determine a second short side reference distance between the position of the pool cleaning robot and the midpoint of the second short side midpoint. The shorter of the first short side reference distance and the second short side reference distance is determined as the distance between the pool cleaning robot and the candidate garbage.

[0445] In one possible embodiment, the four vertices include a first vertex, a second vertex, a third vertex, and a fourth vertex, and the target garbage determination unit is configured to determine a first vertex reference distance between the position of the pool cleaning robot and the vertex position of the first vertex. Determine a second vertex reference distance between the position of the pool cleaning robot and the vertex position of the second vertex. Determine a third vertex reference distance between the position of the pool cleaning robot and the vertex position of the third vertex. Determine a fourth vertex reference distance between the position of the pool cleaning robot and the vertex position of the fourth vertex. The shorter of the first vertex reference distance, the second vertex reference distance, the third vertex reference distance, and the fourth vertex reference distance is determined as the distance between the pool cleaning robot and the candidate garbage.

[0446] In one possible implementation, the target garbage determination unit is configured to multiply the distance between the pool cleaning robot and each candidate garbage item by a distance cost weight, thereby obtaining a distance cost between the pool cleaning robot and each candidate garbage item. The area cost weight is divided by the garbage area of ​​each candidate garbage item to obtain an area cost for each candidate garbage item. The distance cost and area cost corresponding to each candidate garbage item are added together to obtain a cleaning cost between the pool cleaning robot and each candidate garbage item.

[0447] In one possible embodiment, the control unit 2803 is further configured to determine whether candidate waste to be cleaned still exists after the target cleaning path has been completed. If candidate waste to be cleaned still exists, the control unit 2803 is configured to redefine target waste from the candidate waste to be cleaned and to redefine the target cleaning path. Based on the redetermined target cleaning path, the pool cleaning robot is controlled to clean the redetermined target waste.

[0448] In one possible embodiment, the control unit 2803 is further configured to, when no candidate garbage to be cleaned exists, perform path planning based on the location of the pool cleaning robot and the location of the base station to obtain a target recharging path, and control the pool cleaning robot to dock with the base station along the target recharging path.

[0449] In some embodiments, referring to FIG. 29 , the visual information acquisition module 2601 includes an image acquisition unit 2901 , and the application module 2602 includes a feature extraction unit 2902 , a feature matching unit 2903 , a first determination unit 2904 , and a mapping unit 2905 .

[0450] The image acquisition unit 2901 is used to acquire the environment image collected by the underwater cleaning robot.

[0451] The feature extraction unit 2902 is used to perform feature extraction to obtain target structure information in the target space where the underwater cleaning robot is located.

[0452] The feature matching unit 2903 is configured to perform feature matching on the target structure information in the target space based on two adjacent frames of the environment image to obtain a corresponding target feature matching result.

[0453] A first determining unit 2904 is configured to determine target depth information of target structure boundary points in the target structure information within the target space based on the target feature matching result and the target acquisition position information corresponding to the underwater cleaning robot. The target acquisition position information is determined based on positioning information acquired by various sensors of the underwater cleaning robot. The target structure boundary points are matching structure boundary points in the target structure information corresponding to two adjacent frames of the environment image.

[0454] The mapping unit 2905 is configured to construct a target environment map of the target space based on the target structure information and target depth information corresponding to the boundary points of the target structure.

[0455] In one possible implementation, the feature extraction unit 2902 includes:

[0456] The extraction unit is used to extract structural boundary information from the environment image collected by the underwater cleaning robot. The structural boundary information includes first position information corresponding to the structural boundary existing in the environment image and second position information of the area including the contour corresponding to the structural boundary.

[0457] The first determining unit is configured to determine target structure information in the target space based on the structure boundary information.

[0458] In a possible implementation, the feature extraction unit 2902 is configured to:

[0459] Using the multi-task model, the target space component element information and the target space auxiliary element information in the environmental image collected by the underwater cleaning robot are segmented and identified respectively, and the target structure information in the target space where the underwater cleaning robot is located is obtained.

[0460] In a possible implementation, the feature matching unit 2903 includes:

[0461] The acquisition position prediction unit is used to predict the predicted acquisition position information of the next frame of the environmental image between two adjacent frames corresponding to the underwater cleaning robot based on the previous frame acquisition position information of the environmental image between the two adjacent frames corresponding to the underwater cleaning robot.

[0462] The structure information prediction unit is used to predict the predicted structure information in the next frame of the environment image based on the predicted acquisition position information and the target structure information in the previous frame of the environment image.

[0463] The structure information matching unit is used to match the target structure information in the next frame of the environment image with the predicted structure information to obtain the predicted structure matching range corresponding to the target structure information in the next frame of the environment image.

[0464] The second determining unit is used to determine the matching target structure boundary points in the target structure information corresponding to two adjacent frames of the environment image based on the geometric features of each structure boundary point in the previous frame of the environment image and the geometric features of each structure boundary point within the predicted structure matching range.

[0465] In one possible implementation, the multiple sensors include a first sensor, a second sensor, and an ultrasonic sensor. The first sensor is used to capture images of the underwater cleaning robot's environment, and the second sensor is used to capture motion data of the underwater cleaning robot. The second sensor includes at least one of the following: an inertial measurement unit, an encoder, and a tachometer. The ultrasonic sensor is used to capture ultrasonic positioning data of the underwater cleaning robot.

[0466] The application module 2602 also includes:

[0467] An acquisition unit is used to acquire two adjacent frames of environmental images based on the first sensor, acquire actual motion data of the underwater cleaning robot based on one or more of the second sensors, and acquire ultrasonic positioning data of the underwater cleaning robot based on the ultrasonic sensor.

[0468] The second determination unit is used to determine the first acquisition position information corresponding to the latter frame of the two adjacent frames of environmental images based on the two adjacent frames of environmental images and the previous frame acquisition position information corresponding to the previous frame of the two adjacent frames of environmental images, determine the second acquisition position information corresponding to the latter frame of environmental image based on the actual motion data, and determine the third acquisition position information corresponding to the latter frame of environmental image based on the ultrasonic positioning data.

[0469] The third determining unit is configured to determine target acquisition position information corresponding to the next frame of the environment image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information.

[0470] In a possible implementation, the application module 2602 further includes:

[0471] The fourth determining unit is configured to determine the position loss corresponding to the subsequent frame of the environmental image based on the predicted acquisition position information corresponding to the subsequent frame of the environmental image, the first acquisition position information, the second acquisition position information, and the third acquisition position information.

[0472] The fifth determining unit is configured to determine the image loss corresponding to the subsequent frame of the environmental image based on the predicted structure information and the target structure information corresponding to the subsequent frame of the environmental image.

[0473] The correction unit is used to correct the position drift of the multiple sensors based on the position loss and image loss corresponding to the subsequent frame of the environmental image.

[0474] In a possible implementation, the application module 2602 further includes:

[0475] The judgment unit is configured to judge whether the ultrasonic positioning data satisfies a non-line-of-sight condition based on the next frame of the environmental image. The non-line-of-sight condition is used to indicate that there is occlusion in the next frame of the environmental image and the ultrasonic positioning data is abnormally jittery.

[0476] A sixth determining unit is configured to determine, if the ultrasonic positioning data satisfies the non-line-of-sight condition, target acquisition position information corresponding to the subsequent frame of the environment image based on the first acquisition position information and the second acquisition position information.

[0477] The third determining unit is configured to execute the step of determining the target acquisition position information corresponding to the subsequent frame of the environment image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information if the ultrasonic positioning data does not meet the non-line-of-sight condition.

[0478] In a possible implementation, the application module 2602 further includes:

[0479] A counting unit is configured to count the number of target structure boundary points in the subsequent frame of the environmental image that meet a preset constraint condition, and the total number of structure boundary points in the target structure information corresponding to the subsequent frame of the environmental image. The preset constraint condition is that matching target structure boundary points exist in at least two consecutive frames of the environmental image.

[0480] The seventh determining unit is configured to determine a visual constraint weight corresponding to the first sensor based on a ratio between the number of target structure boundary points that meet the preset constraint condition and the total number of structure boundary points.

[0481] An eighth determining unit is configured to determine weights corresponding to the second sensor and the ultrasonic sensor based on covariance matrices corresponding to the second sensor and the ultrasonic sensor, respectively.

[0482] A ninth determining unit is configured to determine a target weight ratio corresponding to the first sensor, the second sensor, and the ultrasonic sensor based on the visual constraint weight corresponding to the first sensor and the weights corresponding to the second sensor and the ultrasonic sensor.

[0483] The third determining unit is configured to:

[0484] The first acquisition position information, the second acquisition position information, and the third acquisition position information are fused according to a target weight ratio to obtain target acquisition position information corresponding to the next frame of the environmental image.

[0485] In a possible implementation, the target feature matching result includes matching target structure boundary points in two adjacent frames of the environment image.

[0486] The first determining unit 2904 is configured to:

[0487] The target depth information of the target structure boundary point in the target space is estimated based on the position of the target structure boundary point in two adjacent frames of the environment image and the target acquisition position information corresponding to the underwater cleaning robot.

[0488] In a possible implementation, the application module 2602 further includes:

[0489] The mobile unit is used to move along the target mapping path.

[0490] The image acquisition unit is used to acquire environmental images within a preset visual range at a preset frequency during movement.

[0491] The mapping unit 2905 is used to:

[0492] After the target mapping path is completed, a target environment map of the target space is constructed based on the target structure information corresponding to each of the multiple environment images on the target mapping path and the target depth information corresponding to the boundary points of the target structure in the multiple environment images.

[0493] In one possible implementation, the target structure information includes target space component element information and target space auxiliary element information in the target space. The target space component element information includes structural boundary information of the target space corresponding component element, where the component element includes at least one of the following: a bottom surface, a wall surface, a step plane, or a step elevation. The target space auxiliary element information includes structural boundary information of the target space corresponding auxiliary element, where the auxiliary element includes at least one of the following auxiliary devices: a wall lamp, a floor lamp, a water inlet, or a water outlet.

[0494] In some embodiments, referring to FIG. 30 , the visual information acquisition module 2601 includes a first acquisition unit 3001 , and the application module 2602 includes a target prediction unit 3002 and a trajectory association unit 3003 .

[0495] The first acquisition unit 3001 is configured to acquire, when the underwater cleaning robot is inspecting the area to be cleaned, a first target detection result corresponding to a first environment image and a second target detection result corresponding to a second environment image within a visual detection range of the underwater cleaning robot; the second environment image is a frame image subsequent to the first environment image;

[0496] The target prediction unit 3002 is configured to generate target prediction information corresponding to the second environment image based on the first environment image and the real motion information of the underwater cleaning robot if the first detection information of the target to be cleaned exists in the first target detection result;

[0497] The trajectory association unit 3003 is configured to determine the trajectory association between the first environment image and the second environment image based on the second target detection result and the target prediction information according to the trajectory association strategy corresponding to the type of the target to be cleared.

[0498] In one possible implementation, the underwater cleaning robot includes a motion sensor;

[0499] The application module 2602 also includes:

[0500] An image registration estimation unit is configured to estimate first motion information of the underwater cleaning robot within a target time period through image registration based on the first environment image and the second environment image; the target time period is a time period between a first acquisition moment corresponding to the first environment image and a second acquisition moment corresponding to the second environment image;

[0501] a second acquiring unit, configured to acquire, through the motion sensor, second motion information of the underwater cleaning robot within the target time period;

[0502] The fusion unit is used to fuse the first motion information with the second motion information to obtain real motion information corresponding to the underwater cleaning robot.

[0503] In one possible implementation, the target prediction unit 3002 includes:

[0504] A Kalman filter unit, configured to generate initial prediction information corresponding to the second environment image based on the first environment image using a Kalman filter or a variant of the Kalman filter;

[0505] The correction unit is used to correct the initial prediction information based on the real motion information of the underwater cleaning robot to obtain the target prediction information corresponding to the second environment image.

[0506] In a possible implementation, the second target detection result includes a detection area corresponding to the target to be cleared in the second environment image; the target prediction information includes a prediction area corresponding to the target to be cleared in the second environment image;

[0507] The trajectory association unit 3003 includes:

[0508] a calculation unit, configured to calculate a first intersection-over-union ratio between the detection area and the prediction area;

[0509] a first determining unit, configured to determine, if the target to be cleared is a rigid target, a trajectory correlation between the first environment image and the second environment image based on the first intersection-over-union ratio;

[0510] a second determining unit configured to, if the target to be cleaned is a non-rigid target, determine a feature correlation between the target to be cleaned in the first environment image and the second environment image based on a first feature corresponding to the target to be cleaned in the first environment image and a second feature corresponding to the target to be cleaned in the second environment image, and determine a trajectory correlation between the first environment image and the second environment image based on the feature correlation and the first intersection-and-union ratio;

[0511] The third determination unit is used to determine the trajectory correlation between the first environmental image and the second environmental image based on the first intersection-and-union ratio and the first intersection-and-union ratio coefficient corresponding to the first intersection-and-union ratio if the target to be cleaned is a small target; the small target is used to represent the target to be cleaned whose imaging area in the first environmental image is smaller than a preset value or the short side size is smaller than a preset size, and the first intersection-and-union ratio coefficient is greater than or equal to 1.

[0512] In a possible implementation, the application module 2602 further includes:

[0513] The first determining unit is configured to determine a first intersection-and-union coefficient corresponding to the first intersection-and-union ratio based on a proportional relationship between a predicted area of ​​the prediction area and a detected area of ​​the detection area if the target to be cleared is a small target.

[0514] In one possible implementation, the underwater cleaning robot includes an image acquisition device;

[0515] The first acquisition unit 3001 includes:

[0516] An acquisition unit, configured to acquire a first environment image and a second environment image within a visual detection range of the underwater cleaning robot through the image acquisition device;

[0517] A first target detection unit is configured to input the first environment image into a target detection model and output a first target detection result corresponding to the first environment image; the target detection model is trained based on a sample training set; the sample training set includes a plurality of target sample images with known sample cleanup targets corresponding to real target information;

[0518] The second target detection unit is used to input the second environment image into the target detection model and output the first target detection result corresponding to the second environment image.

[0519] In a possible implementation, the target real information includes the type of the sample cleaning target in the target sample image, and the real position and real area of ​​the real area corresponding to the sample cleaning target;

[0520] The application module 2602 also includes:

[0521] A third acquisition unit is used to acquire the sample training set;

[0522] A candidate region generating unit, configured to generate a candidate region corresponding to the sample cleaning target based on the target sample image using an initial detection model;

[0523] a second determining unit, configured to determine a target determination parameter corresponding to the target sample image based on a second intersection-over-union ratio between the real area and the candidate area;

[0524] A judging unit, configured to judge whether the target sample image is a positive target sample image based on the target judgment parameter; the positive target sample image is used to represent a target sample image in the sample training set whose corresponding target judgment parameter is greater than a target judgment threshold;

[0525] The training unit is configured to train the initial detection model based on each positive target sample image in the sample training set to obtain a target detection model.

[0526] In a possible implementation, the second determining unit includes:

[0527] a fourth determining unit, configured to determine, if the real area of ​​the real region is smaller than a preset area, a second intersection-and-union coefficient corresponding to the second intersection-and-union ratio based on a ratio of a real area mean to the real area of ​​the real region; the real area mean is used to represent an average value of the real areas of the real regions corresponding to the sample cleaning targets of each target sample image in the sample training set; and the second intersection-and-union coefficient is greater than 1;

[0528] The fifth determining unit is configured to determine a target determination parameter corresponding to the target sample image based on the second IoU and a second IoU coefficient corresponding to the second IoU.

[0529] In a possible implementation, the second intersection-over-combination coefficient is:

[0530] Wherein, t is the second intersection-over-combination coefficient; s mean is the true area mean; s is the true area of ​​the sample cleaning target corresponding to the i-th target sample image in the sample training set The actual area; The real area corresponding to the i-th target sample image With the candidate region b i m is the second intersection-and-union ratio coefficient threshold corresponding to the second intersection-and-union ratio.

[0531] In a possible implementation, if the first target detection result contains first detection information of the target to be cleared, the application module 2602 further includes:

[0532] The first control unit is used to control the underwater cleaning robot to move toward the target to be cleaned based on the first detection information, so as to track and clean the target to be cleaned.

[0533] In a possible implementation, the application module 2602 further includes:

[0534] a third determining unit, configured to determine that the object to be cleaned is a stubborn stain if the object to be cleaned is still not cleaned after the underwater cleaning robot has cleaned the object to be cleaned a number of times reaching a target value;

[0535] a path optimization unit, configured to optimize a target inspection and cleaning path of the underwater cleaning robot in the area to be cleaned based on a target position of the stubborn stain in the area to be cleaned; and / or,

[0536] a model optimization unit, configured to optimize a target detection model corresponding to the underwater cleaning robot based on the first environment image corresponding to the stubborn stain; and / or,

[0537] The odometer optimization unit is used for performing loop detection based on the target position of the stubborn stain during the mapping or positioning process of the underwater cleaning robot in the area to be cleaned, so as to eliminate the cumulative error of the odometer of the underwater cleaning robot.

[0538] In a possible implementation, the application module 2602 further includes:

[0539] A fourth determining unit, configured to determine a target cleaning trajectory corresponding to the target to be cleaned based on the trajectory correlation;

[0540] A path planning unit is used to perform path planning based on the target cleaning trajectory to obtain a cleaning path to be inspected corresponding to the underwater cleaning robot;

[0541] The second control unit is used to control the underwater cleaning robot to continue to inspect and clean the area to be cleaned based on the inspection and cleaning path.

[0542] In a possible implementation, the application module 2602 further includes:

[0543] A trajectory optimization unit, configured to optimize the target cleaning trajectory using an interpolation algorithm when the computing power of the underwater cleaner is greater than a preset computing power;

[0544] The path planning unit is specifically used for:

[0545] Path planning is performed based on the optimized target cleaning trajectory and the inspected cleaning path of the underwater cleaning robot to obtain the cleaning path to be inspected corresponding to the underwater cleaning robot.

[0546] In a possible implementation, the application module 2602 further includes:

[0547] a first trajectory initialization unit configured to initialize the cleaning trajectory corresponding to N frames of environmental images corresponding to the specified type after the first environmental image as a new cleaning trajectory corresponding to the underwater cleaning robot, if the trajectory correlation is less than a correlation threshold and the target to be cleaned belongs to a specified type, and if the trajectory correlation between the N frames of environmental images corresponding to the specified type after the first environmental image and the first environmental image is less than the correlation threshold; where N is a positive integer greater than 1;

[0548] The second trajectory initialization unit is used to determine a new cleaning trajectory corresponding to the underwater cleaning robot based on the cleaning position corresponding to the third environmental image when the target to be cleaned appears in the third environmental image if the trajectory correlation is less than the correlation threshold and the target to be cleaned belongs to a non-specified category; the third environmental image is an environmental image collected after the first environmental image; wherein the target to be cleaned belonging to the specified category is used to represent the target to be cleaned that can be cleaned by the underwater cleaning robot in one go.

[0549] It should be noted that the visual information utilization device provided in the above embodiments utilizes visual information only by exemplifying the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be distributed among different functional modules as needed, i.e., the internal structure of the robot controller can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the visual information utilization device provided in the above embodiments and the visual information utilization method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0550] The present invention also provides a pool cleaning robot. FIG31 is a schematic diagram of a robot controller provided by the present invention. Generally, the pool cleaning robot includes a robot controller 3100 , which includes one or more processors 3101 and one or more memories 3102 .

[0551] The processor 3101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 3101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 3101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 3101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 3101 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0552] Memory 3102 may include one or more computer-readable storage media, which may be non-transitory. Memory 3102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 3102 is used to store at least one computer program, which is executed by processor 3101 to implement the method for using visual information provided in the method embodiment of the present application.

[0553] In some embodiments, the pool cleaning robot 3100 may optionally include a peripheral device interface 3103 and at least one peripheral device. The processor 3101, memory 3102, and peripheral device interface 3103 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 3103 via a bus, signal lines, or circuit boards.

[0554] Those skilled in the art will appreciate that the structure shown in FIG. 31 does not limit the pool cleaning robot 3100 , and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0555] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The computer program can be executed by a processor to implement the method for utilizing visual information in the above-described embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0556] In an exemplary embodiment, a computer program product or computer program is also provided, which includes a program code, which is stored in a computer-readable storage medium. The processor of the robot controller reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the robot controller executes the above-mentioned method of applying visual information.

[0557] In some embodiments, the computer program involved in the embodiments of the present application can be deployed and executed on a robot controller, or on multiple robot controllers located at one location, or on multiple robot controllers distributed at multiple locations and interconnected through a communication network. Multiple robot controllers distributed at multiple locations and interconnected through a communication network can form a blockchain system.

[0558] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0559] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for using visual information, the method comprising: Obtaining visual information collected by a pool cleaning robot; Controlling the pool cleaning robot based on the visual information, or processing the visual information in a preset manner.

2. The method according to claim 1, wherein The obtaining of the visual information collected by the pool cleaning robot includes: When the pool cleaning robot is performing a cleaning task in the pool, obtaining an environmental image around the pool cleaning robot; The controlling of the pool cleaning robot based on the visual information includes: Inputting the environmental image into an image recognition model, performing object detection and semantic segmentation on the environmental image through the image recognition model, to obtain a first recognition frame and a second recognition frame on the environmental image, and a position type of the position where the pool cleaning robot is located, where the first recognition frame is used to indicate the position of a target to be cleaned on the environmental image, and the second recognition frame is used to indicate the position of an obstacle on the environmental image; Controlling the pool cleaning robot to clean the target to be cleaned based on the environmental image, the first recognition frame, the second recognition frame, and the position type of the position where the pool cleaning robot is located.

3. The method according to claim 2, wherein The performing of object detection and semantic segmentation on the environmental image through the image recognition model, to obtain a first recognition frame and a second recognition frame on the environmental image, and a position type of the position where the pool cleaning robot is located, includes: Performing feature extraction on the environmental image through the image recognition model to obtain environmental image features of the environmental image; Performing border detection on the environmental image based on the environmental image features through the image recognition model to obtain the first recognition frame and the second recognition frame; Performing classification on multiple pixel points in the environmental image based on the environmental image features through the image recognition model to obtain the position type of the position where the pool cleaning robot is located.

4. The method according to claim 3, wherein, The performing of border detection on the environmental image based on the environmental image features to obtain the first recognition frame and the second recognition frame includes: Controlling a first candidate recognition frame and a second candidate recognition frame to slide on the environmental image, where the first candidate recognition frame is used to identify a target to be cleaned, and the second candidate recognition frame is used to identify an obstacle; Determining the first recognition frame and the second recognition frame based on first sub-image features corresponding to multiple first image regions in the environmental image features and second sub-image features corresponding to multiple second image regions in the environmental image features; Wherein, the first image region is an image region covered by the first candidate recognition frame on the environmental image, and the second image region is an image region covered by the second candidate recognition frame on the environmental image.

5. The method according to claim 4, wherein The determining of the first recognition frame and the second recognition frame based on first sub-image features corresponding to multiple first image regions in the environmental image features and second sub-image features corresponding to multiple second image regions in the environmental image features includes: For any one of the multiple first image regions, perform fully connected and normalization operations on the first sub-image features corresponding to the first image region to determine whether the first image region contains a target to be cleaned; in the case where the first image region contains a target to be cleaned, determine the border of the first image region as the first reference recognition frame; based on the multiple first reference recognition frames on the environmental image, determine the first recognition frame; For any one of the multiple second image regions, perform fully connected and normalization operations on the second sub-image features corresponding to the second image region to determine whether the second image region contains an obstacle; in the case where the second image region contains an obstacle, determine the border of the second image region as the second reference recognition frame; based on the multiple second reference recognition frames on the environmental image, determine the second recognition frame.

6. The method according to claim 3, wherein, The classifying the multiple pixel points in the environmental image based on the environmental image features to obtain the position type of the position where the pool cleaning robot is located includes: Perform multiple upsampling operations on the environmental image features to obtain an upsampled feature map of the environmental image, the size of the upsampled feature map being the same as that of the environmental image, the upsampled feature map including multiple channels, the multiple channels corresponding to multiple candidate position types, and the pixel values of the pixel points in each channel representing the confidence that the corresponding pixel points in the environmental image are in the candidate position corresponding to the channel; Based on the upsampled feature map, determine the position type corresponding to each pixel point in the environmental image; Based on the position types corresponding to the multiple pixel points, determine the position type of the position where the pool cleaning robot is located.

7. The method according to claim 6, wherein, The determining the position type of the position where the pool cleaning robot is located based on the position types corresponding to the multiple pixel points includes: Based on the position types corresponding to the multiple pixel points, determine the position type in front of the traveling direction of the pool cleaning robot; Use the position type in front of the traveling direction of the pool cleaning robot to determine the position type of the position where the pool cleaning robot is located.

8. The method according to claim 2, wherein, The controlling the pool cleaning robot to clean the target to be cleaned based on the environmental image, the first recognition frame, the second recognition frame, and the position type of the position where the pool cleaning robot is located includes: Determine the cleaning mode corresponding to the position type of the position where the pool cleaning robot is located; Based on the environmental image, the first recognition frame, the second recognition frame, and the cleaning mode, perform path planning on the pool cleaning robot to obtain a target movement trajectory; Control the pool cleaning robot to move along the target movement trajectory in the cleaning mode to clean the target to be cleaned in the pool.

9. The method according to claim 8, wherein The performing path planning on the pool cleaning robot based on the environmental image, the first recognition frame, the second recognition frame, and the cleaning mode to obtain a target movement trajectory includes: Determine the preset cleaning trajectory corresponding to the cleaning mode; determining a reference movement trajectory based on positions of the first identification frame and the second identification frame in the environment image, an obstacle type of the obstacle in the second identification frame, and a size of the second identification frame; The preset cleaning trajectory and the reference moving trajectory are combined to obtain the target moving trajectory.

10. The method according to claim 9, wherein, The determining of the reference movement trajectory based on the positions of the first identification frame and the second identification frame in the environment image, the obstacle type of the obstacle in the second identification frame, and the size of the second identification frame includes: Determining whether the pool cleaning robot needs to avoid the obstacle based on the obstacle type of the obstacle, the position of the second identification frame in the environment image, and the size of the second identification frame; In the case where the pool cleaning robot needs to avoid the obstacle, determining a first moving direction and a first moving distance of the pool cleaning robot based on a relative positional relationship between the first identification frame, the second identification frame and a center point of the environment image; The reference movement trajectory is generated based on a first movement direction and a first movement distance of the pool cleaning robot.

11. The method according to claim 10, wherein, The method further comprises: In the case where the pool cleaning robot does not need to avoid the obstacle, determining a second moving direction and a second moving distance of the pool cleaning robot based on a relative position relationship between the first recognition frame and a center point of the environment image; The reference movement trajectory is generated based on a second movement direction and a second movement distance of the pool cleaning robot.

12. The method according to claim 8, wherein Before performing path planning for the pool cleaning robot based on the environment image, the first recognition frame, the second recognition frame, and the cleaning mode to obtain a target moving trajectory, the method further includes: Determining an obstacle climbing level of the pool cleaning robot, wherein the obstacle climbing level is positively correlated with the obstacle climbing capability; The performing path planning for the pool cleaning robot based on the environment image, the first recognition frame, the second recognition frame, and the cleaning mode to obtain a target movement trajectory includes: When the obstacle climbing level of the pool cleaning robot is less than or equal to a preset level, segmenting the obstacle in the second identification frame to obtain a contour line of the obstacle in the environment image; The path of the pool cleaning robot is planned based on the environment image, the first recognition frame, the outline of the obstacle, and the cleaning mode to obtain a target movement trajectory.

13. The method according to claim 8, wherein, The step of determining the cleaning mode corresponding to the position type of the pool cleaning robot comprises: When the pool cleaning robot is located at the pool wall of the pool, the cleaning mode of the pool cleaning robot is determined to be a roller brush cleaning mode; When the pool cleaning robot is located at the bottom of the pool, the cleaning mode of the pool cleaning robot is determined to be a water pumping mode; When the pool cleaning robot is located at a boundary between the bottom and the wall of the pool, the cleaning mode of the pool cleaning robot is determined to be a side cleaning mode; When the position of the pool cleaning robot is on the step of the pool, determine the cleaning mode of the pool cleaning robot as the basic cleaning mode or the deep cleaning mode. In the basic cleaning mode, clean the step plane and do not clean the step vertical surface. In the deep cleaning mode, clean the step vertical surface and the step plane.

14. The method according to claim 1, wherein The obtaining of the visual information collected by the pool cleaning robot includes: Obtain the position of the pool cleaning robot and the position of the target garbage to be cleaned around the pool cleaning robot; The controlling of the pool cleaning robot based on the visual information includes: Generate a target cleaning path of the pool cleaning robot based on the position of the pool cleaning robot, the position of the target garbage, and the target garbage type of the target garbage. The target cleaning path is the cleaning path used when cleaning the target garbage, and the target garbage type is used to represent the shape of the target garbage; Control the pool cleaning robot to clean the target garbage based on the target cleaning path.

15. The method according to claim 14, wherein, The target cleaning path is the first cleaning path, the second cleaning path, or the third cleaning path. The generating of the target cleaning path of the pool cleaning robot based on the position of the pool cleaning robot, the position of the target garbage, and the target garbage type of the target garbage includes: When the target garbage type is the point type, determine the central position of the center point of the target garbage based on the position of the target garbage; perform path planning between the position of the pool cleaning robot and the central position to obtain the first cleaning path; When the target garbage type is the line segment type, determine the line segment position of the line segment corresponding to the target garbage based on the position of the target garbage; generate the second cleaning path based on the position of the pool cleaning robot and the line segment position; When the target garbage type is the surface type, determine the area where the target garbage is located based on the position of the target garbage; generate the third cleaning path based on the position of the pool cleaning robot and the area where the target garbage is located.

16. The method according to claim 15, wherein, The generating of the second cleaning path based on the position of the pool cleaning robot and the line segment position includes: Determine the starting position of the line segment corresponding to the target garbage from the line segment positions; Perform path planning between the position of the pool cleaning robot and the starting position to obtain a first initial cleaning path; Add a second initial cleaning path of the line segment corresponding to the target garbage to the first initial cleaning path to obtain the second cleaning path, and the second initial cleaning path passes through the starting point and the ending point of the line segment corresponding to the target garbage.

17. The method according to claim 15, wherein, The generating of the third cleaning path based on the position of the pool cleaning robot and the area where the target garbage is located includes: Determine the vertex position of any vertex of the target garbage from the area where the target garbage is located; Perform path planning between the position of the pool cleaning robot and the vertex position to obtain a third initial cleaning path; Starting from the vertex position, perform path planning based on the area where the target garbage is located to obtain a fourth initial cleaning path that covers the area where the target garbage is located; Combine the third initial cleaning path and the fourth initial cleaning path to obtain the third cleaning path.

18. The method according to claim 14, wherein, Before obtaining the position of the pool cleaning robot and the position of the target garbage to be cleaned around the pool cleaning robot, the method further includes: Determine the garbage types of multiple candidate garbage around the pool cleaning robot; Based on the position of the pool cleaning robot, the garbage types, positions, and garbage areas of each candidate garbage, determine the cleaning cost between the pool cleaning robot and each candidate garbage; Based on the cleaning cost between the pool cleaning robot and each candidate garbage, determine the target garbage from the multiple candidate garbage, where the target garbage is the candidate garbage with the minimum corresponding cleaning cost.

19. The method according to claim 18, wherein The determining the garbage types of multiple candidate garbage around the pool cleaning robot includes: Determine the contours of each candidate garbage; Based on the contours of each candidate garbage, determine the circumscribed rectangles of each candidate garbage; Based on the size information of the circumscribed rectangles of each candidate garbage, determine the garbage types of each candidate garbage.

20. The method according to claim 19, wherein, The obtaining the contours of each candidate garbage includes: Obtain the environmental image around the pool cleaning robot; Perform perspective transformation on the environmental image to obtain the BEV image corresponding to the environmental image; Perform object detection on the BEV image to obtain the contours of each candidate garbage.

21. The method according to claim 20, wherein, The obtaining the environmental image around the pool cleaning robot includes: Control the pool cleaning robot to rotate and obtain multiple initial environmental images during the rotation; Stitch the multiple initial environmental images to obtain the environmental image around the pool cleaning robot.

22. The method according to claim 19, wherein, The size information includes length and width. The determining the garbage types of each candidate garbage based on the size information of the circumscribed rectangles of each candidate garbage includes: For any candidate garbage among the multiple candidate garbage, when both the length and width of the circumscribed rectangle of the candidate garbage are less than or equal to a preset threshold, determine the target garbage type of the target garbage as the point type, where the preset threshold is associated with the single - time cleaning width of the pool cleaning robot; When the length of the circumscribed rectangle of the candidate garbage is less than or equal to the preset threshold and the width is greater than the preset threshold, or when the width of the circumscribed rectangle of the candidate garbage is less than or equal to the preset threshold and the length is greater than the preset threshold, determine the target garbage type of the target garbage as the line segment type; When both the length and width of the circumscribed rectangle of the candidate garbage are greater than the preset threshold, determine the target garbage type of the target garbage as the surface type.

23. The method according to claim 19, wherein The method for obtaining the garbage area of each candidate garbage includes: Determine the area of the circumscribed rectangle of each candidate garbage as the garbage area of each candidate garbage; Alternatively, the area enclosed by the contour of each of the candidate wastes is determined as the waste area of each of the candidate wastes.

24. The method according to claim 18, wherein Determining the cleaning cost between the pool cleaning robot and each of the candidate wastes based on the position of the pool cleaning robot, the waste type, position, and waste area of each of the candidate wastes includes: Based on the position of the pool cleaning robot, the waste type, and position of each of the candidate wastes, determining the distance between the pool cleaning robot and each of the candidate wastes; Based on the distance between the pool cleaning robot and each of the candidate wastes and the waste area of each of the candidate wastes, determining the cleaning cost between the pool cleaning robot and each of the candidate wastes.

25. The method according to claim 24, wherein The determining the distance between the pool cleaning robot and each of the candidate wastes based on the position of the pool cleaning robot, the waste type, and position of each of the candidate wastes includes: For any one of the multiple candidate wastes, when the candidate waste type is a point type, based on the position of the candidate waste, determining the central position of the center point of the candidate waste; and determining the distance between the position of the pool cleaning robot and the central position as the distance between the pool cleaning robot and the candidate waste; When the candidate waste type is a line segment type, based on the position of the candidate waste, determining the midpoint position of the midpoints of the two short sides of the circumscribed rectangle of the candidate waste; and based on the position of the pool cleaning robot and the midpoint position of the two short sides, determining the distance between the pool cleaning robot and the candidate waste; When the candidate waste type is a surface type, based on the position of the candidate waste, determining the vertex positions of the four vertices of the circumscribed rectangle of the candidate waste; and based on the position of the pool cleaning robot and the vertex positions of the four vertices, determining the distance between the pool cleaning robot and the candidate waste.

26. The method according to claim 25, wherein, The two short side midpoints include a first short side midpoint and a second short side midpoint, and the determining the distance between the pool cleaning robot and the candidate waste based on the position of the pool cleaning robot and the midpoint position of the two short sides includes: Determining a first short side reference distance between the position of the pool cleaning robot and the midpoint position of the first short side; Determining a second short side reference distance between the position of the pool cleaning robot and the midpoint position of the second short side; Determining the shorter distance among the first short side reference distance and the second short side reference distance as the distance between the pool cleaning robot and the candidate waste.

27. The method according to claim 25, wherein, The four vertices include a first vertex, a second vertex, a third vertex, and a fourth vertex, and the determining the distance between the pool cleaning robot and the candidate waste based on the position of the pool cleaning robot and the vertex positions of the four vertices includes: Determining a first vertex reference distance between the position of the pool cleaning robot and the vertex position of the first vertex; Determine a second vertex reference distance between the position of the pool cleaning robot and the vertex position of the second vertex; Determine a third vertex reference distance between the position of the pool cleaning robot and the vertex position of the third vertex; Determine a fourth vertex reference distance between the position of the pool cleaning robot and the vertex position of the fourth vertex; Determine the shorter distance among the first vertex reference distance, the second vertex reference distance, the third vertex reference distance, and the fourth vertex reference distance as the distance between the pool cleaning robot and the candidate garbage.

28. The method according to claim 24, wherein The determining the cleaning cost between the pool cleaning robot and each candidate garbage based on the distance between the pool cleaning robot and each candidate garbage and the garbage area of each candidate garbage includes: Multiply the distance cost weight by the distance between the pool cleaning robot and each candidate garbage to obtain the distance cost between the pool cleaning robot and each candidate garbage; Divide the area cost weight by the garbage area of each candidate garbage to obtain the area cost of each candidate garbage; Add the distance cost and the area cost corresponding to each candidate garbage to obtain the cleaning cost between the pool cleaning robot and each candidate garbage.

29. The method according to claim 14, wherein After the method controls the pool cleaning robot to clean the target garbage based on the target cleaning path, the method further includes: When the movement of the target cleaning path is completed, determine whether there is still candidate garbage to be cleaned; When there is still candidate garbage to be cleaned, re-determine the target garbage and the target cleaning path from the candidate garbage to be cleaned; Control the pool cleaning robot to clean the re-determined target garbage based on the re-determined target cleaning path.

30. The method according to claim 29, wherein, The method further includes: When there is no candidate garbage to be cleaned, perform path planning based on the position of the pool cleaning robot and the position of the base station to obtain a target recharge path; Control the pool cleaning robot to dock with the base station according to the target recharge path.

31. The method according to claim 1, wherein, The obtaining the visual information collected by the pool cleaning robot includes: Obtain the environmental image collected by the pool cleaning robot; The processing the visual information in a preset manner includes: Extract features from the environmental image collected by the pool cleaning robot to obtain the target structure information in the target space where the pool cleaning robot is located; Perform feature matching based on the target structure information corresponding to two adjacent frames of the environmental image in the target space to obtain a corresponding target feature matching result; Determine the target depth information of the target structure boundary point in the target space in the target structure information based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot; the target acquisition position information is comprehensively determined based on the positioning information obtained by various sensors of the pool cleaning robot; the target structure boundary point is the matching structure boundary point in the target structure information corresponding to two adjacent frames of the environmental image. Construct a target environment map of the target space based on the target structure information and the target depth information corresponding to the target structure boundary points.

32. The method according to claim 31, wherein, The feature extraction of the environmental image collected by the pool cleaning robot to obtain the target structure information in the target space where the pool cleaning robot is located includes: Extract the structural boundary information in the environmental image collected by the pool cleaning robot; the structural boundary information includes the first position information corresponding to the structural boundary existing in the environmental image and the second position information of the area included in the contour corresponding to the structural boundary. Determine the target structure information in the target space based on the structural boundary information.

33. The method according to claim 31, wherein, The feature extraction of the environmental image collected by the pool cleaning robot to obtain the target structure information in the target space where the pool cleaning robot is located includes: Use a multi-task model to separately segment and identify the target space composition element information and the target space auxiliary element information in the environmental image collected by the pool cleaning robot, and obtain the target structure information in the target space where the pool cleaning robot is located.

34. The method according to claim 31, wherein, The feature matching based on the target structure information corresponding to two adjacent frames of the environmental image in the target space to obtain the corresponding target feature matching result includes: Predict the predicted acquisition position information of the pool cleaning robot corresponding to the latter frame of the environmental image in two adjacent frames of the environmental image based on the acquisition position information of the previous frame of the pool cleaning robot corresponding to the previous frame of the environmental image in two adjacent frames of the environmental image. Predict the predicted structure information in the latter frame of the environmental image based on the predicted acquisition position information and the target structure information in the previous frame of the environmental image. Match the target structure information in the latter frame of the environmental image with the predicted structure information to obtain the predicted structure matching range corresponding to the target structure information in the latter frame of the environmental image. Determine the matching target structure boundary points in the target structure information corresponding to each of the two adjacent frames of the environmental image based on the geometric features of each structural boundary point in the previous frame of the environmental image and the geometric features of each structural boundary point within the predicted structure matching range.

35. The method according to claim 31, wherein, The multiple sensors include a first sensor, a second sensor, and an ultrasonic sensor. The first sensor is used to collect the environmental image of the pool cleaning robot, and the second sensor is used to collect the motion data of the pool cleaning robot; the second sensor includes at least one of the following: an inertial measurement unit, an encoder, and a flow meter; the ultrasonic sensor is used to collect the ultrasonic positioning data of the pool cleaning robot. Before determining the target depth information of the target structure boundary points in the target structure information in the target space based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot, the method further includes: Obtain two adjacent frames of environmental images based on the first sensor, obtain the actual motion data of the pool cleaning robot based on one or more of the second sensors, and obtain the ultrasonic positioning data of the pool cleaning robot based on the ultrasonic sensor. Determine the first acquisition position information corresponding to the latter frame of the adjacent two frames of environmental images based on the adjacent two frames of environmental images and the previous frame acquisition position information corresponding to the previous frame of environmental image in the adjacent two frames of environmental images, determine the second acquisition position information corresponding to the latter frame of environmental image based on the actual motion data, and determine the third acquisition position information corresponding to the latter frame of environmental image based on the ultrasonic positioning data; Determine the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information.

36. The method according to claim 35, wherein, After determining the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information, the method further includes: Determine the position loss corresponding to the latter frame of environmental image based on the predicted acquisition position information, the first acquisition position information, the second acquisition position information, and the third acquisition position information corresponding to the latter frame of environmental image; Determine the image loss corresponding to the latter frame of environmental image based on the predicted structure information and the target structure information corresponding to the latter frame of environmental image; Correct the position drift of multiple sensors based on the position loss and the image loss corresponding to the latter frame of environmental image.

37. The method according to claim 35, wherein, Before determining the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information, the method further includes: Judge whether the ultrasonic positioning data meets the non-line-of-sight condition based on the latter frame of environmental image; the non-line-of-sight condition is used to characterize the situation that there is occlusion in the latter frame of environmental image and the ultrasonic positioning data belongs to abnormal jitter; If the ultrasonic positioning data meets the non-line-of-sight condition, determine the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information and the second acquisition position information; If the ultrasonic positioning data does not meet the non-line-of-sight condition, execute the step of determining the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information.

38. The method according to claim 35, wherein, Before determining the target acquisition position information corresponding to the latter frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information, the method further includes: Count the number of target structure boundary points that meet the preset constraint conditions in the latter frame of environmental image, and the total number of structure boundary points in the target structure information corresponding to the latter frame of environmental image; the preset constraint condition is that there are matching target structure boundary points in at least two consecutive frames of environmental images; Determine the visual constraint weight corresponding to the first sensor based on the ratio between the number of target structure boundary points that meet the preset constraint conditions and the total number of structure boundary points; Determine the weights corresponding to the second sensor and the ultrasonic sensor respectively based on the covariance matrices corresponding to the second sensor and the ultrasonic sensor respectively; Determine the target weight ratios corresponding to the first sensor, the second sensor, and the ultrasonic sensor based on the visual constraint weight corresponding to the first sensor and the weights corresponding to the second sensor and the ultrasonic sensor respectively; The determining the target acquisition position information corresponding to the subsequent frame of environmental image based on the first acquisition position information, the second acquisition position information, and the third acquisition position information includes: Fuse the first acquisition position information, the second acquisition position information, and the third acquisition position information according to the target weight ratios to obtain the target acquisition position information corresponding to the subsequent frame of environmental image.

39. The method according to claim 31, wherein, The target feature matching result includes the target structure boundary points that match in two adjacent frames of the environmental image; The determining the target depth information of the target structure boundary points in the target space in the target structure information based on the target feature matching result and the target acquisition position information corresponding to the pool cleaning robot includes: Estimate the target depth information of the target structure boundary points in the target space based on the positions of the target structure boundary points in two adjacent frames of the environmental image and the target acquisition position information corresponding to the pool cleaning robot.

40. The method according to claim 31, wherein, Before constructing the target environmental map of the target space based on the target structure information and the target depth information corresponding to the target structure boundary points, the method further includes: Move along the target mapping path and collect environmental images within a preset visual range at a preset frequency during the movement; The constructing the target environmental map of the target space based on the target structure information and the target depth information corresponding to the target structure boundary points includes: After the travel along the target mapping path is completed, construct the target environmental map of the target space based on the target structure information corresponding to multiple environmental images on the target mapping path and the target depth information corresponding to the target structure boundary points in the multiple environmental images.

41. The method according to claim 31, wherein, The target structure information includes the target space composition element information and the target space auxiliary element information in the target space; The target space composition element information includes the structural boundary information of the corresponding composition elements in the target space, and the composition elements include at least one of the following: bottom surface, wall surface, step plane, step elevation; The target space auxiliary element information includes the structural boundary information of the corresponding auxiliary elements in the target space, and the auxiliary elements include at least one of the following auxiliary devices: wall lamp, floor lamp, water inlet, water outlet.

42. The method according to claim 1, wherein, The obtaining the visual information collected by the pool cleaning robot includes: When the pool cleaning robot conducts inspections in the area to be cleaned, obtain the first target detection result corresponding to the first environmental image and the second target detection result corresponding to the second environmental image within the visual detection range of the pool cleaning robot; the second environmental image is the next frame image of the first environmental image; Processing the visual information in a preset manner includes: If there is first detection information of a target to be cleaned in the first target detection result, generating target prediction information corresponding to the second environmental image based on the first environmental image and the true motion information of the pool cleaning robot; Determining the trajectory relevance between the first environmental image and the second environmental image based on the second target detection result and the target prediction information according to the trajectory association strategy corresponding to the type of the target to be cleaned.

43. The method according to claim 42, wherein, The pool cleaning robot includes a motion sensor; Before generating the target prediction information corresponding to the second environmental image based on the first environmental image and the true motion information of the pool cleaning robot, the method further includes: Estimating the first motion information of the pool cleaning robot within a target time period based on the first environmental image and the second environmental image through image registration; the target time period is the time period between the first acquisition moment corresponding to the first environmental image and the second acquisition moment corresponding to the second environmental image; Obtaining the second motion information of the pool cleaning robot within the target time period through the motion sensor; Fusing the first motion information and the second motion information to obtain the true motion information corresponding to the pool cleaning robot.

44. The method according to claim 42, wherein, Generating the target prediction information corresponding to the second environmental image based on the first environmental image and the true motion information of the pool cleaning robot includes: Using a Kalman filter or a variant of the Kalman filter to generate initial prediction information corresponding to the second environmental image based on the first environmental image; Correcting the initial prediction information based on the true motion information of the pool cleaning robot to obtain the target prediction information corresponding to the second environmental image.

45. The method according to claim 42, wherein, The second target detection result includes the detection area corresponding to the target to be cleaned in the second environmental image; the target prediction information includes the prediction area corresponding to the target to be cleaned in the second environmental image; Determining the trajectory relevance between the first environmental image and the second environmental image based on the second target detection result and the target prediction information according to the trajectory association strategy corresponding to the type of the target to be cleaned includes: Calculating the first intersection over union between the detection area and the prediction area; If the target to be cleaned belongs to a rigid target, determining the trajectory relevance between the first environmental image and the second environmental image based on the first intersection over union; If the target to be cleaned is a non-rigid target, determining the feature correlation degree of the target to be cleaned between the first environmental image and the second environmental image based on the first feature corresponding to the target to be cleaned in the first environmental image and the second feature corresponding to the target to be cleaned in the second environmental image, and determining the trajectory relevance between the first environmental image and the second environmental image based on the feature correlation degree and the first intersection over union; If the target to be cleaned is a small target, determine the trajectory correlation between the first environmental image and the second environmental image based on the first intersection over union (IoU) and the first IoU coefficient corresponding to the first IoU; the small target is used to represent a target to be cleaned with an imaging area smaller than a preset value or a short side dimension smaller than a preset dimension in the first environmental image, and the first IoU coefficient is greater than or equal to 1.

46. The method according to claim 45, wherein, Before determining the trajectory correlation between the first environmental image and the second environmental image based on the first IoU and the first IoU coefficient corresponding to the first IoU, the method further includes: If the target to be cleaned is a small target, determine the first IoU coefficient corresponding to the first IoU based on the proportional relationship between the predicted area of the predicted region and the detected area of the detected region.

47. The method according to claim 42, wherein, The pool cleaning robot includes an image acquisition device; The obtaining the first target detection result corresponding to the first environmental image and the second target detection result corresponding to the second environmental image within the visual detection range of the pool cleaning robot includes: Obtain the first environmental image and the second environmental image within the visual detection range of the pool cleaning robot through the image acquisition device; Input the first environmental image into the target detection model, and output the first target detection result corresponding to the first environmental image; the target detection model is trained based on a sample training set; the sample training set includes multiple target sample images with target ground truth information corresponding to known sample cleaning targets; Input the second environmental image into the target detection model, and output the first target detection result corresponding to the second environmental image.

48. The method according to claim 47, wherein, The target ground truth information includes the category to which the sample cleaning target belongs in the target sample image, as well as the real position and real area of the real region corresponding to the sample cleaning target; Before inputting the first environmental image into the target detection model and outputting the first target detection result corresponding to the first environmental image, the method further includes: Obtain the sample training set; Using an initial detection model, generate candidate regions corresponding to the sample cleaning targets based on the target sample images; Determine the target determination parameter corresponding to the target sample image based on the second IoU between the real region and the candidate region; Judge whether the target sample image is a positive target sample image based on the target determination parameter; the positive target sample image is used to represent a target sample image in the sample training set corresponding to the target determination parameter greater than a target determination threshold; If so, train the initial detection model based on each positive target sample image in the sample training set to obtain a target detection model.

49. The method according to claim 48, wherein The determining the target determination parameter corresponding to the target sample image based on the second IoU between the real region and the candidate region includes: If the true area of the true region is less than the preset area, determine the second intersection over union coefficient corresponding to the second intersection over union based on the ratio of the average true area to the true area of the true region; the average true area is used to represent the average of the true areas of the true regions corresponding to the sample cleaning targets in the sample training set; the second intersection over union coefficient is greater than 1; Determine the target determination parameter corresponding to the target sample image based on the second intersection over union and the second intersection over union coefficient corresponding to the second intersection over union.

50. The method according to claim 49, wherein, The second intersection-over-union coefficient is as follows: where t is the second intersection-over-union coefficient; s mean is the average real area; s is the real region of the sample cleaning target corresponding to the i-th target sample image in the sample training set The true area; corresponding to the i-th target sample image for the true region with the candidate region b i the second intersection over union ratio therebetween; m is the second intersection over union ratio coefficient threshold corresponding to the second intersection over union ratio.

51. The method according to claim 42, wherein, If there is the first detection information of the target to be cleaned in the first target detection result, the method further includes: Control the pool cleaning robot to move towards the target to be cleaned based on the first detection information to perform tracking cleaning on the target to be cleaned.

52. The method according to claim 51, wherein, After performing the tracking cleaning on the target to be cleaned, the method further includes: If the target to be cleaned is still not cleaned after the number of times the pool cleaning robot cleans the target to be cleaned reaches the target value, determine that the target to be cleaned belongs to stubborn stains; Optimize the target inspection and cleaning path of the pool cleaning robot in the area to be cleaned based on the target position of the stubborn stains in the area to be cleaned; and / or, optimize the target detection model corresponding to the pool cleaning robot based on the first environmental image corresponding to the stubborn stains; and / or, during the process of map building or positioning of the pool cleaning robot in the area to be cleaned, perform loop detection based on the target position of the stubborn stains to eliminate the cumulative error of the odometer of the pool cleaning robot.

53. The method according to claim 42, wherein, After determining the trajectory relevance between the first environmental image and the second environmental image according to the trajectory association strategy corresponding to the type of the target to be cleaned based on the second target detection result and the target prediction information, the method further includes: Determine the target cleaning trajectory corresponding to the target to be cleaned based on the trajectory relevance; Perform path planning based on the target cleaning trajectory to obtain the target inspection and cleaning path corresponding to the pool cleaning robot; Control the pool cleaning robot to continue performing inspection and cleaning in the area to be cleaned based on the target inspection and cleaning path.

54. The method according to claim 53, wherein, After determining the target cleaning trajectory corresponding to the target to be cleaned based on the trajectory relevance and before performing path planning based on the target cleaning trajectory to obtain the target inspection and cleaning path corresponding to the pool cleaning robot, the method further includes: Optimize the target cleaning trajectory using an interpolation algorithm when the computing power of the pool cleaning robot is greater than the preset computing power; Performing path planning based on the target cleaning trajectory to obtain the target inspection and cleaning path corresponding to the pool cleaning robot includes: Perform path planning based on the optimized target cleaning trajectory and the inspected and cleaned path of the pool cleaning robot to obtain the target inspection and cleaning path corresponding to the pool cleaning robot.

55. The method according to claim 42, wherein After determining the trajectory relevance between the first environmental image and the second environmental image based on the trajectory association strategy corresponding to the type to which the target to be cleaned belongs according to the second target detection result and the target prediction information, the method further includes: If the trajectory relevance is less than the relevance threshold and the target to be cleaned belongs to a specified type, when the trajectory relevance between each of the N environmental images corresponding to the specified type after the first environmental image and the first environmental image is less than the relevance threshold, initialize the cleaning trajectory corresponding to the N environmental images as a new cleaning trajectory corresponding to the pool cleaning robot; N is a positive integer greater than 1; If the trajectory relevance is less than the relevance threshold and the target to be cleaned belongs to a non-specified type, when the target to be cleaned appears in the third environmental image, determine a new cleaning trajectory corresponding to the pool cleaning robot based on the cleaning position corresponding to the third environmental image; the third environmental image is an environmental image collected after the first environmental image; Among them, the target to be cleaned belonging to the specified type is used to represent the target to be cleaned that the pool cleaning robot can complete cleaning in one go.

56. An apparatus for using visual information, the apparatus includes: A visual information acquisition module, configured to acquire visual information collected by a pool cleaning robot; An application module, configured to control the pool cleaning robot based on the visual information, or process the visual information in a preset manner.

57. A pool cleaning robot, characterized in that, The pool cleaning robot includes a robot controller, the robot controller includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the method for using visual information according to any one of claims 1 to 55.

Citation Information

Patent Citations

  • Target tracking method and device for robot and medium

    CN114004863A

  • Swimming pool cleaning robot and swimming pool cleaning method

    CN114109095A

  • Inspection cleaning method, device and equipment and storage medium

    CN115933685A

  • Map building and visual positioning method and device of robot

    CN116164728A

  • Sweeping path control method and device of swimming pool robot and swimming pool robot

    CN116540702A

Cited By

  • Control method and device for cleaning robot for underwater steps of swimming pool

    CN120909298A

  • Path planning method and system and cleaning robot

    CN121028842A

  • Water garbage automatic collection method and device based on dynamic path planning

    CN121147734A

  • An automatic water garbage collection method and device based on dynamic path planning

    CN121147734B