Cleaning method and cleaning robot
By using deep learning algorithms and multi-frame image fusion discrimination strategies in cleaning robots, combined with temporal behavior recognition algorithms, the problem of traditional cleaning robots being unable to accurately identify pet behavior has been solved, enabling intelligent adjustment of cleaning strategies and precise monitoring of pet behavior.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DREAM INNOVATION TECH (SUZHOU) CO LTD
- Filing Date
- 2025-08-29
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional cleaning robots cannot accurately identify pet behavior or perform intelligent analysis to adjust cleaning logic, resulting in cleaning strategies that are not dynamic or timely enough.
By acquiring a single-frame image and performing initial behavior recognition based on a preset deep learning algorithm, and combining a multi-frame image fusion discrimination strategy and a temporal behavior recognition algorithm, the target behavior recognition result is determined to formulate a cleaning strategy.
It improves the accuracy and stability of pet behavior recognition, accurately reflects the pet's real behavioral state, provides a reliable basis for cleaning robots, and enables intelligent cleaning processes.
Smart Images

Figure CN121963306A_ABST
Abstract
Description
A cleaning method and a cleaning robot
[0001] This application is a divisional application of application number 2025112236437 (A cleaning method, apparatus, cleaning robot and readable storage medium, application date: 2025-08-29). Technical Field
[0002] This invention belongs to the field of image recognition technology, and specifically relates to a cleaning method, apparatus, cleaning robot, and readable storage medium. Background Technology
[0003] With the increasing rate of pet ownership in households, the demand for cleaning robots among pet-owning families is becoming more and more prominent.
[0004] In traditional technologies, cleaning robots have developed a variety of pet-related functions, including pet video monitoring via front-end cameras, cat and dog recognition for obstacle avoidance, capturing exciting moments of pets, identifying pet activity hotspots, and recognizing pet supplies.
[0005] However, in traditional technologies, pet monitoring is based solely on video surveillance. This not only requires human control or relies on synchronous monitoring during work, but also treats pets as ordinary obstacles, resulting in low accuracy in recognizing pet behavior and an inability to intelligently analyze their behavior to adjust cleaning logic. Summary of the Invention
[0006] Therefore, the technical problem to be solved by this invention is that the accuracy of pet behavior recognition is low, and it is impossible to intelligently analyze the behavior to adjust the cleaning logic.
[0007] To address the aforementioned technical problems, in a first aspect, the present invention provides a cleaning method applied to a cleaning robot, the method comprising:
[0008] A single-frame image is acquired, and the single-frame image is recognized based on a preset deep learning algorithm to obtain an initial behavior recognition result;
[0009] Based on a multi-frame image fusion and discrimination strategy, the initial behavior recognition results corresponding to each single-frame image are fused and judged to obtain the first behavior recognition result;
[0010] If the first behavior recognition result is an abnormality, each of the single-frame images is accumulated, and the accumulated multi-frame images are subjected to temporal behavior recognition according to the temporal behavior recognition algorithm to determine the second behavior recognition result.
[0011] Based on the first behavior recognition result and the second behavior recognition result, a target behavior recognition result is determined; the target behavior recognition result is used to determine a target cleaning strategy for cleaning treatment.
[0012] In one embodiment, prior to acquiring the single-frame image, the method further includes:
[0013] The image acquisition unit captures a single frame of a pet activity scene.
[0014] A single frame image is input into a pre-trained deep learning model, and feature extraction and pet behavior category prediction are performed on the single frame image based on the deep learning model, outputting the probability values of each type of pet behavior;
[0015] The initial behavior recognition results are obtained based on the probability values of the various pet behaviors.
[0016] In one embodiment, the multi-frame image fusion discrimination strategy, which fuses and judges the initial behavior recognition results corresponding to each single frame image to obtain a first behavior recognition result, includes:
[0017] If the initial behavior recognition results corresponding to multiple consecutive single-frame images are all the same, then the initial behavior recognition result is determined as the first behavior recognition result;
[0018] If the initial behavior recognition results corresponding to multiple consecutive single-frame images are not all the same, then the first behavior recognition result is determined as the default value.
[0019] In one embodiment, the first behavior identification result being an anomaly includes: the first behavior identification result being a default value, and / or the first behavior identification result being that the pet exhibits abnormal behavior.
[0020] In one embodiment, the accumulation of each of the single-frame images, and the performance of temporal behavior recognition on the accumulated multi-frame images according to a temporal behavior recognition algorithm to determine a second behavior recognition result, includes:
[0021] Based on the accumulated single-frame images within the target duration in a preset time sequence, a multi-frame image sequence is obtained;
[0022] Spatial features are extracted from the multi-frame image sequence using a temporal behavior recognition model to obtain feature vectors;
[0023] Based on temporal correlation, the feature vectors are fused to obtain a fused feature vector;
[0024] The fused feature vector is used to perform pet behavior recognition to determine the second behavior recognition result.
[0025] In one embodiment, the step of accumulating each single frame image within a preset time sequence to obtain a multi-frame image sequence includes:
[0026] Based on the accumulated single-frame images within a preset time sequence within the target duration, a video segment is obtained;
[0027] The video segment is input into a fast-slow network, and dual-path frame sampling is performed through the fast-slow network to obtain the first keyframe image sequence corresponding to the slow path and the second keyframe image sequence corresponding to the fast path.
[0028] In one embodiment, the step of extracting spatial features from the multi-frame image sequence using a temporal behavior recognition model to obtain a feature vector includes:
[0029] The fast and slow network is used to perform dual-path feature extraction on the first keyframe image sequence and the second keyframe image sequence to obtain the static feature vector corresponding to the slow path and the dynamic feature vector corresponding to the fast path.
[0030] In one embodiment, the step of fusing the feature vectors based on temporal correlation to obtain a fused feature vector includes:
[0031] Adjust the number of channels in the dynamic feature vector corresponding to the fast path, and align the adjusted dynamic feature vector with the static feature vector corresponding to the slow path by channel.
[0032] The fast path is dimensionality-reduced by using the time-averaged pooling layer or convolutional layer in the fast and slow network, so that the dynamic feature vector of the dimensionality-reduced fast path is time-aligned with the static feature vector of the corresponding slow path.
[0033] The dynamic feature vectors and static feature vectors that are aligned by channels and time are fused together to obtain a fused feature vector.
[0034] In one embodiment, the step of performing pet behavior recognition on the fused feature vector to determine the second behavior recognition result includes:
[0035] The fused feature vector is processed by the fully connected layer in the fast and slow network to identify pet behaviors and output the probability values of various pet behaviors.
[0036] The second behavior recognition result is determined based on the probability values of the various pet behaviors.
[0037] In one embodiment, the method further includes:
[0038] Based on the target behavior recognition results, a target cleaning strategy is determined, which includes one or more of cleaning, avoidance, and early warning.
[0039] In one embodiment, determining the target cleaning strategy based on the target behavior recognition result includes:
[0040] When the target behavior identification result is eating behavior, a virtual eating area is constructed based on the pet's location or the location of the food items, and the virtual eating area is marked as the first cleaning area;
[0041] Clean all other cleaning areas except the first cleaning area, and continuously monitor the cleaning progress and the pet's activity status in the first cleaning area;
[0042] Cleaning of the first cleaning area is performed when the other cleaning areas have been cleaned and / or the pet has left the first cleaning area.
[0043] In one embodiment, performing the cleaning of the first cleaning area includes:
[0044] Identify the cleaning area and / or degree of dirt in the first cleaning area, and clean the first cleaning area based on the cleaning area and / or degree of dirt.
[0045] In one embodiment, identifying the cleaning area and / or degree of dirt in the first cleaning area, and cleaning the first cleaning area based on the cleaning area and / or the degree of dirt, includes:
[0046] When the cleaning area of the first cleaning area is greater than the area threshold or the degree of dirtiness is greater than the dirtiness threshold, the cleaning robot is instructed to return to the base station to clean the cleaning parts of the cleaning robot.
[0047] Based on the cleaned component after cleaning treatment, the first cleaning area is cleaned.
[0048] In one embodiment, the cleaning process for the cleaning components of the cleaning robot includes: drying the cleaning components of the cleaning robot and washing the cleaning components of the cleaning robot.
[0049] In one embodiment, instructing the cleaning robot to return to the base station to clean the cleaning components of the cleaning robot includes:
[0050] Before cleaning the first cleaning area, the cleaning robot is instructed to return to the base station; the base station is used to dry the cleaning parts of the cleaning robot.
[0051] In one embodiment, the cleaning component includes a side brush, a roller brush, and a cloth. Cleaning the first cleaning area based on the cleaned area using the cleaning component includes:
[0052] The side brush is controlled to be in a non-cleaning position, and the roller brush in the cleaning position is controlled to clean the first cleaning area at a first cleaning speed, and the cloth in the cleaning position is controlled to clean at a second cleaning speed.
[0053] In one embodiment, when solid particles are present in the first cleaning area, the first cleaning speed of the roller brush is determined based on the particle density of the solid particles; or, the first cleaning speed of the roller brush is inversely proportional to the second cleaning speed of the side brush.
[0054] In one embodiment, after cleaning the first cleaning area based on the cleaning component that has undergone cleaning treatment, the method further includes:
[0055] Based on the degree of dirtiness of the first cleaned area after the initial cleaning, it is determined whether the cleaning component needs to be cleaned, and the first cleaned area is cleaned repeatedly.
[0056] In one embodiment, the repeated cleaning of the first cleaning area includes:
[0057] The side brush is controlled to be in a non-cleaning position, and the roller brush in the cleaning position is controlled to clean the first cleaning area at a third cleaning speed and the cloth in the cleaning position is controlled to clean at a first cleaning speed.
[0058] In one embodiment, identifying the cleaning area and / or degree of dirt in the first cleaning area, and cleaning the first cleaning area based on the cleaning area and / or degree of dirt, includes:
[0059] When the cleaning area of the first cleaning area is less than the area threshold and / or the degree of dirt is less than the dirt threshold, the first cleaning area is cleaned based on the cleaning component of the cleaning robot.
[0060] Based on the degree of dirtiness of the first cleaned area after the initial cleaning, it is determined whether the cleaning component needs to be cleaned, and the first cleaned area is cleaned repeatedly.
[0061] In one embodiment, the cleaning component includes a side brush, a roller brush, and a cloth. The cleaning component based on the cleaning robot performs cleaning treatment on the first cleaning area, including:
[0062] The side brush is controlled to be in a non-cleaning position, and the roller brush and the cloth in the cleaning position are both controlled to clean the first cleaning area at a second cleaning speed.
[0063] In one embodiment, determining the target cleaning strategy based on the target behavior recognition result includes:
[0064] When the target behavior identification result is sleeping behavior, grooming behavior, or staying behavior, a second cleaning area is marked based on the pet's location;
[0065] Continuously monitor the pet's cleaning progress and / or activity status in the second cleaning area until cleaning of other cleaning areas besides the second cleaning area is completed and / or the pet has left the second cleaning area, then perform cleaning on the second cleaning area.
[0066] In one embodiment, determining the target cleaning strategy based on the target behavior recognition result includes:
[0067] When the target behavior identification result is abnormal behavior, a warning notification message is sent to the user based on the number of times the abnormal behavior occurs and / or the duration of occurrence; and / or,
[0068] When the target behavior identification result is abnormal behavior, a virtual obstacle avoidance zone is constructed based on the pet's location, and an obstacle avoidance cleaning mode is executed to avoid the virtual obstacle avoidance zone.
[0069] In one embodiment, sending a warning notification message to the user based on the number of occurrences and duration of the abnormal behavior includes:
[0070] When the number of occurrences and / or duration of the same abnormal behavior exceed a preset threshold, a warning notification message will be sent to the user.
[0071] In one embodiment, the abnormal behavior corresponds to a reporting priority, and the method further includes:
[0072] When the target behavior identification result contains multiple abnormal behaviors, an early warning notification message is sent to the user based on the abnormal behavior with the highest reporting priority.
[0073] In a second aspect, the present invention also provides a cleaning device applied to a cleaning robot, comprising: a first identification module, a second identification module, a third identification module, and a first determination module, wherein:
[0074] The first recognition module is used to acquire a single frame image and recognize the single frame image based on a preset deep learning algorithm to obtain the initial behavior recognition result.
[0075] The second recognition module is used to perform fusion judgment on the initial behavior recognition results corresponding to each single frame image based on a multi-frame image fusion and discrimination strategy to obtain the first behavior recognition result.
[0076] The third recognition module is used to accumulate each single frame image when the recognition result of the first behavior is an abnormality, and to perform time-series behavior recognition on the accumulated multi-frame images according to the time-series behavior recognition algorithm to determine the recognition result of the second behavior.
[0077] The first determining module is used to determine the target behavior identification result based on the first behavior identification result and the second behavior identification result; the target behavior identification result is used to determine the target cleaning strategy for cleaning treatment.
[0078] Thirdly, this application also provides a cleaning robot, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any one of the first aspects.
[0079] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0080] A single-frame image is acquired, and the single-frame image is recognized based on a preset deep learning algorithm to obtain an initial behavior recognition result;
[0081] Based on a multi-frame image fusion and discrimination strategy, the initial behavior recognition results corresponding to each single-frame image are fused and judged to obtain the first behavior recognition result;
[0082] If the first behavior identification result is an abnormality, each of the single-frame images is accumulated, and the accumulated multi-frame images are subjected to temporal behavior identification according to the temporal behavior identification algorithm to determine the second behavior identification result;
[0083] Based on the first behavior recognition result and the second behavior recognition result, a target behavior recognition result is determined; the target behavior recognition result is used to determine a target cleaning strategy for cleaning treatment.
[0084] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0085] A single-frame image is acquired, and the single-frame image is recognized based on a preset deep learning algorithm to obtain an initial behavior recognition result;
[0086] Based on a multi-frame image fusion and discrimination strategy, the initial behavior recognition results corresponding to each single-frame image are fused and judged to obtain the first behavior recognition result;
[0087] If the first behavior recognition result is an abnormality, each of the single-frame images is accumulated, and the accumulated multi-frame images are subjected to temporal behavior recognition according to the temporal behavior recognition algorithm to determine the second behavior recognition result.
[0088] Based on the first behavior recognition result and the second behavior recognition result, a target behavior recognition result is determined; the target behavior recognition result is used to determine a target cleaning strategy for cleaning treatment.
[0089] The technical solution provided by this invention has the following advantages:
[0090] Initial behavior recognition results are obtained by recognizing single-frame images. The first behavior recognition result is obtained by combining multi-frame image fusion and discrimination strategies. This effectively filters out single-frame misjudgments and improves the stability and reliability of conventional behavior recognition. If no abnormal behavior is detected in the initial behavior recognition, the first behavior recognition result is directly used as the final target behavior recognition result, which can save computing resources and improve computing efficiency. If possible abnormal behavior is detected in the initial behavior recognition result, the second behavior recognition result is determined by accumulating multiple frames of images and using a temporal behavior recognition algorithm, which greatly improves the recognition accuracy of complex abnormal behaviors. Finally, the target behavior recognition result obtained by fusing the two types of results can not only accurately reflect the pet's real behavior state, but also provide a reliable basis for cleaning robots to formulate targeted cleaning strategies. Attached Figure Description
[0091] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0092] Figure 1 is a schematic flowchart of a cleaning method in one embodiment;
[0093] Figure 2 is a flowchart illustrating the steps for obtaining the initial behavior recognition result in one embodiment;
[0094] Figure 3 is a schematic diagram of the internal structure of a cleaning robot in one embodiment;
[0095] Figure 4 is a flowchart illustrating the steps for determining the first behavior recognition result in one embodiment;
[0096] Figure 5 is a flowchart illustrating the steps for determining the second behavior recognition result in one embodiment;
[0097] Figure 6 is a flowchart illustrating the dual-path sampling step in a fast and slow network in one embodiment;
[0098] Figure 7 is a flowchart illustrating the dual-path feature extraction step in a fast and slow network in one embodiment;
[0099] Figure 8 is a flowchart illustrating the steps for obtaining the fused feature vector in one embodiment;
[0100] Figure 9 is a schematic diagram of the fast and slow network processing flow in one embodiment;
[0101] Figure 10 is a flowchart illustrating the steps for determining the second behavior recognition result based on a fast and slow network in one embodiment;
[0102] Figure 11 is a flowchart illustrating the process of determining a target cleaning strategy in one embodiment;
[0103] Figure 12 is a flowchart illustrating the steps of marking a first cleaning area and cleaning the first cleaning area in one embodiment.
[0104] Figure 13 is a flowchart illustrating the steps of cleaning a first cleaning area by identifying the cleaning area and / or degree of dirt in one embodiment.
[0105] Figure 14 is a flowchart illustrating the cleaning steps for the first cleaning area when the cleaning area or degree of dirt exceeds a threshold in one embodiment.
[0106] Figure 15 is a flowchart illustrating the cleaning steps of each cleaning component in a first cleaning area in one embodiment.
[0107] Figure 16 is a flowchart illustrating the process of determining repeated cleaning steps for a first cleaning area in one embodiment.
[0108] Figure 17 is a flowchart illustrating the repeated cleaning process of the first cleaning area by each cleaning component in one embodiment.
[0109] Figure 18 is a flowchart illustrating the cleaning steps for the first cleaning area when the cleaning area or degree of dirt is less than a threshold in one embodiment.
[0110] Figure 19 is a flowchart illustrating the process of cleaning a first cleaning area with a cleaning area or degree of dirt less than a threshold in one embodiment.
[0111] Figure 20 is a flowchart illustrating the cleaning steps for the second cleaning area in one embodiment;
[0112] Figure 21 is a flowchart illustrating the steps of implementing a target cleaning and handling strategy for early warning and avoidance in one embodiment;
[0113] Figure 22 is a flowchart illustrating the steps of sending an early warning notification message to a user in one embodiment;
[0114] Figure 23 is a flowchart illustrating the steps of sending a warning notification message to the user based on the priority of abnormal behavior in one embodiment;
[0115] Figure 24 is a schematic diagram of the cleaning device in one embodiment;
[0116] Figure 25 is a structural schematic diagram of a cleaning robot in one embodiment;
[0117] Figure 26 is a schematic diagram of the cleaning component of a cleaning robot in one embodiment. Detailed Implementation
[0118] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0119] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0120] In this invention, unless otherwise stated, directional terms such as "upper," "lower," "top," and "bottom" are generally used in relation to the direction shown in the accompanying drawings, or in relation to the vertical, perpendicular, or gravitational direction of the component itself; similarly, for ease of understanding and description, "inner" and "outer" refer to the inner and outer contours of each component itself, but the above directional terms are not intended to limit this invention.
[0121] In the daily lives of pet-owning families, cleaning robots have evolved to include a variety of pet-related functions, such as pet video monitoring via front-end cameras, cat and dog recognition for obstacle avoidance, capturing exciting moments of pets, identifying pet activity hotspots, and recognizing pet supplies. However, traditional technologies rely solely on video monitoring as a basic monitoring tool for pet surveillance. This not only requires human control or simultaneous monitoring during operation but also treats pets merely as ordinary obstacles, resulting in low accuracy in recognizing pet behavior. Furthermore, it cannot intelligently analyze pet behavior to adjust cleaning logic, making it difficult to dynamically optimize cleaning strategies and respond promptly to abnormal pet conditions.
[0122] Based on this, in an exemplary embodiment, as shown in FIG1, a cleaning method is provided, which is applied to a cleaning robot, and the method includes:
[0123] Step 101: Acquire a single-frame image and perform recognition on the single-frame image based on a preset deep learning algorithm to obtain the initial behavior recognition result.
[0124] Deep learning algorithms can be implemented based on deep learning models.
[0125] In implementation, for pet-owning households, the cleaning robot's cleaning strategy needs to be tailored to the pet's behavior. Therefore, during the cleaning process, the robot uses image acquisition units and sensors to collect information about the cleaning environment, recognizing pet behavior to set targeted cleaning strategies, ensuring accurate pet monitoring and thorough cleaning of the area. The cleaning robot also includes a processing unit. The image acquisition unit acquires individual frames of the cleaning environment. These frames are then transmitted to the processing unit, which uses a deep learning model to perform image recognition and filtering, identifying frames containing the pet. Based on these pet frames and a pre-defined deep learning algorithm, the robot performs behavior recognition to obtain the initial pet behavior recognition result for that frame.
[0126] Specifically, the image acquisition unit can be, but is not limited to, a monocular camera, a multi-view camera, a depth camera, or a panoramic camera. This disclosure does not limit the type or number of image acquisition units. Taking a monocular camera as an example, the cleaning robot acquires single-frame images of a pet in real time using its onboard monocular camera. To ensure the image quality of each frame and improve the accuracy of pet behavior recognition, the processing unit can preprocess each frame. For example, the image can be cropped to an effective field of view of 800×600, and then Gaussian filtering for noise reduction and adaptive histogram equalization to enhance image contrast, thus obtaining a preprocessed image. The processing unit then inputs the preprocessed single-frame image into a pre-trained deep learning model, which performs recognition processing on the single-frame image, ultimately outputting the initial behavior recognition result for that single-frame image.
[0127] The deep learning model adopts a two-stage architecture of "object detection" and "behavior classification". It is pre-trained using labeled training samples. The trained deep learning model can identify and judge one or more pet behaviors. Optionally, the deep learning model can be a YOLO model. Using the YOLOv8 algorithm, the pet target in a single frame image can be located, and the pet in the image can be labeled, outputting the bounding box coordinates (x1, y1, x2, y2) of the labeled pet. Furthermore, the YOLO model can identify pet species using the YOLOv8 algorithm and perform pet behavior recognition within the bounding box area. For example, during model training, nine pet behaviors are labeled in the training samples for recognition training. The fully connected layer of the trained YOLO model can output the probability distribution of these nine pet behaviors, including the probability values corresponding to normal (pet staying behavior), drinking, eating, sleeping, licking fur, scratching, squatting, vomiting, and excretion. A behavior recognition probability threshold is set for each single frame image, and the initial behavior recognition result is determined based on this threshold. For example, the behavior category with the highest probability (≥0.7) is taken as the initial behavior recognition result.
[0128] Optionally, if the highest probability value among the probability values corresponding to each behavior is less than the behavior recognition probability threshold, i.e., the highest probability value is <0.7, then the processing unit of the cleaning robot marks the single frame image as "unrecognized".
[0129] Step 102: Based on the multi-frame image fusion and discrimination strategy, the initial behavior recognition results corresponding to each single frame image are fused and judged to obtain the first behavior recognition result;
[0130] In implementation, to ensure the accuracy of pet behavior recognition results, the cleaning robot is pre-configured with a multi-frame image fusion and discrimination strategy. Therefore, when performing pet behavior recognition, the cleaning robot uses this multi-frame image fusion and discrimination strategy to fuse and judge the initial behavior recognition results corresponding to multiple consecutive single-frame images to obtain the first behavior recognition result. Specifically, the cleaning robot's processing unit can use a sliding window mechanism to fuse and judge the initial recognition results corresponding to three consecutive image frames. Centered on the current frame, it selects the previous and next frames to form a window containing three images, with a time span of 0.1 seconds for fusion judgment. The essence of this fusion judgment is a consistency judgment of the initial behavior recognition results. If the initial behavior recognition results corresponding to the three frames are completely consistent and all are valid behaviors (i.e., there are no unrecognized behaviors), then the initial behavior recognition result is output as the first behavior recognition result. If at least one or more frames have inconsistent initial behavior recognition results, or if there are unrecognized valid behaviors, then a "default value" is output as the first behavior recognition result.
[0131] Optionally, for various pet behaviors, such as "excretion" and "sleeping" which usually last for more than 5 seconds, a time decay mechanism can be set. That is, if a certain behavior is output in the first 3 frames and not detected in the next 2 frames, the behavior label will still be retained for 2 seconds in the next 2 frames to avoid misjudgment of behavior due to short interruptions.
[0132] In one optional embodiment, the initial behavior recognition result identified in a single frame image represents the pet's routine behavior, such as the pet staying, drinking water, eating, sleeping, licking its fur, excreting, etc., and the cleaning robot fuses and judges multiple initial behavior recognition results based on a continuous multi-frame image fusion discrimination strategy, and outputs a consistent first behavior recognition result. Then, the first behavior recognition result can be used as the target behavior recognition result for the pet behavior recognition, without triggering the time-series behavior recognition process, saving the processing flow of pet behavior recognition and improving the efficiency of pet behavior recognition.
[0133] Step 103: If the first behavior identification result is an abnormality, accumulate each single frame image, and perform temporal behavior identification on the accumulated multi-frame images according to the temporal behavior identification algorithm to determine the second behavior identification result.
[0134] In implementation, after fusing the initial behavior recognition results from multiple consecutive single-frame images to obtain the first behavior recognition result, if the pet behavior output by this first behavior recognition result represents a simple behavior, such as eating, drinking, or grooming, then this first behavior recognition result is used as the final target behavior recognition result for the pet to improve computational efficiency and save computational resources. However, if the first behavior recognition result is abnormal, it indicates that relying solely on static single-frame images, or using multiple single-frame images for pet behavior recognition, may be inaccurate. Therefore, a temporal behavior recognition algorithm needs to be introduced to comprehensively judge the pet behavior. Abnormalities in the first behavior recognition result include two situations: 1. The first behavior recognition result is a default value; 2. The first behavior recognition result indicates abnormal pet behavior. For both situations, it is necessary to enable the detection of temporal behavior recognition. In the first situation, where the first behavior recognition result is a default value, it proves that the pet behavior corresponding to the three consecutive frames in the fusion judgment is inconsistent, and it may not be possible to identify the same type of behavior in all three consecutive frames, thus leading to an inability to obtain accurate pet behavior. In scenario two, the initial recognition results of the first behavior identification result for three consecutive frames show abnormal behaviors such as "scratching" and "vomiting." Abnormal pet behaviors often involve complex and continuous actions, indicating that complex pet behaviors are difficult to identify. Relying solely on a single frame may not be sufficient to identify this type of pet behavior. Therefore, a temporal behavior recognition process is initiated. The cleaning robot accumulates each single frame image and performs temporal behavior recognition on the accumulated multi-frame images using a temporal behavior recognition algorithm to determine the second behavior identification result. Specifically, starting from the image frame where the abnormal behavior was first detected, the processing unit continuously accumulates 64 frames, spanning approximately 2.1 seconds. These frames are uniformly scaled to a resolution of 256×340 and data augmented. Subsequently, an optimized temporal behavior recognition algorithm is used to perform temporal behavior recognition on the accumulated multi-frame images to determine the second behavior identification result. The temporal behavior recognition algorithm can be implemented through a temporal behavior recognition model, which may include, but is not limited to, SlowFast networks, Temporal Segment Networks (TSNs: Towards good practices for deep action recognition), etc. This disclosure does not limit the specific implementation of this model. The detailed processing procedure of the temporal behavior recognition model will be described in detail in the following embodiments of this application, and will not be elaborated upon further here.
[0135] Step 104: Determine the target behavior recognition result based on the first behavior recognition result and the second behavior recognition result.
[0136] In implementation, for cases where abnormal pet behavior is identified in the initial behavior recognition results, the first behavior recognition result obtained by fusing multiple frames of images is combined with the second behavior recognition result determined by the temporal behavior recognition algorithm to determine the final target behavior recognition result. Specifically, when the first behavior recognition result labels a single frame image with abnormal behavior tags such as "scratching," "vomiting," or "hen squatting," or when an abnormal behavior is detected during the fusion judgment of multiple initial behavior recognition results but fails to pass the 3-frame fusion, the temporal recognition process is triggered to determine the second behavior recognition result. In this way, when the cleaning robot combines the first and second behavior recognition results to determine the final target behavior recognition result, it follows a priority rule: if the second behavior recognition result is "scratching," "vomiting," or "hen squatting" with a probability ≥ 0.75, it is directly taken as the target result because abnormal behavior takes priority; if the second behavior recognition result is "unrecognized" or normal behavior, the first behavior recognition result that satisfies the requirement of 3-frame fusion is used; if there is a conflict between the first behavior recognition result being a normal behavior and the second behavior recognition result being an abnormal behavior, the probability of the two is compared to determine the outcome. When the probability of abnormal behavior is ≥ the probability of normal behavior, the second behavior recognition result is used; otherwise, the first behavior recognition result is used. The final output may include, but is not limited to, the behavior category, the confidence level in the range of 0-1, the behavior occurrence timestamp, and the cleaning robot's localization coordinates based on the SLAM (Simultaneous Localization and Mapping) map, providing data support for subsequent motion planning.
[0137] In this embodiment, an initial behavior recognition result is obtained through single-frame image recognition, and a first behavior recognition result is obtained by combining a multi-frame image fusion discrimination strategy. This effectively filters out single-frame misjudgments and improves the stability and reliability of conventional behavior recognition. At the same time, for possible abnormal behaviors detected in the initial behavior recognition result, a second behavior recognition result is determined by accumulating multiple frames of images and using a temporal behavior recognition algorithm, which greatly improves the recognition accuracy of complex abnormal behaviors. Finally, the target behavior recognition result obtained by fusing the two types of results can not only accurately reflect the pet's real behavior state, but also provide a reliable basis for the cleaning robot to formulate targeted cleaning strategies, making the cleaning process more intelligent and more in line with actual needs.
[0138] In an exemplary embodiment, as shown in FIG2, prior to step 101, the method further includes:
[0139] Step 201: Acquire a single-frame image of the pet activity scene through the image acquisition unit.
[0140] As shown in Figure 3, which is a schematic diagram of the internal structure of the cleaning robot, the robot includes an image acquisition unit. This unit can capture real-time images of the cleaning environment. For pet-owning households, where pets are active in the cleaning environment, images of these activities can be further captured during the cleaning process. Specifically, the image acquisition unit of the cleaning robot can be a monocular camera, a depth camera, a low-light camera, etc. This embodiment does not limit the type of device used in the image acquisition unit of the cleaning robot; the choice is based on factors such as lighting conditions and field of view requirements in the actual application scenario. During the acquisition process, as the cleaning robot moves, the image acquisition unit continuously captures images of the pet's activity at a preset frame rate, capturing one frame at fixed intervals as a single frame of the pet's activity scene. Simultaneously, to ensure the accuracy of subsequent recognition, the acquired images must contain clear information about the pet and its surrounding environment, avoiding the loss of pet features due to excessive occlusion, blurring, or insufficient lighting. After acquisition, the images are temporarily stored in the cleaning robot's local cache for further processing.
[0141] Step 202: Input the single-frame image into the pre-trained deep learning model, and perform feature extraction and pet behavior category prediction on the single-frame image based on the deep learning model, and output the probability value of each type of pet behavior.
[0142] In practice, the cleaning robot is pre-integrated with a trained deep learning model. The cleaning robot can input the acquired single-frame images into the pre-trained deep learning model, and perform feature extraction and pet behavior category prediction based on the deep learning model. Finally, it outputs the probability values of various pet behaviors.
[0143] Specifically, the cleaning robot first preprocesses the input single-frame image, including adjusting its size to the model's required specifications (e.g., 224×224 pixels), normalizing it, and mapping pixel values to the 0-1 range to adapt to the model's input requirements. Then, the preprocessed image is input into the trained deep learning model. The preprocessed image is further processed: first, the object detection network locates the pet's position and crops the pet area; then, the behavior classification network extracts features from the pet area, capturing key features such as the pet's limb posture, movement contours, and fur texture. Finally, based on these features, the robot predicts the pet's behavior category and outputs the probability values for various preset pet behaviors (e.g., normal (staying), drinking water, eating, scratching, vomiting, etc.), with the sum of the probability values for all behavior categories being 1.
[0144] The pre-trained deep learning model can use a two-stage architecture of "YOLOv8 object detection + EfficientNet-B4 behavior classification" to process the pre-processed single-frame image, or it can use Faster RCNN network (Faster Region-based Convolutional Neural Networks, a faster region-based convolutional neural network) to identify pet behavior. This disclosure does not limit the specific implementation of the model.
[0145] Step 203: Obtain the initial behavior recognition results based on the probability values of various pet behaviors.
[0146] In practice, based on the probability values of various pet behaviors output by the deep learning model, the cleaning robot determines the pet behavior corresponding to the highest probability value among the various pet behavior probability values as the initial behavior recognition result.
[0147] Specifically, after obtaining the probability values of various pet behaviors, the processing unit of the cleaning robot analyzes and judges them, selecting the behavior category with the highest probability value as the candidate result. If the highest probability value is greater than or equal to the preset behavior recognition probability threshold, then the behavior category is determined as the initial behavior recognition result; if the highest probability value is less than the preset threshold, then it is considered that the current image cannot accurately identify the pet behavior, and the initial behavior recognition result is marked as "unrecognized". For example, if the deep learning model outputs a probability of "drinking water" of 0.85, and this value is the highest among all behavior categories and greater than the preset behavior recognition probability threshold of 0.7, then the initial behavior recognition result is "drinking water"; if the highest probability value is 0.65 and less than the preset behavior recognition probability threshold of 0.7, then the initial behavior recognition result is "unrecognized".
[0148] In this embodiment, a single-frame image of a pet's activity scene is accurately acquired through an image acquisition unit. Then, a pre-trained deep learning model is used to extract features and predict behavior categories from the image. With the powerful learning ability of the model, key features such as the pet's limb posture and movement details can be fully captured, and the probability values of various behaviors can be accurately output. Finally, the initial behavior recognition result is determined based on the probability values, which improves the real-time performance and accuracy of pet behavior recognition.
[0149] In an exemplary embodiment, as shown in FIG4, the specific processing procedure of step 102 includes:
[0150] Step 401: If the initial behavior recognition results corresponding to multiple consecutive single-frame images are all the same, then the initial behavior recognition result is determined as the first behavior recognition result.
[0151] In practice, to improve the accuracy of pet behavior recognition, after initial behavior recognition for a single frame image, the cleaning robot can also perform a fusion judgment on the initial behavior recognition results corresponding to multiple consecutive single frame images based on a multi-frame image fusion and discrimination strategy. This is also known as a consistency judgment. If all the consecutive initial behavior recognition results are the same, the cleaning robot will determine the pet behavior corresponding to the initial behavior recognition result as the first behavior recognition result.
[0152] Specifically, if three consecutive single-frame images with a time span of approximately 0.1 seconds correspond to the same initial behavior recognition result, and all are valid behaviors, the cleaning robot directly determines this initial behavior recognition result as the first behavior recognition result. For example, if the initial recognition result of three consecutive frames is the pet "sleeping" behavior, and the behavior probability value of each frame meets the preset behavior recognition probability threshold (e.g., ≥0.7), it indicates that the behavior has high stability and reliability, so "sleeping" is taken as the current first behavior recognition result.
[0153] Step 402: If the initial behavior recognition results corresponding to multiple consecutive single-frame images are not all the same, then the first behavior recognition result is determined as the default value.
[0154] In practice, if the initial behavior recognition results corresponding to multiple consecutive single-frame images are not all the same, and inconsistencies exist, the cleaning robot will determine the first behavior recognition result as a default value or as a null value. This indicates that the fusion judgment of multiple frames of images did not output a valid result. For example, if the initial behavior recognition results corresponding to three consecutive frames with a time span of 0.1 seconds are used for fusion judgment, and if the initial behavior recognition results corresponding to two frames are "drinking water" and the initial behavior recognition result corresponding to one frame is "running," and the initial behavior recognition results of the three frames are not completely the same, then the first behavior recognition result will be determined as a default value.
[0155] In this embodiment, the first behavior recognition result is determined by performing a consistency judgment on the initial behavior recognition results of multiple consecutive single-frame images. When all results are identical, the consistent initial behavior recognition result is determined as the first behavior recognition result, which effectively utilizes the behavioral continuity of consecutive frames to ensure the stability and reliability of the recognition results. When the initial behavior recognition results are not all identical, the first behavior recognition result is set to a default value, which avoids the adoption of erroneous results caused by single-frame misjudgment, instantaneous action interference, etc., and reduces recognition errors. This processing method not only fully leverages the advantages of multi-frame image fusion and filters unstable recognition information, but also lays a reliable foundation for subsequent determination of the final target behavior by combining temporal recognition results, thereby improving the accuracy and rigor of the overall pet behavior recognition.
[0156] In an exemplary embodiment, if abnormal pet behavior is detected in the initial behavior recognition results, to avoid the complexity of the abnormal behavior, relying solely on a single frame image is insufficient for accurate identification. Therefore, when abnormal behavior occurs, as shown in Figure 5, the specific processing procedure of step 103 includes:
[0157] Step 501: Based on the preset time sequence, accumulate each single frame image within the target duration to obtain a multi-frame image sequence.
[0158] In implementation, the cleaning robot accumulates multiple single-frame images of a target duration based on a preset time sequence. This allows the robot to obtain a video segment of the target duration. Further, image preprocessing is performed on this video segment, including spatial adjustment and scaling of each frame to a fixed resolution, such as 224x224. Data augmentation methods are then applied, such as random cropping and horizontal flipping, to enhance the data of each frame, thereby improving the model's generalization ability. Finally, based on the preprocessed video segment and the sampling requirements of the temporal behavior recognition model, data sampling is performed, enabling the temporal behavior recognition model to process multi-frame image sequences.
[0159] Specifically, the target duration for accumulating multiple frames of images can be preset based on the normal duration of abnormal behavior, and this embodiment does not limit this. For behaviors such as "scratching" and "vomiting," it is set to approximately 2.1 seconds, corresponding to 64 frames of images captured at a frame rate of 30fps. During the accumulation process, the single-frame images are arranged strictly in the chronological order of image acquisition to ensure that the multi-frame image sequence can completely reflect the temporal development process of the pet's behavior. At the same time, timestamp information is added to each frame of image to provide a temporal reference for subsequent temporal correlation analysis. After accumulation, the multi-frame image sequence is uniformly stored in video clip format for easy input into the temporal behavior recognition model for processing.
[0160] Step 502: Spatial features are extracted from the multi-frame image sequence using a temporal behavior recognition model to obtain feature vectors.
[0161] In implementation, the cleaning robot uses a temporal behavior recognition model to extract spatial features from a multi-frame image sequence, obtaining feature vectors. The temporal behavior recognition model can be, but is not limited to, an optimized SlowFast network, a TSN time-segmentation network, etc. Thus, the feature extraction network in the temporal behavior recognition model of the cleaning robot performs deep feature mining on the pet region contained in each frame image, extracting high-dimensional spatial features containing information such as the pet's limb posture, movement contours, and fur texture. These features are encoded to form corresponding feature vectors, and each feature vector can accurately represent the spatial feature attributes of the pet in a single frame image.
[0162] Step 503: Based on temporal correlation, the feature vectors are fused to obtain a fused feature vector.
[0163] In implementation, the temporal behavior recognition model in the cleaning robot fuses the obtained feature vectors based on the temporal correlation between multiple frames of images, resulting in a fused feature vector. During feature fusion, the temporal behavior recognition model utilizes temporal modeling mechanisms to analyze the temporal correlation of feature vectors in the multi-frame image sequence, capturing the dynamic changes in pet behavior over time. For example, for the behavior of "vomiting," the temporal behavior recognition model focuses on the correlation between the "head down, back arched" feature vectors in the previous frames and the "body twitching" feature vectors in subsequent frames. It integrates these temporally correlated feature vectors into a single fused feature vector through one or more fusion methods, such as weighting, concatenation, or direct addition. This fused feature vector not only contains the spatial feature information of each individual frame but also encapsulates the dynamic changes in pet behavior over time, thus providing a more comprehensive representation of the pet's complete behavioral process.
[0164] Step 504: Perform pet behavior recognition on the fused feature vector to determine the second behavior recognition result.
[0165] In implementation, the temporal behavior recognition model in the cleaning robot uses the fused feature vector to identify pet behaviors, thereby determining the second behavior recognition result. Specifically, the cleaning robot inputs the fused feature vector into the classification layer (e.g., a fully connected layer) of the temporal behavior recognition model. The classification layer predicts the pet behavior category based on the spatial and temporal features contained in the fused feature vector and outputs the probability value of each pet behavior category. Then, the temporal behavior recognition model selects the behavior category with the highest probability that meets the behavior recognition probability threshold requirement according to a preset behavior recognition probability threshold (e.g., the current behavior recognition probability threshold is 0.75), and determines it as the second behavior recognition result.
[0166] In this embodiment, single-frame images within a target duration are accumulated in a preset time sequence to form a multi-frame image sequence. Feature vectors are then extracted using a temporal behavior recognition model. The resulting fused feature vector, obtained by fusing the feature vectors based on temporal correlation, contains both spatial information and integrates dynamic changes in the temporal dimension, comprehensively representing the complete process of pet behavior. Finally, the fused feature vector is used to identify the second behavior recognition result, significantly improving the accuracy of pet abnormal behavior recognition and effectively compensating for the limitations of single-frame recognition and simple multi-frame fusion.
[0167] In an exemplary embodiment, taking a fast-slow network as an example of a temporal behavior recognition model, the process of determining the second behavior recognition result is described as shown in Figure 6. The specific processing steps of step 501 include:
[0168] Step 601: Based on the preset time sequence, accumulate each single frame image within the target duration to obtain a video segment.
[0169] In practice, the cleaning robot accumulates individual frames of images within a preset time sequence for a target duration to obtain a video segment. The target duration is set based on the typical and persistent characteristics of the pet's abnormal behavior. This process has been described in step 401 of the above embodiment and will not be repeated here.
[0170] Optionally, during the image frame accumulation process, single-frame images are arranged strictly according to the timestamp order of image acquisition to ensure that the video segment can fully reflect the time process of the pet's behavior from occurrence to development. At the same time, each frame image is initially verified, for example, frames that are severely blurry or where the pet has completely moved out of the frame are removed, ultimately forming a continuous and complete video segment.
[0171] Step 602: Input the video segment into the fast-slow network, and perform dual-path frame sampling through the fast-slow network to obtain the first keyframe image sequence corresponding to the slow path and the second keyframe image sequence corresponding to the fast path.
[0172] In practice, after obtaining the video segment, the cleaning robot inputs it into a slow-fast network (e.g., a SlowFast network). The network performs dual-path frame sampling to obtain the first keyframe image sequence corresponding to the slow path and the second keyframe image sequence corresponding to the fast path. Specifically, the slow path in the network has high spatial resolution and low temporal resolution, while the fast path has low spatial channels and high temporal resolution. In this way, the slow path in the fast-slow network samples the video segment with a larger time step (e.g., τ=16), selecting one frame every 16 frames, extracting 4 key images from the 64-frame video segment to form the first keyframe image sequence. This sequence has a large time span and focuses on capturing the static spatial features of pet behavior, such as limb posture and body shape. The fast path in the fast-slow network samples with a smaller time step (e.g., β=2), selecting one frame every 2 frames, extracting 32 key images from the 64-frame video segment to form the second keyframe image sequence. This sequence has high temporal resolution and focuses on capturing the dynamic changes in pet behavior, such as the amplitude of movements and movement trajectories. In this way, through this dual-path sampling method, two sets of keyframe image sequences are obtained, which can preserve the overall spatiotemporal context of the behavior while also taking into account detailed motion information.
[0173] In this embodiment, single-frame images within a target duration are accumulated in a preset time sequence to form a video segment, which fully preserves the continuous change process of pet behavior in the time dimension. The video segment is then input into a fast and slow network for dual-path frame sampling. The first keyframe image sequence obtained by the slow path can capture the overall spatial features and long-term change trend of the behavior, while the second keyframe image sequence obtained by the fast path can accurately capture the rapid and dynamic details. This dual-path sampling method takes into account both the static and dynamic features of the behavior, which reduces the amount of redundant data processing to improve efficiency, while fully preserving key information, laying the foundation for subsequent accurate extraction and fusion of features and improving the accuracy of complex behavior recognition.
[0174] In an exemplary embodiment, as shown in FIG7, the specific processing procedure of step 502 includes:
[0175] Step 701: Perform dual-path feature extraction on the first keyframe image sequence and the second keyframe image sequence using a fast and slow network to obtain the static feature vector corresponding to the slow path and the dynamic feature vector corresponding to the fast path.
[0176] In implementation, dual-path feature extraction is performed on the first and second keyframe image sequences using a fast and slow network to obtain the static feature vector corresponding to the slow path and the dynamic feature vector corresponding to the fast path. Specifically, the initial convolutional layer of the slow path uses a 1×7×7 3D convolution (time×height×width) with a stride of (1,2,2) to process the first keyframe image sequence. For example, it can be obtained by sampling 4 frames from a continuous 64-frame video segment at τ=16. A deep 3D convolutional neural network, such as 3D ResNet, is used to gradually extract spatial detail features from the image through multiple residual stages, including static information such as the pet's limb structure, torso posture, and facial expressions. Finally, these features are encoded into a fixed-dimensional static feature vector to represent the stable spatial attributes of the behavior. The 3D ResNet (three-dimensional residual network) composed of residual block structures of the slow path is divided into multiple continuous stages according to function. Taking ResNet-50 (a network with 50 trainable layers) as an example, it is typically divided into conv1 (initial convolutional layer), conv2_x to conv5_x (four residual stages). Each stage consists of several stacked residual blocks, and the output feature map size of each stage gradually decreases while the number of channels gradually increases. For example, the number of channels changes from 64 to 256 to 512 to 1024, and the spatial resolution decreases. The initial convolutional layer of the fast path uses a 5×1×1 3D convolution with a stride of (1,2,2), and the number of channels is only 1 / 8 of that of the slow path. Lightweight residual blocks are also used. Thus, for the second keyframe image sequence, for example, 32 frames sampled with β=2, a lightweight 3D convolutional structure is used to capture the dynamic changes of pet behavior at a higher temporal resolution, such as rapid paw swings, body twitching amplitude, and continuous movement trajectories. After feature compression, a dynamic feature vector is formed, and its number of channels is usually 1 / 8 of that of the slow path to balance the computational cost. This dual-path parallel extraction method can not only deeply explore the static basic features of behavior, but also accurately capture the dynamic change patterns, providing comprehensive feature support for subsequent feature fusion.
[0177] In this embodiment, a fast and slow network is used to perform dual-path feature extraction on the first and second keyframe image sequences. The slow path focuses on extracting static features such as the pet's limb structure and torso posture from keyframes with a large time span, forming static feature vectors that can accurately capture the stable spatial attributes of behavior. The fast path, on the other hand, extracts dynamic features such as paw swings and body twitches from keyframes with high temporal resolution, forming dynamic feature vectors that can meticulously depict the instantaneous movement changes of behavior. This dual-path parallel extraction method not only preserves the basic static information of behavior but also captures the details of dynamic changes, achieving multi-dimensional and comprehensive feature representation of pet behavior and effectively improving the ability to identify complex behaviors such as abnormal behaviors.
[0178] In an exemplary embodiment, as shown in FIG8, the specific processing procedure of step 503 includes:
[0179] Step 801: Adjust the number of channels in the dynamic feature vector corresponding to the fast path, and align the adjusted dynamic feature vector with the static feature vector corresponding to the slow path.
[0180] In implementation, the cleaning robot adjusts the number of channels in the dynamic feature vector corresponding to the fast path using a fast-slow network, and then aligns the channels of the static feature vector corresponding to the adjusted slow path. Specifically, due to the different network structures of the fast and slow paths, the extracted dynamic and static feature vectors typically have different numbers of channels; for example, a dynamic feature vector might have 64 channels, while a static feature vector has 1024 channels. Therefore, a 1×1 convolutional layer is needed to adjust the channel dimension of the dynamic feature vector, mapping its channel number to match that of the static feature vector. For example, the number of channels is gradually increased from 64 to 256, then further to 512, and finally to 1028. Channel alignment is then performed based on the same number of channels. After alignment, the two types of feature vectors have the same dimensionality in terms of channels, laying a structural foundation for subsequent feature fusion and ensuring that dynamic and static features can be effectively integrated within the same dimensional space.
[0181] Step 802: The fast path is dimensionality-reduced by using a time-averaged pooling layer or convolutional layer in the fast and slow networks, so that the dynamic feature vector of the dimensionality-reduced fast path is time-aligned with the static feature vector of the corresponding slow path.
[0182] In implementation, the fast path is dimensionality-reduced by using temporal average pooling layers or convolutional layers in both the fast and slow networks. This ensures that the dynamic feature vectors of the dimensionality-reduced fast path are temporally aligned with the corresponding static feature vectors of the slow path. Since the fast path samples more keyframes (32 frames, as shown in Figure 9), its dynamic feature vector typically has a larger temporal dimension than the slow path (e.g., 4 frames). Therefore, a temporal average pooling layer or a temporal convolutional operation with a stride of 4 is needed to reduce the temporal dimension of the dynamic feature vectors from 32 to 4, aligning it with the temporal dimension of the static feature vectors. This achieves temporal alignment between the dynamic and static feature vectors. Temporal alignment ensures the correspondence between the two types of feature vectors on the time axis, avoiding feature misalignment due to differences in time scales, and ensuring that the fused features accurately reflect the correlation between static and dynamic attributes at the same time point.
[0183] Step 803: Perform feature fusion on the channel-aligned and time-aligned dynamic feature vectors and static feature vectors to obtain a fused feature vector.
[0184] In implementation, the cleaning robot fuses the dynamic and static feature vectors from channel alignment and time alignment to obtain a fused feature vector. The specific fusion location can be performed after each residual block, achieving multi-level feature interaction. Specifically, after completing the dual alignment of the channel and time dimensions, feature fusion is performed using element-wise addition or channel concatenation.
[0185] Feature fusion can be achieved through two main methods: element-wise addition, which strengthens key consistent information between the two types of features through weighted fusion (e.g., superimposing the static feature of "arched back posture" related to "vomiting" in pet abnormal behavior with the dynamic feature of "body twitching"); and channel concatenation, which preserves the complete information of both types of features and forms a higher-dimensional feature vector (e.g., 1024 + 1024 = 2048 channels). The fused feature vector contains both static spatial and dynamic temporal features of pet behavior, comprehensively representing the overall attributes and detailed changes of the behavior, providing richer and more robust feature inputs for subsequent behavior recognition.
[0186] In this embodiment, the number of channels in the fast path dynamic feature vector is adjusted to achieve channel alignment; then, the fast path is dimensionality-reduced and sampled using a temporal average pooling layer or a convolutional layer to achieve temporal alignment; finally, the double-aligned feature vectors are fused, which retains stable spatial information such as pet limb posture in static features and integrates instantaneous motion information such as action changes in dynamic features, enabling the fused feature vector to comprehensively and accurately represent the spatiotemporal attributes of pet behavior. This improves the accuracy of pet behavior recognition.
[0187] In an exemplary embodiment, as shown in FIG10, the specific processing procedure of step 504 includes:
[0188] Step 1001: Use the fully connected layer in the fast and slow network to perform pet behavior recognition on the fused feature vector and output the probability values of various pet behaviors.
[0189] In implementation, after feature vector fusion, the cleaning robot uses a fully connected layer and a softmax activation function layer in a fast-slow network to recognize pet behaviors from the fused feature vector, outputting probability values for various pet behaviors. Specifically, after obtaining the fused feature vector, it is input into the fully connected layer of the fast-slow network. This fully connected layer performs non-linear transformation and dimensionality compression on the fused feature vector, mapping the high-dimensional fused features to a preset pet behavior category space. The number of neurons in the fully connected layer is the same as the number of behavior categories, with each neuron corresponding to a score for one behavior category. After processing by the softmax activation function, these scores are converted into probability values for each category, and the sum of the probability values for all categories is 1. For example, if the fused feature vector contains typical spatiotemporal features of the "vomiting" behavior, the fully connected layer will output a high probability value for the "vomiting" category while reducing the probability values of other irrelevant categories, thereby quantifying the degree of matching between the fused features and various behaviors.
[0190] Step 1002: Determine the second behavior recognition result based on the probability values of various pet behaviors.
[0191] In implementation, after the probability values of various pet behaviors are output by the fully connected layer and activation function layer of the fast and slow network, the classification layer in the cleaning robot's fast and slow network determines the second behavior recognition result based on the probability values of various pet behaviors. Specifically, after obtaining the probability values of various pet behaviors, the behavior category with the highest probability is first selected, and then it is determined whether the highest probability value reaches the preset behavior recognition probability threshold (usually set to ≥0.75 for abnormal behaviors to reduce false alarms). If the highest probability value meets the requirement of the behavior recognition probability threshold, the corresponding behavior category is determined as the second behavior recognition result; if the highest probability value is lower than the behavior recognition probability threshold, or the probability distribution of all categories is relatively even (no obviously dominant category), the second behavior recognition result is marked as "unrecognized". For example, when the probability value of "hen squatting" is 0.82 ≥ the behavior recognition probability threshold of 0.75 and is the highest value, the second behavior recognition result is "hen squatting".
[0192] In this embodiment, a fully connected layer of a fast-slow network is used to identify pet behaviors from the fused feature vector and output probability values for various behaviors. The second behavior identification result is then determined based on these probability values. By setting a reasonable threshold, the most probable behavior category is selected, ensuring both accuracy and reducing false alarms. The overall process fully utilizes the spatiotemporal information contained in the fused feature vector, improving the accuracy of identifying complex pet behaviors.
[0193] In an optional embodiment, the second behavior recognition result can also be determined using a Time Period Network (TSN). Specifically, a Time Period Network is pre-trained in the cleaning robot. The accumulated multi-frame images are input into the pre-trained Time Period Network, which divides the accumulated multi-frame images into several segments according to a preset time period. Keyframe features are extracted from each segment. Subsequently, the spatial branch of the Time Period Network extracts static features from each keyframe, while the temporal branch captures the temporal evolution of the behavior by fusing the inter-frame dynamic information of different segments. Next, pooling operations are performed using the pooling layer of the Time Period Network to aggregate the feature vectors of all segments, obtaining a comprehensive feature representation of the entire behavior sequence, i.e., a fused feature vector. Finally, after classifier processing, the second behavior recognition result is determined by combining a multi-segment voting strategy, i.e., selecting the behavior category with the second highest confidence and satisfying temporal logical coherence from all possible behavior categories. This embodiment does not limit the specific processing method for determining the second behavior recognition result.
[0194] In one exemplary embodiment, based on the recognition of pet behavior, the cleaning robot provides a precise cleaning strategy, thereby achieving deep cleaning of the environment and improving the cleaning coverage and efficiency for pet-owning households, as shown in Figure 11. The method further includes:
[0195] Step 1101: Determine the target cleaning strategy based on the target behavior recognition results.
[0196] The target cleaning strategy includes one or more of cleaning, avoidance, and early warning.
[0197] In implementation, the cleaning robot determines the target cleaning strategy based on the target behavior recognition results. Specifically, the cleaning robot dynamically matches the corresponding processing strategy based on the pre-obtained target behavior recognition results and the scene characteristics corresponding to the behavior. For example, when the target behavior recognition result is "the pet is drinking water and has left the drinking area," the cleaning robot will trigger a targeted cleaning strategy for the pet's drinking behavior, calling a refined mopping and sweeping scheme for the wastewater area to clean the pet's drinking area; if it recognizes "the pet is eating," it will prioritize the "avoidance" strategy, marking the eating area as a to-be-cleaned area and bypassing it to avoid disturbing the pet's eating. Furthermore, this target cleaning strategy also includes a function to warn the user. Therefore, when it recognizes "the pet is frequently scratching and has reached a threshold," it will activate the "warning" strategy, sending an abnormal pet reminder to the user via the APP; when it recognizes "the pet is vomiting," it will send a manual cleaning reminder to the user via the APP, while pausing the automatic cleaning of that area to prevent secondary contamination. For complex scenarios, such as a pet scratching itself excessively while drinking water, the cleaning robot can integrate multiple cleaning strategies. It can first avoid disturbing the pet, and then, after the pet leaves, combine cleaning and warning strategies to both treat wastewater and remind the user to pay attention to the pet's health, ensuring the comprehensiveness and adaptability of the strategy. The following embodiments of this application will provide specific examples of the cleaning strategy execution for different target behavior recognition results, which will not be elaborated upon here.
[0198] In this embodiment, based on the target behavior recognition results, a target cleaning strategy including one or more of cleaning, avoidance, and early warning is determined, enabling precise adaptation between the cleaning robot and pet behavior. This ensures cleaning efficiency while demonstrating pet-friendly care. Furthermore, the flexible combination of multiple strategies can address diverse needs in complex scenarios, reduce ineffective cleaning and misoperation, and improve the overall cleaning effect.
[0199] In an exemplary embodiment, as shown in FIG12, the specific processing procedure of step 1101 includes:
[0200] Step 1201: When the target behavior identification result is eating behavior, a virtual eating area is constructed based on the pet's location or the location of the food items, and the virtual eating area is marked as the first cleaning area.
[0201] In implementation, when the target behavior is identified as a pet's eating behavior (i.e., the pet is drinking or eating in the area), the cleaning robot constructs a virtual eating area based on the pet's location or the location of food items (e.g., water bowl, food bowl) to avoid disturbing the pet's normal eating. This virtual eating area is then designated as the first cleaning area to be cleaned. Specifically, the cleaning robot uses its sensors, such as an RGB camera or infrared sensor, to accurately locate the pet's current position or the placement of food items like the water bowl and food bowl. Centered on this location and based on the pet's size (e.g., a radius of 50cm for small dogs and 80cm for large dogs) or the typical movement range of the food items (water bowl, food bowl), it automatically delineates a virtual rectangular or circular area as the virtual eating area and designates it as the first cleaning area. Simultaneously, the cleaning robot stores the boundary information and center coordinates of this first cleaning area in its map module, marking it as a "delayed cleaning area" on its internal map and setting its cleaning priority to ensure accurate identification of this area during subsequent cleaning path planning.
[0202] Step 1202: Clean the other cleaning areas besides the first cleaning area, and continuously monitor the cleaning progress and the pet's activity status in the first cleaning area.
[0203] During implementation, since the pet was drinking or eating in the primary cleaning area, the cleaning robot, based on the refined mopping and sweeping plan planned according to the target cleaning strategy, selected to clean other cleaning areas besides the primary cleaning area first, i.e., the area surrounding the primary cleaning area, to avoid disturbing the pet's eating. Simultaneously, the cleaning robot continuously monitored the cleaning progress of the areas surrounding the primary cleaning area, as well as the pet's activity status in the primary cleaning area, to issue movement planning instructions for the next cleaning step.
[0204] Specifically, the cleaning robot first follows a preset global cleaning path, prioritizing the cleaning of areas outside the first cleaning zone. During this time, the side brushes, roller brush, and mop are all in normal working order to efficiently complete the basic cleaning of other areas. Simultaneously, the robot's sensors collect images and infrared data from the first cleaning zone in real time. Using behavioral recognition algorithms, it determines whether the pet is still active within the area—for example, whether the pet's head is near the food bowl or its body is within the zone's boundaries. It also records the cleaning completion rate of other areas, such as the percentage of the total area that has been cleaned. When the cleaning progress reaches 90% or the pet is detected leaving the first cleaning zone, the robot triggers the next cleaning preparation command.
[0205] Step 1203: When other cleaning areas have been completed and / or the pet has left the first cleaning area, perform cleaning on the first cleaning area.
[0206] In practice, when other cleaning areas have been cleaned and / or the pet has left the first cleaning area, the cleaning robot will then perform cleaning on the first cleaning area.
[0207] Specifically, if all other areas have been cleaned, meaning the cleaning robot determines the cleaning progress to be 100%, and / or the cleaning robot's sensors detect that the pet has completely left the first cleaning area (meaning no pet outline is captured in the images taken in the first cleaning area for 5 consecutive seconds), the cleaning robot adjusts its cleaning path and moves towards the first cleaning area to perform cleaning. During this process, the cleaning robot automatically switches to the corresponding cleaning mode based on the previously identified feeding behavior of the pet: for the drinking area, the side brushes are raised and inactive, while the roller brush and mop operate according to the wastewater cleaning parameters; for the feeding area, the side brushes are inactive, while the roller brush increases its speed to clean particles. During the cleaning process, sensors continuously monitor the area for any pet returning. If a pet re-enters, cleaning is paused and the robot exits the area, resuming operation only after the pet leaves again, until the first cleaning area is completely cleaned.
[0208] The following examples will describe in detail the cleaning modes corresponding to different pet diets, which will not be repeated here.
[0209] In this embodiment, by constructing a virtual feeding area and designating it as the first cleaning area, the scope requiring special handling during the cleaning process of the cleaning robot is precisely defined. Priority is given to cleaning other areas outside the first cleaning area to avoid interfering with the pet's diet. The cleaning progress and the pet's activity status are continuously monitored, which can not only efficiently promote the overall cleaning work, but also ensure that the cleaning rhythm is reasonably planned without disturbing the pet. At the same time, it also avoids possible disturbance or secondary pollution when the pet is present, which greatly improves the accuracy and efficiency of intelligent cleaning.
[0210] In an exemplary embodiment, as shown in FIG13, the specific process of performing the cleaning step of the first area in step 1203 includes:
[0211] Step 1301: Identify the cleaning area and / or degree of dirt in the first cleaning area, and clean the first cleaning area based on the cleaning area and / or degree of dirt.
[0212] In practice, after determining that the first cleaning area needs cleaning, the cleaning robot identifies the cleaning area and / or degree of dirt in the first cleaning area, and formulates a corresponding cleaning mode and sweeping plan based on the cleaning area and / or degree of dirt to clean the first cleaning area. Specifically, the cleaning robot uses RGB and infrared sensors mounted on its body to perform a comprehensive scan of the first cleaning area, determines the actual area within the area boundary through image segmentation technology, and compares it with a preset area threshold (e.g., 0.5 square meters) to determine the size of the cleaning area; the degree of dirt in the first cleaning area is identified by analyzing the grayscale value and texture features of stains in the image, such as reflective areas of sewage and the distribution density of food particles, converting them into a quantified dirt index, and comparing it with the dirt threshold to determine the degree of dirt in the area. Then, based on the identification results, the cleaning robot automatically matches cleaning parameters and determines a specific cleaning mode and sweeping plan to clean the first cleaning area.
[0213] Optionally, the process of identifying the degree of dirt in the first cleaning area is as follows: by processing a preset dirt index, a quantified dirt score can be obtained. For example, the dirt score is from 1 to 10 points. Thus, a preset dirt threshold of 6 points is set. Based on the relationship between the dirt score obtained from the quantified dirt index and the dirt threshold, the degree of dirt in the first cleaning area can be determined.
[0214] In this embodiment, the cleaning range is determined based on the size of the cleaning area, and the cleaning intensity is adjusted according to the degree of dirt. This ensures the cleaning effect of the first cleaning area while avoiding increased energy consumption and component wear caused by over-cleaning, thus further improving the efficiency and rationality of cleaning.
[0215] In an exemplary embodiment, since the first cleaning area is the pet's eating area, the target cleaning strategy of the cleaning robot is described in detail, taking the identification of the pet drinking water or eating in the first cleaning area as examples, as shown in Figure 14. The specific processing procedure of step 1301 includes:
[0216] Step 1401: When the cleaning area of the first cleaning area is greater than the area threshold or the degree of dirtiness is greater than the dirtiness threshold, the cleaning robot is instructed to return to the base station to clean the cleaning parts of the cleaning robot.
[0217] In practice, because the cleaning robot's cloth is quite dirty after cleaning other areas, directly cleaning the first cleaning area could easily lead to the spread of dirt and secondary contamination. Therefore, to ensure successful cleaning, before the cleaning robot begins cleaning the first cleaning area, the state of the cleaning components can be assessed based on the required level of cleanliness. Specifically, if the pet behavior in the first cleaning area is drinking water, the cleaning robot uses sensor data to determine if the cleaning area exceeds a preset area threshold or the degree of dirt exceeds a dirt threshold. If the cleaning area or the degree of dirt in the first cleaning area exceeds the threshold, it indicates that there may be a lot of dirty water in the first cleaning area. The cleaning robot can then be instructed to return to the base station to clean its cleaning components, for example, by cleaning and drying them, to prevent the dirt area from spreading and to avoid secondary contamination. Then, the first cleaning area can be cleaned. In other words, the cleaning robot's processing unit will send a command to the cleaning robot to return to the base station. Upon receiving the instruction, the cleaning robot pauses its cleaning preparation work in the current area and autonomously returns to the base station along the planned optimal path. The base station then initiates the cleaning process for the cleaning components. Optionally, for the initial cleaning area generated by drinking water, the primary cleaning objective of the cleaning robot is to clean the wastewater. Therefore, the cleaning of the cleaning components can mainly involve drying, and the degree of drying of the cloth should be greater than that of a cloth during regular cleaning.
[0218] For the first cleaning area, which is identified as a pet eating behavior, the cleaning robot uses sensor data to determine if the area exceeds a preset threshold or if the level of dirt exceeds a preset threshold. Before officially starting to clean this area, the robot will proactively return to its base station to thoroughly clean the mop to remove residual dirt and perform a moderate drying process. Considering that cleaning in this first cleaning area typically involves particulate matter, the core purpose of cleaning the cleaning components is to prevent residual dirt on the mop from causing secondary contamination during subsequent cleaning, ensuring effective cleaning. Therefore, the cleaning components are primarily cleaned, for example, to remove tangled fibers and particles embedded in crevices. After completing this preparation, the cleaning robot proceeds to the area to be cleaned to perform the first complete cleaning.
[0219] In an optional embodiment, cleaning the cleaning components of the cleaning robot includes drying and washing. Specifically, the cleaning components include side brushes, roller brushes, and cloths. In the washing process, for components such as cloths and roller brushes that are prone to dirt accumulation, the base station first uses high-pressure water to rinse the surface of the cloths to remove wastewater, food residue, or hair. Simultaneously, the built-in brushes are activated to deeply clean the roller brushes, removing tangled fibers and particles embedded in crevices. For stubborn stains (such as dried oil stains), a neutral detergent is added to aid dissolution, ensuring no residual dirt remains on the surface of the cleaning components. After washing, the cleaning process begins: the base station uses a hot air circulation system to heat and dry the cleaning components. The drying temperature is set according to the material of the cleaning components, and airflow disturbance is used to ensure even drying and prevent localized dampness that could breed bacteria. Furthermore, the drying time is dynamically adjusted according to the degree of dirt on the cleaning components: if a small amount of water remains after washing, the drying time is extended until completely dry. By combining washing and drying, stains on the cleaning items can be thoroughly removed, preventing secondary contamination, while keeping the items dry and clean, providing a reliable guarantee for subsequent cleaning of the first cleaning area, and improving overall cleaning efficiency and hygiene standards.
[0220] Step 1402: Based on the cleaned parts after cleaning treatment, clean the first cleaning area.
[0221] In practice, after the cleaning robot's cleaning components have been cleaned and dried by the base station and restored to a clean and dry state, the robot will replan its path to return to the first cleaning area and initiate a targeted cleaning process. For the cleaning cloth, due to deep cleaning and enhanced drying, residual stains are effectively avoided from being carried to the area to be cleaned. During the cleaning process, the robot adjusts the coverage density of the cleaning path according to the cleaning area and degree of dirt in the first cleaning area, such as large areas of sewage identified earlier. For heavily soiled areas, a reciprocating sweeping mode is used to increase the contact time between the cloth and the ground; for edges and corners, side brushes are used to help gather stains, which are then handled by the roller brush and cloth in conjunction. Simultaneously, the robot monitors the status of the cleaning components in real time. If it detects that the cloth has absorbed a lot of stains again or the roller brush is slightly tangled, it will automatically shorten the round-trip interval with the base station and clean the cleaning components again, ensuring that the cleaning tasks are always performed with clean components, ultimately achieving thorough cleaning of the first cleaning area and avoiding secondary pollution.
[0222] In this embodiment, for situations where the cleaning area of the first cleaning zone is too large or the degree of dirt is high, the cleaning robot is first instructed to return to the base station to clean the cleaning parts, and then the cleaned parts are used for cleaning. This avoids secondary pollution caused by cleaning parts with dirt when cleaning large areas or highly dirty areas, and ensures that the cleaning parts are put into cleaning work in optimal condition. This significantly improves the cleaning effect.
[0223] In an exemplary embodiment, as shown in FIG15, the cleaning components include a side brush, a roller brush, and a cloth. The specific process for cleaning the first cleaning area in step 1402 includes:
[0224] Step 1501: Control the side brush to be in the non-cleaning position, and control the roller brush in the cleaning position to clean the first cleaning area at a first cleaning speed and the cloth in the cleaning position to clean at a second cleaning speed.
[0225] In practice, when the cleaning robot cleans the first cleaning area where the cleaning area is greater than the area threshold or the degree of dirt is greater than the dirt threshold, it can control the side brush to be in a non-cleaning position, that is, the side brush is lifted and not working. At the same time, it controls the roller brush in the cleaning position to clean the first cleaning area at a first cleaning speed and the mop in the cleaning position to clean at a second cleaning speed.
[0226] Specifically, the cleaning strategy of the cleaning robot varies depending on the type of dirt in the first cleaning area. For example, for the first cleaning area identified by pet drinking behavior, the main issue is cleaning dirty water. In this case, if the cleaning area of the first cleaning area is larger than the area threshold, or the degree of dirt is greater than the dirt threshold, the cleaning robot issues a control command to raise or retract the side brush to a non-working position. For example, the side brush is raised to the target position on the chassis to prevent it from stirring up the dirty water and other dirt in the first cleaning area during the cleaning process, thus preventing the spread of stains and secondary pollution. At the same time, it ensures that the roller brush and the mop are in contact with the ground: the roller brush rotates at a preset first cleaning speed, which is the normal cleaning speed of the roller brush. The cleaning robot's mop rotates at a second cleaning speed to clean the dirty water stains in the first cleaning area. The second cleaning speed is the normal cleaning speed of the mop. While ensuring effective wiping of stains on the ground, the normal cleaning speed reduces the splashing of dirty water caused by high-speed friction, and the normal cleaning speed allows the mop to have more contact with the ground, improving the adsorption effect on stubborn stains. At this point, the second cleaning speed can be the same as the first cleaning speed.
[0227] For example, in the first cleaning area identified based on pet eating behavior, the main issue is cleaning solid particles such as cat food crumbs and dog food pellets. If the cleaning area of the first cleaning area exceeds a certain threshold, the cleaning robot issues a control command to raise or retract the side brush to a non-working position. For instance, the side brush is raised to a target position on the chassis to prevent it from agitating food particles and other dirt during cleaning, thus preventing secondary pollution. Simultaneously, the cleaning robot ensures the roller brush and mop are in contact with the ground: the roller brush rotates at a preset first cleaning speed, which is its high-speed cleaning capability. This first cleaning speed allows it to quickly gather and suck up dispersed dirt, such as cat food crumbs and dog food pellets, into the dust collection box. Then, the cleaning robot's mop rotates at a second cleaning speed, lower than the first cleaning speed of the roller brush, to clean the ground in the first cleaning area. In this case, the first cleaning speed is greater than the second cleaning speed.
[0228] In this embodiment, placing the side brush in a non-cleaning position can prevent food particles, sewage, and other dirt from spreading outwards from the first cleaning area when it rotates, thus reducing the risk of secondary pollution. Furthermore, the roller brush and the cloth in the cleaning position are controlled to perform targeted cleaning treatment on the first cleaning area at a first cleaning speed and a second cleaning speed, respectively, making the cleaning of the first cleaning area more thorough and precise.
[0229] In an exemplary embodiment, as shown in FIG16, after step 1402, the method further includes:
[0230] Step 1601: Based on the degree of dirt in the first cleaned area after the initial cleaning, determine whether to clean the cleaning parts and repeat the cleaning of the first cleaned area.
[0231] In practice, after the initial cleaning of the first cleaning area, the cleaning robot performs a second scan using its body sensors to re-evaluate the degree of residual dirt within the area. For example, image analysis can be used to calculate indicators such as the percentage of dirt area and particle density to detect the degree of dirt in the first cleaning area. If the degree of residual dirt is still higher than a preset threshold, the first cleaning area is determined to require further processing. At this point, the cleaning robot will assess the current state of the cleaning components to determine whether to perform further cleaning. If further cleaning is required, the robot will be instructed to return to the base station to perform targeted cleaning of the roller brush and cloth, and then replan its path to return to the first cleaning area to repeat the cleaning process. This process is repeated at least once or multiple times until the cleanliness of the first cleaning area meets the preset cleaning standard, at which point the cleaning of the first cleaning area is complete.
[0232] In this embodiment, by assessing the dirt level of the area after the initial cleaning, if the dirt level is still insufficient, timely cleaning of the cleaning device can prevent secondary pollution caused by residual dirt. The first cleaning area can then be cleaned again. This ensures the cleaning device is in a clean state and, with specifically adjusted cleaning parameters, can more efficiently handle remaining stains. If the dirt level has reached the required standard, no additional operation is needed, reducing unnecessary energy consumption and cleaning device wear. This approach ensures the cleaning quality of the first cleaning area while also balancing efficiency and resource conservation, making the cleaning process more intelligent and efficient.
[0233] In an optional embodiment, when the first cleaning area is repeatedly cleaned, for example, when the first cleaning area is mopped again, the degree of drying of the cloth will gradually increase with the increase of the number of moppings, forming a dynamic adjustment mechanism in which the number of moppings is proportional to the degree of drying. Specifically, during the first mopping, since a small amount of sewage or wet stains may still remain on the ground, the degree of drying of the cloth is lower to enhance the adsorption capacity of liquid dirt and ensure that the residual sewage traces can be effectively wiped away; during the second mopping, the dirt on the ground has been greatly reduced, with only a small amount of water stains remaining. At this time, the base station appropriately increases the degree of drying of the cloth to reduce the moisture carried by the cloth and avoid re-wetting the ground; as the number of moppings continues to increase (such as the third time and above), the ground approaches cleanliness, and only the surface water stains need to be treated. The degree of drying of the cloth will be further increased, even approaching complete dryness, so that the dry cloth can quickly absorb the residual moisture and accelerate the drying of the ground. This design, which dynamically adjusts the drying level according to the mopping stage, ensures efficient removal of stains in the early stages while avoiding dampness on the floor caused by overly wet mops in the later stages. Ultimately, it achieves a balance between cleaning effect and drying efficiency, allowing the first cleaning area to be thoroughly cleaned and quickly restored to a dry state.
[0234] In an exemplary embodiment, as shown in FIG17, the specific process of repeatedly cleaning the first cleaning area in step 1601 includes:
[0235] Step 1701: Control the side brush to be in the non-cleaning position, and control the roller brush in the cleaning position to clean the first cleaning area at the third cleaning speed and the rag in the cleaning position to clean at the first cleaning speed.
[0236] In implementation, after the initial cleaning of the first cleaning area, if the cleaning robot determines that repeated cleaning of the first cleaning area is necessary, to ensure the cleaning effect of the first cleaning area, the cleaning components can be cleaned repeatedly first to ensure that the cleaning components do not cause secondary contamination to the first cleaning area. Then, after the cleaning components have been cleaned, the cleaning robot will adjust parameters according to the type of residual dirt to determine the cleaning mode of the cleaning components, thereby realizing repeated cleaning of the first cleaning area. Specifically, when the remaining dirt in the first cleaning area is mainly solid particles, regardless of whether the cleaning area of the first cleaning area is greater than or less than the area threshold, the cleaning robot controls the side brush to be in a non-cleaning position, that is, raised and not working, to avoid the side brush disturbing the particles and causing secondary contamination. At the same time, the roller brush operates at high speed (third cleaning speed) and the rotation speed of the mop is appropriately increased (first cleaning speed) to perform one or more repeated cleanings of the first cleaning area. At this time, the third cleaning speed is equal to the first cleaning speed.
[0237] When the remaining dirt in the first cleaning area is mainly sewage, if the cleaning area of the first cleaning area is greater than the area threshold, or the degree of dirt is greater than the dirt threshold, the cleaning robot controls the side brush to be in a non-cleaning position, i.e., raised and not working, to avoid the side brush disturbing the sewage area and causing secondary pollution. At the same time, the roller brush reduces its rotation speed (third cleaning speed) and appropriately increases the rotation speed of the mop (first cleaning speed) to clean the first cleaning area once or multiple times until the first cleaning area is cleaned. In this case, the third cleaning speed is lower than the first cleaning speed. If the cleaning area of the first cleaning area is less than the area threshold, or the degree of dirt is less than the dirt threshold, the cleaning robot adopts the same cleaning mode as in the case where it is greater, cleaning the first cleaning area once or multiple times until the first cleaning area is cleaned. In addition, if the cleaning area of the first cleaning area is greater than the area threshold, but the degree of dirt is less than the dirt threshold, the cleaning robot controls both the side brush and the roller brush to be in a non-cleaning position, i.e., raised and not working, to avoid the side brush disturbing the sewage area and causing secondary pollution, and cleans the first cleaning area once or multiple times with a mop at normal rotation speed until the first cleaning area is cleaned.
[0238] In an exemplary embodiment, a first cleaning area is identified based on the pet's eating behavior. This first cleaning area may contain solid particles, such as cat food crumbs, dog food pellets, etc. When solid particles are present in the first cleaning area, the first cleaning speed of the cleaning robot's roller brush can be determined based on the particle density of the solid particles during the cleaning process; or, the first cleaning speed of the roller brush can be inversely proportional to the second cleaning speed of the side brush. Specifically, when solid particles, such as pet food crumbs or granular residue, are present in the first cleaning area, the first cleaning speed of the cleaning robot's roller brush will be dynamically adjusted according to the actual scenario. For example, the cleaning robot identifies the particle density of the solid particles through sensors. When a high density of solid particles is detected in the area, for example, more than 5 particles per square centimeter, the first cleaning speed of the roller brush is automatically increased to a higher level, utilizing stronger centrifugal force and entrainment ability to quickly gather and remove a large amount of dirt. If the number of solid particles is small, for example, less than 2 particles per square centimeter, the rotation speed is appropriately reduced. In this way, energy consumption and component wear are reduced while ensuring cleaning effectiveness. In another optional implementation, the first cleaning speed of the roller brush can be inversely proportional to the second cleaning speed of the mop: when solids need to be cleaned first, the roller brush rotates at a high speed (e.g., 1800 rpm), while the mop speed is reduced to a low speed (e.g., 400 rpm) to prevent the mop from spinning up particles too quickly; when the amount of solids decreases and the focus is on wiping residual stains, the roller brush speed is reduced (e.g., 1000 rpm), and the mop speed is increased accordingly (e.g., 800 rpm). By dynamically distributing the speeds to balance the cleaning and wiping functions, it is ensured that solid particles are efficiently removed while floor stains are also adequately treated. At the same time, by controlling the inverse relationship between the first cleaning speed of the roller brush and the second cleaning speed of the mop, the power is balanced, which can increase the cleaning robot's battery life and thus improve the overall cleaning quality and efficiency.
[0239] In an exemplary embodiment, since the first cleaning area is the pet's eating area, the target cleaning strategy of the cleaning robot is explained by taking the identification of the pet drinking or eating in the first cleaning area as an example. As shown in Figure 18, the specific processing procedure of step 1301 includes:
[0240] Step 1801: When the cleaning area of the first cleaning area is less than the area threshold and / or the degree of dirt is less than the dirt threshold, the first cleaning area is cleaned based on the cleaning component of the cleaning robot.
[0241] In practice, when the cleaning area of the first cleaning zone is less than the area threshold or the degree of dirt is less than the dirt threshold, the cleaning robot does not need to return to the base station because the dirt is light. It can directly use the current cleaning component to clean the first cleaning zone with gentler parameters. Specifically, if the cleaning robot determines that the first cleaning zone is small or lightly dirty, it does not need to initiate an enhanced cleaning process and directly calls the regular cleaning mode to perform the first cleaning treatment on the first cleaning zone.
[0242] Step 1802: Based on the degree of dirt in the first cleaned area after the initial cleaning, determine whether to clean the cleaning parts and repeat the cleaning of the first cleaned area.
[0243] In practice, after the initial cleaning of the first cleaning area, the cleaning robot re-inspects the degree of dirt in the first cleaning area. If the degree of dirt in the first cleaning area still does not meet the cleaning standards, the cleaning robot first determines whether a second cleaning of the current cleaning component is needed based on the degree of dirt in the first cleaning area. If cleaning of the cleaning component is required, the first cleaning area is cleaned again after the cleaning component is cleaned. If the current cleaning component is still relatively dry, the first cleaning area can be cleaned again directly based on the current state of the cleaning component.
[0244] Specifically, after the cleaning robot performs its first cleaning of the first cleaning area, its sensors quickly scan for residual dirt. If the level of dirt is below a preset secondary cleaning threshold, no further cleaning is needed, and the cleaning process is not repeated. If a small amount of dirt remains, such as undried water stains, the robot determines that the first cleaning area needs to be cleaned again. However, because the dirt is light, the robot does not need to return to the base station and can directly use the current cleaning component with gentler parameters for a supplementary cleaning, ensuring the area is clean while avoiding efficiency losses caused by frequent returns to the base station. If it detects that the cleaning component needs to be cleaned at this point, the robot will clean the first cleaning area again after processing the component.
[0245] In this embodiment, for cases where the first cleaning area is small and / or the degree of dirt is low, the cleaning parts of the cleaning robot can be used directly for cleaning, which can quickly and efficiently complete basic cleaning and avoid wasting time and energy by starting a complex process; if the dirt has reached the standard, no additional operation is required, reducing unnecessary resource consumption; if there is still a small amount of residue, it can be treated in a targeted manner, which not only ensures the cleaning effect, but also avoids increasing the wear and tear and energy consumption of the cleaning parts due to over-cleaning.
[0246] In an exemplary embodiment, the cleaning components of the cleaning robot include side brushes, roller brushes, and mops, as shown in Figure 19. Step 1801, based on the cleaning components of the cleaning robot, cleans the first cleaning area, specifically including:
[0247] Step 1901: Control the side brush to be in the non-cleaning position, and control the roller brush and the cloth in the cleaning position to clean the first cleaning area at the second cleaning speed.
[0248] In practice, when the cleaning area of the first cleaning zone is less than the area threshold and / or the degree of dirt is less than the dirt threshold, the cleaning robot controls the side brush to be in a non-cleaning position, that is, lifted and not working, while the roller brush and the cloth in the cleaning position are controlled to clean the first cleaning zone at the second cleaning speed.
[0249] Specifically, the cleaning robot first issues a command to raise or retract the side brush to a non-working position on the side of the machine, preventing it from spreading small amounts of dirt to the already cleaned area when rotating in the first cleaning area, which may be small or lightly soiled. Simultaneously, it ensures that the roller brush and cloth remain in contact with the floor and operate at a uniform second cleaning speed—the normal cleaning speed of the roller brush and cloth. This allows the roller brush to clean small amounts of food particles or debris at this speed, avoiding energy waste from high-speed operation and excessive friction on the floor. The cloth, operating at the same speed, wipes away any remaining water or oil stains with a moderate friction frequency, achieving a cleaning effect without needing to increase the speed.
[0250] In this embodiment, by controlling the side brush to be in a non-cleaning position, the rotation of the side brush can prevent the spread of the small amount of dirt in the first cleaning area, thus preventing secondary pollution. The roller brush and cloth in the cleaning position operate at a second cleaning speed, which can meet the cleaning needs of small areas with light dirt without causing unnecessary energy consumption and component wear due to high-speed operation. This approach balances energy saving and equipment protection, and improves the cleaning efficiency of small, lightly soiled areas.
[0251] In an exemplary embodiment, as shown in FIG20, step 1101, based on the target behavior recognition result, determines the target cleaning strategy, specifically including:
[0252] Step 2001: When the target behavior identification result is sleeping behavior, grooming behavior, or staying behavior, mark the second cleaning area based on the pet's location.
[0253] In practice, when the target behavior of a pet is identified as sleeping, grooming, or lingering, the cleaning robot designates a second cleaning zone based on the pet's current location. Specifically, the cleaning robot uses cameras, infrared sensors, and other devices to capture the pet's behavior in real time. Once it detects that the pet is sleeping, grooming, or lingering in one place for an extended period, it delineates a second cleaning zone centered on the pet's current location, taking into account the pet's size and activity range. For example, for small pets, a circular area with a radius of 60cm is defined with the pet's body center as the origin, while for large pets, the radius is expanded to 100cm. This area is also marked as a "delayed cleaning zone" on the internal map to prevent the cleaning robot from entering while the pet is active.
[0254] Step 2002: Continuously monitor the pet's activity status in the second cleaning area until cleaning of other cleaning areas outside the second cleaning area is completed and / or the pet has left the second cleaning area, and then perform cleaning of the second cleaning area.
[0255] During implementation, the cleaning robot cleans the areas (surrounding areas) other than the second cleaning area. At the same time, it continuously monitors the cleaning progress of other areas and / or the activity status of the pet in the second cleaning area until the cleaning of other cleaning areas is completed and / or the pet has left the second cleaning area, and then performs cleaning of the second cleaning area.
[0256] Specifically, the cleaning robot prioritizes routine cleaning of areas outside the second cleaning zone to avoid disturbing sleeping, grooming, or loitering pets. Simultaneously, sensors track the pet's movements within the second cleaning zone in real time: if the pet is still sleeping, grooming, or loitering, the robot will bypass that area; if it detects the pet has left or that other areas have been completely cleaned, the robot can adjust its path and proceed to the second cleaning zone. During cleaning in the second cleaning zone, parameters are adjusted based on the type of dirt that may be present. For example, the roller brush is improved to handle tangled hair, and the mop operates at low humidity to avoid startling any returning pets, ensuring thorough cleaning of the area without disturbing the pet's rest or activity.
[0257] In this embodiment, by marking a second cleaning area and delaying cleaning for behaviors such as pets sleeping, licking their fur, or lingering, the cleaning robot is prevented from entering when the pet is resting or quietly active, thus reducing disturbance to the pet. The second area is cleaned only after the other areas have been cleaned or the pet has left, which ensures that the overall cleaning work proceeds in an orderly manner, allows the pet to be in a comfortable environment without being disturbed, and ensures that dirt in the second area is cleaned up in a timely manner, thus realizing intelligent and precise cleaning of the second cleaning area.
[0258] In an exemplary embodiment, as shown in FIG21, step 1101, based on the target behavior recognition result, determines the target cleaning strategy, specifically including:
[0259] Step 2101: When the target behavior identification result is abnormal behavior, a warning notification message is sent to the user based on the number of times the abnormal behavior occurs and / or the duration of occurrence.
[0260] In practice, when the target behavior is identified as abnormal, the cleaning robot can send a warning notification to the user based on the frequency and / or duration of the abnormal behavior. Specifically, the cleaning robot can identify abnormal pet behavior through devices such as cameras and sound sensors, such as persistent restless pacing or frequent scratching of the same spot. Furthermore, to avoid false alarms or triggers due to abnormal behavior, the robot can limit the notification to only when the frequency and / or duration of the abnormal behavior reaches a preset threshold. For example, if the abnormal behavior occurs more than three times within one hour, or if a single occurrence lasts more than five minutes, the cleaning robot will automatically trigger a warning mechanism, pushing a notification message to the user via a mobile app that includes the type of abnormal behavior, the time of occurrence, and real-time captured footage, reminding the user to pay attention to the pet's condition and check for potential health risks or environmental discomfort.
[0261] Step 2102: When the target behavior identification result is abnormal behavior, construct a virtual obstacle avoidance area based on the pet's location, execute the obstacle avoidance cleaning mode, and avoid the virtual obstacle avoidance area.
[0262] In practice, when the target behavior is identified as abnormal, the cleaning robot can construct a virtual obstacle avoidance zone based on the pet's location and execute an obstacle avoidance cleaning mode to avoid the virtual obstacle avoidance zone. Specifically, once the cleaning robot detects that the pet is in an abnormal state, it will construct a virtual obstacle avoidance zone centered on the pet's current location and its activity trajectory. For example, if the pet is continuously pacing anxiously in the sofa area, the sofa and a 1-meter radius around it will be designated as an obstacle avoidance zone and marked as a prohibited area on the cleaning map. In this way, when performing cleaning tasks, the cleaning robot will automatically plan a path to avoid this virtual obstacle avoidance zone, avoiding approaching the pet and causing further emotional distress; at the same time, the cleaning robot will reduce its own operating noise to minimize additional stimulation to the pet, ensuring that cleaning work in other areas is completed without disturbing the pet, until the pet's abnormal behavior is resolved or the user manually adjusts the cleaning strategy.
[0263] In this embodiment, when abnormal pet behavior is detected, an alert is sent based on the frequency and duration of the abnormal behavior, allowing users to be aware of the pet's abnormal state in a timely manner and take quick countermeasures. By constructing a virtual obstacle avoidance zone based on the pet's location and performing obstacle avoidance cleaning, the cleaning robot can be prevented from approaching and interfering with the pet, thus preventing the abnormal state from being aggravated. At the same time, it ensures that the cleaning work in other areas can be carried out normally, ensuring the pet's comfort without affecting the overall cleaning effect, and achieving a balance between pet care and cleaning needs.
[0264] In an exemplary embodiment, as shown in FIG22, the specific process of sending a warning notification message to the user based on the number of abnormal behaviors and / or the duration of occurrence in step 2101 includes:
[0265] Step 2201: When the number of occurrences and / or duration of the same abnormal behavior exceed a preset threshold, a warning notification message is sent to the user.
[0266] In practice, to avoid accidental touches by pets or false alarms, the cleaning robot can also be configured with thresholds for the number of times and duration of abnormal behavior. In this way, the cleaning robot will only send a warning notification message to the user when the number of times and / or duration of the same abnormal behavior exceeds the preset threshold.
[0267] Specifically, the cleaning robot continuously records and analyzes abnormal pet behaviors, such as frequent collisions with objects, prolonged periods without eating or drinking, and persistent hiding. It also has preset thresholds for detection; for example, a threshold of 5 occurrences of the same abnormal behavior within 24 hours and a threshold of 30 minutes for each occurrence. Thus, if the pet is detected to have collided with furniture 6 times in a day, or to have refused to eat or drink for more than 40 minutes, both exceeding the preset thresholds, the cleaning robot will integrate specific information about the abnormal behavior, including behavior type, cumulative number of occurrences, total duration, time of occurrence, and associated environmental images, and push a warning notification through the user's linked mobile application. Optionally, the notification will clearly display key data about the abnormal behavior, such as "Your pet has collided with objects 6 times today, exceeding the threshold of 5 times," and suggest that the user check for injuries, stress, or health problems in their pet, thereby helping the user to promptly identify potential risks and provide timely attention and care for their pet.
[0268] In this embodiment, a preset threshold is set to determine the frequency and duration of the same abnormal behavior. When the threshold is exceeded, an alert notification is sent to the user. This effectively avoids misjudging occasional or brief abnormal behaviors of pets and ensures the accuracy of the alert. This mechanism prevents users from being distracted by irrelevant information, and ensures that they receive timely reminders when their pets exhibit persistent or frequent abnormal behaviors. This allows for rapid intervention to understand the pet's condition, investigate potential health problems or environmental discomfort, and provide timely care and assistance to ensure the pet's safety and health.
[0269] In an exemplary embodiment, as shown in FIG23, abnormal behavior corresponds to a reporting priority, and the method further includes:
[0270] Step 2301: When the target behavior identification result contains multiple abnormal behaviors, a warning notification message is sent to the user based on the abnormal behavior with the highest reporting priority.
[0271] During implementation, when the target behavior recognition result contains multiple abnormal behaviors, the cleaning robot will send a warning notification message to the user based on the abnormal behavior with the highest reporting priority.
[0272] Specifically, the cleaning robot pre-sets clear reporting priority levels for different types of abnormal pet behaviors. For example, behaviors that may endanger the pet's life are listed as the highest priority, such as a hen squatting or vomiting. Behaviors that affect the pet's health are listed as medium priority, such as scratching. And behaviors that only indicate emotional abnormalities are listed as low priority. When the cleaning robot detects multiple abnormal behaviors simultaneously, it automatically compares the priority levels of each behavior, selects the highest priority behavior, and generates a warning notification message around that behavior. This warning notification message details the type of the highest priority abnormal behavior, the time of occurrence, the duration, and a screenshot of the scene, while also briefly mentioning the existence of other low-priority abnormal behaviors. This ensures that users can pay attention to the most urgent situation immediately and also allow them to fully understand the pet's overall condition so that they can take targeted measures.
[0273] In this embodiment, when the target behavior identification result contains multiple abnormal behaviors, an early warning notification message is sent based on the abnormal behavior with the highest reporting priority. This avoids users being distracted by a large amount of information and ignoring the core situation, ensuring that users can know the problem that poses the greatest threat to the pet's safety and health as soon as possible and deal with it in a timely manner. At the same time, it also allows users to pay attention to other abnormalities after dealing with high-priority issues, thereby improving response efficiency and better protecting the pet's safety and health.
[0274] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0275] Based on the same inventive concept, this application also provides a cleaning apparatus for implementing the cleaning method described above. The solution provided by this apparatus is similar to the solution described in the above method; therefore, the specific limitations of one or more obstacle cleaning apparatus embodiments provided below can be found in the limitations of the obstacle cleaning method described above, and will not be repeated here.
[0276] In an exemplary embodiment, as shown in FIG24, a cleaning device 2400 is provided, which is applied to a cleaning robot and includes: a first identification module 2401, a second identification module 2402, a third identification module 2403, and a first determination module 2404, wherein:
[0277] The first recognition module 2401 is used to acquire a single frame image and recognize the single frame image based on a preset deep learning algorithm to obtain an initial behavior recognition result.
[0278] The second recognition module 2402 is used to perform fusion judgment on the initial behavior recognition results corresponding to each single frame image based on a multi-frame image fusion and discrimination strategy to obtain the first behavior recognition result.
[0279] The third recognition module 2403 is used to accumulate each single frame image when the recognition result of the first behavior is an abnormality, and to perform time-series behavior recognition on the accumulated multi-frame images according to the time-series behavior recognition algorithm to determine the recognition result of the second behavior.
[0280] The first determining module 2404 is used to determine the target behavior identification result based on the first behavior identification result and the second behavior identification result; the target behavior identification result is used to determine the target cleaning strategy for cleaning treatment.
[0281] In one embodiment, the first identification module 2401 is specifically used for:
[0282] The image acquisition unit captures a single frame of a pet activity scene.
[0283] A single frame image is input into a pre-trained deep learning model, and feature extraction and pet behavior category prediction are performed on the single frame image based on the deep learning model, outputting the probability value of each type of pet behavior;
[0284] Initial behavior recognition results are obtained based on the probability values of various pet behaviors.
[0285] In one embodiment, the second identification module 2402 is specifically used for:
[0286] If the initial behavior recognition results corresponding to multiple consecutive single-frame images are all the same, then the initial behavior recognition result is determined as the first behavior recognition result.
[0287] If the initial behavior recognition results corresponding to multiple consecutive single-frame images are not all the same, then the first behavior recognition result is determined as the default value.
[0288] In one embodiment, the first behavior identification result being identified as abnormal includes: the first behavior identification result being a default value, and / or, the first behavior identification result being that the pet exhibits abnormal behavior.
[0289] In one embodiment, the third identification module 2403 is specifically used for:
[0290] Based on the cumulative single-frame images within the target duration in a preset time sequence, a multi-frame image sequence is obtained;
[0291] Spatial features are extracted from a multi-frame image sequence using a temporal behavior recognition model to obtain feature vectors;
[0292] Based on temporal correlation, the feature vectors are fused to obtain a fused feature vector;
[0293] The fused feature vectors are used to identify pet behavior, and the second behavior identification result is determined.
[0294] In one embodiment, the third identification module 2403 is specifically used for:
[0295] A video segment is obtained by accumulating individual frames within a target duration in a preset time sequence.
[0296] The video segment is input into a fast-slow network, and dual-path frame sampling is performed through the fast-slow network to obtain the first keyframe image sequence corresponding to the slow path and the second keyframe image sequence corresponding to the fast path.
[0297] In one embodiment, the third identification module 2403 is specifically used for:
[0298] By using a fast and slow network, dual-path feature extraction is performed on the first keyframe image sequence and the second keyframe image sequence to obtain the static feature vector corresponding to the slow path and the dynamic feature vector corresponding to the fast path.
[0299] In one embodiment, the third identification module 2403 is specifically used for:
[0300] Adjust the number of channels in the dynamic feature vector corresponding to the fast path, and align the adjusted dynamic feature vector with the static feature vector corresponding to the slow path by channel.
[0301] By using time-averaged pooling layers or convolutional layers in fast and slow networks, the fast path is dimensionality-reduced and sampled in time so that the dynamic feature vector of the dimensionality-reduced fast path is aligned with the static feature vector of the corresponding slow path.
[0302] The dynamic and static feature vectors, which are aligned by channels and time, are fused together to obtain a fused feature vector.
[0303] In one embodiment, the third identification module 2403 is used for:
[0304] Pet behavior recognition is performed by fusing feature vectors through fully connected layers in fast and slow networks, and the probability values of various pet behaviors are output.
[0305] The second behavior identification result is determined based on the probability values of various pet behaviors.
[0306] In one embodiment, the cleaning device 2400 further includes:
[0307] The second determination module is used to determine the target cleaning strategy based on the target behavior recognition results. The target cleaning strategy includes one or more of cleaning, avoidance and early warning.
[0308] In one embodiment, the second determining module is specifically used for:
[0309] When the target behavior is identified as eating behavior, a virtual eating area is constructed based on the pet's location or the location of the food items, and the virtual eating area is marked as the first cleaning area.
[0310] Clean all areas except the first cleaning area, and continuously monitor the cleaning progress and the pet's activity status in the first cleaning area;
[0311] Cleaning of the first cleaning area is performed once other cleaning areas have been completed and / or the pet has left the first cleaning area.
[0312] In one embodiment, the second determining module is specifically used for:
[0313] Identify the cleaning area and / or degree of dirt in the first cleaning area, and clean the first cleaning area based on the cleaning area and / or degree of dirt.
[0314] In one embodiment, the second determining module is specifically used for:
[0315] When the cleaning area of the first cleaning zone is greater than the area threshold or the degree of dirtiness is greater than the dirtiness threshold, the cleaning robot is instructed to return to the base station to clean the cleaning parts of the cleaning robot.
[0316] Based on the cleaned parts after the cleaning process, the first cleaning area is cleaned.
[0317] In one embodiment, the second determining module is specifically used for: drying the cleaning parts of the cleaning robot and cleaning the cleaning parts of the cleaning robot.
[0318] In one embodiment, the second determining module is specifically used for:
[0319] Before cleaning the first cleaning area, the cleaning robot is instructed to return to the base station; the base station is used to dry the cleaning parts of the cleaning robot.
[0320] In one embodiment, the cleaning component includes a side brush, a roller brush, and a cloth; the second determining module is specifically used for:
[0321] The side brush is controlled to be in a non-cleaning position, and the roller brush in the cleaning position is controlled to clean the first cleaning area at a first cleaning speed, and the cloth in the cleaning position is controlled to clean at a second cleaning speed.
[0322] In one embodiment, when solid particles are present in the first cleaning area, the first cleaning speed of the roller brush is determined based on the particle density of the solid particles; or, the first cleaning speed of the roller brush is inversely proportional to the second cleaning speed of the side brush.
[0323] In one embodiment, the cleaning device 2400 further includes:
[0324] The third determining module is used to determine whether to clean the cleaning parts based on the degree of dirt in the first cleaning area after the first cleaning, and to repeat the cleaning of the first cleaning area.
[0325] In one embodiment, the third determining module is specifically used for:
[0326] The side brush is controlled to be in a non-cleaning position, and the roller brush in the cleaning position is controlled to clean the first cleaning area at a third cleaning speed and the cloth in the cleaning position is controlled to clean at a first cleaning speed.
[0327] In one embodiment, the second determining module is specifically used for:
[0328] When the cleaning area of the first cleaning zone is less than the area threshold and / or the degree of dirt is less than the dirt threshold, the first cleaning zone is cleaned based on the cleaning components of the cleaning robot.
[0329] Based on the degree of dirtiness in the first cleaned area after the initial cleaning, determine whether to clean the items and repeat the cleaning process on the first cleaned area.
[0330] In one embodiment, the cleaning component includes a side brush, a roller brush, and a cloth; the second determining module is specifically used for:
[0331] The side brush is controlled to be in a non-cleaning position, and the roller brush and the cloth in the cleaning position are both controlled to clean the first cleaning area at a second cleaning speed.
[0332] In one embodiment, the second determining module is specifically used for:
[0333] When the target behavior is identified as sleeping, grooming, or loitering, a second cleaning area is marked based on the pet's location.
[0334] Continuously monitor the pet's cleaning progress and / or activity status in the second cleaning area until cleaning of other cleaning areas outside the second cleaning area is completed and / or the pet has left the second cleaning area, then perform cleaning of the second cleaning area.
[0335] In one embodiment, the second determining module is specifically used for:
[0336] When the target behavior is identified as abnormal, a warning notification message is sent to the user based on the number of times the abnormal behavior occurs and / or the duration of the occurrence; and / or,
[0337] When the target behavior is identified as abnormal, a virtual obstacle avoidance zone is constructed based on the pet's location, and an obstacle avoidance cleaning mode is executed to avoid the virtual obstacle avoidance zone.
[0338] In one embodiment, the second determining module is specifically used for:
[0339] When the number of occurrences and / or duration of the same abnormal behavior exceed a preset threshold, a warning notification message will be sent to the user.
[0340] In one embodiment, abnormal behavior corresponds to a reporting priority, and the cleaning device 2400 further includes:
[0341] The sending module is used to send a warning notification message to the user based on the abnormal behavior with the highest reporting priority when the target behavior identification result contains multiple abnormal behaviors.
[0342] Obviously, the embodiments described above are merely some, not all, embodiments of the present invention. Based on the embodiments of the present invention, those skilled in the art can make other variations or modifications without creative effort, and all such variations or modifications should fall within the scope of protection of the present invention.
[0343] Each module in the aforementioned cleaning device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0344] In an exemplary embodiment, a cleaning robot 1000 is provided, as shown in Figure 25. The cleaning robot can be a sweeping robot, mopping robot, sweeping and mopping robot, window cleaning robot, etc. The cleaning robot can include a body, a walking system, a sensing system, cleaning components, a processing unit, etc. As shown in Figure 26, the cleaning components of the cleaning robot 1000 include a side brush 210, a roller brush 220, and a mop 230. For example, the mop 230 can be a disc mop, a vibrating disc mop, a tracked mop, a rotary mop, or a roller mop. Taking a rotary mop as an example, its projection along the height direction is rectangular. Taking a tracked mop as an example, along the height direction, the longitudinal section of the tracked mop is an elongated hole shape, while the longitudinal section of the rotary mop is a circular hole shape. The interior of the rotary mop is hollow, and an internal support bracket is disposed in the cavity of the rotary mop to tension the rotary mop. One side of the inner support bracket along its length is rotatably connected to a drive shaft, and the other side of the inner support bracket along its length is rotatably connected to a driven shaft. The drive shaft and the driven shaft are parallel to each other, and both the drive shaft and the driven shaft are parallel to the width direction of the machine body. The rotary wet cleaning component is tensioned outside the drive shaft and the driven shaft.
[0345] In one exemplary embodiment, a cleaning robot is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0346] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0347] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0348] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
[0349] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc. The storage medium can also include combinations of the above types of memory.
[0350] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application may include at least one of relational databases and non-relational databases. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0351] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0352] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A cleaning method, characterized in that, The method is applied to a cleaning robot and includes: acquiring a single-frame image of a pet activity scene through an image acquisition unit; inputting the single-frame image into a pre-trained deep learning model, and performing feature extraction and pet behavior category prediction on the single-frame image based on the deep learning model, outputting probability values for various pet behaviors; obtaining an initial behavior recognition result based on the probability values of various pet behaviors; performing a fusion judgment on the initial behavior recognition results corresponding to each single-frame image based on a multi-frame image fusion discrimination strategy to obtain a first behavior recognition result; if the first behavior recognition result is an anomaly, accumulating each single-frame image, and performing time-series behavior recognition on the accumulated multi-frame images according to a time-series behavior recognition algorithm to determine a second behavior recognition result; determining a target behavior recognition result based on the first behavior recognition result and the second behavior recognition result; the target behavior recognition result is used to determine a target cleaning strategy for cleaning processing.
2. The method according to claim 1, characterized in that, The multi-frame image fusion discrimination strategy, which fuses and judges the initial behavior recognition results corresponding to each single frame image to obtain a first behavior recognition result, includes: if the initial behavior recognition results corresponding to multiple consecutive single frames are all the same, then the initial behavior recognition result is determined as the first behavior recognition result; if the initial behavior recognition results corresponding to multiple consecutive single frames are not all the same, then the first behavior recognition result is determined as a default value.
3. The method according to claim 1, characterized in that, The first behavior recognition result being an abnormality includes: the first behavior recognition result being a default value, and / or the first behavior recognition result being that the pet exhibits abnormal behavior.
4. The method according to claim 1, characterized in that, The method further includes: determining a target cleaning strategy based on the target behavior recognition result, wherein the target cleaning strategy includes one or more of cleaning, avoidance, and early warning.
5. The method according to claim 4, characterized in that, The step of determining the target cleaning strategy based on the target behavior recognition result includes: when the target behavior recognition result is eating behavior, constructing a virtual eating area based on the pet's location or the location of the food items, and marking the virtual eating area as the first cleaning area; cleaning other cleaning areas besides the first cleaning area, and continuously monitoring the cleaning progress and the pet's activity status in the first cleaning area; when the cleaning of other cleaning areas is completed and / or the pet has left the first cleaning area, performing cleaning on the first cleaning area.
6. The method according to claim 4, characterized in that, The step of determining the target cleaning strategy based on the target behavior recognition result includes: when the target behavior recognition result is sleeping behavior, grooming behavior, or staying behavior, marking a second cleaning area based on the pet's location; continuously monitoring the pet's cleaning progress and / or activity status in the second cleaning area until cleaning of other cleaning areas besides the second cleaning area is completed and / or the pet has left the second cleaning area, and then performing cleaning on the second cleaning area.
7. The method according to claim 4, characterized in that, The step of determining the target cleaning strategy based on the target behavior recognition result includes: when the target behavior recognition result is abnormal behavior, sending a warning notification message to the user based on the number of occurrences and / or duration of the abnormal behavior; and / or, when the target behavior recognition result is abnormal behavior, constructing a virtual obstacle avoidance zone based on the pet's location, executing an obstacle avoidance cleaning mode, and avoiding the virtual obstacle avoidance zone.
8. The method according to claim 7, characterized in that, The step of sending a warning notification message to the user based on the number of occurrences and duration of the abnormal behavior includes: sending a warning notification message to the user when the number of occurrences and / or duration of the same abnormal behavior exceeds a preset threshold.
9. The method according to claim 7, characterized in that, The abnormal behavior corresponds to a reporting priority, and the method further includes: when the target behavior identification result contains multiple abnormal behaviors, sending a warning notification message to the user based on the abnormal behavior with the highest reporting priority.
10. A cleaning robot, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.