Sweeping robot control method and system
By fusing image, spatial, and sound data into a three-stream neural network, the robotic vacuum cleaner can identify obstacles and execute adaptive cleaning strategies, solving the problem of poor cleaning performance in existing technologies and achieving intelligent cleaning results and improved efficiency.
Patent Information
- Application Number
- CN202511050782.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-11
AI Technical Summary
Existing robotic vacuum cleaners are unable to take accurate measures when encountering obstacles in the cleaning area, resulting in poor cleaning performance and low level of intelligence.
A three-stream neural network with a fusion attention mechanism processes image, 3D space, and sound data. It performs adaptive fusion through a cross-modal attention mechanism, identifies obstacles, and performs hierarchical scene adaptive planning. It adopts an adaptive variational cleaning strategy, adjusts cleaning parameters and paths in real time, and combines visual verification to identify residual stains and trigger a re-cleaning procedure.
It enables the robot vacuum to intelligently respond to obstacles, improving cleaning effectiveness and efficiency, ensuring cleaning quality, and reducing repetitive cleaning and energy consumption.
Smart Images

Figure CN120918529A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, specifically to a control method and system for a sweeping robot. Background Technology
[0002] A robotic vacuum cleaner is an intelligent device that can move autonomously and perform cleaning. Current robotic vacuum cleaners generally perform cleaning tasks based on a single cleaning strategy. When they encounter obstacles in the cleaning area, they cannot take accurate measures or guarantee the cleaning effect, and their level of intelligence is relatively low. Summary of the Invention
[0003] Based on the above-mentioned problems, this invention proposes a control method and system for a sweeping robot. Through the solution of this invention, when the sweeping robot encounters obstacles in the cleaning area, it can intelligently take accurate countermeasures and achieve excellent cleaning results.
[0004] In view of this, one aspect of the present invention proposes a control method for a robotic vacuum cleaner, comprising:
[0005] Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner;
[0006] The first image data, the first three-dimensional spatial data, and the first sound data are processed using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment.
[0007] The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result;
[0008] Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task;
[0009] After the cleaning task is completed, a second image of the cleaned area is collected;
[0010] Based on the second image data and the preset cleaning effect recognition model, identify residual stains on the ground;
[0011] If the area of the residual stain exceeds a preset first area threshold, a re-cleaning procedure is triggered, and the corresponding cleaning mode is automatically switched according to the type of residual stain.
[0012] Optionally, the step of processing the first image data, the first three-dimensional spatial data, and the first sound data using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment includes:
[0013] The first image data is input into a first-stream convolutional neural network, the first three-dimensional spatial data is input into a second-stream point cloud processing network, and the first sound data is input into a third-stream temporal neural network, and the initial feature representations of each modality are extracted in parallel.
[0014] The initial feature representations output by each stream network are processed to unify the feature dimensions, mapping the feature vectors of different modalities to the same feature space dimension, ensuring the compatibility of subsequent fusion processing;
[0015] The importance weights of features within each modality are calculated separately using a self-attention mechanism, thereby enhancing the expression strength of key features within each modality and suppressing interference from redundant information.
[0016] A cross-modal attention mechanism is used to calculate the correlation matrix between different modalities, and the fusion weights of each modal feature are adaptively allocated according to the current environmental scene to achieve the integration of complementary information between modalities.
[0017] The weighted multimodal features are cascaded and fused, and a comprehensive environmental feature representation containing object attribute features, spatial features, and sound attribute features is output through a fully connected layer, which serves as the input data for subsequent object recognition and classification.
[0018] Optionally, the step of adaptively fusing object attribute features, spatial features, and sound attribute features through a cross-modal attention mechanism to obtain the object recognition result includes:
[0019] A cross-modal correlation mapping table is constructed, and the correlation strength matrix between different modal features is established by calculating the semantic similarity between object attribute features, spatial features and sound attribute features;
[0020] Based on the correlation strength matrix, attention weight coefficients for each modality are dynamically generated, and the fusion ratio of object attribute features, spatial features, and sound attribute features is adaptively adjusted according to the reliability and importance of information of each modality in the current environment.
[0021] The weighted multimodal features are fused at the feature level, and a unified feature representation containing multimodal information is generated through feature concatenation and dimensionality reduction.
[0022] The unified feature representation is input into a pre-trained object recognition classifier, which performs feature parsing and pattern matching through a deep neural network, and outputs the object category probability distribution and confidence score.
[0023] By combining the confidence score and the preset recognition threshold, the final object recognition result is determined, and structured recognition information containing object type, location coordinates, motion state and risk level is generated.
[0024] Optionally, the step of performing hierarchical scene adaptive planning and adopting an adaptive variational cleaning strategy to perform the cleaning task based on the object recognition result includes:
[0025] Based on the object type, location coordinates, and risk level information in the object recognition results, the cleaning environment is divided into high-risk areas, medium-risk areas, and low-risk areas.
[0026] Differentiated route planning strategies are generated for areas with different risk levels. For high-risk areas, detour routes are generated and alarm information is sent to user terminals. For medium-risk areas, deceleration routes are planned. For low-risk areas, efficient coverage routes are planned.
[0027] Based on the identified floor material type and stain type, an adaptive cleaning parameter adjustment mechanism is built to dynamically adjust cleaning parameters including suction power, brush head speed and brush head height;
[0028] During the cleaning task, the cleaning effect and environmental changes are continuously monitored based on real-time sensor feedback information. When new obstacles are detected or the cleaning parameters deviate from expectations, the cleaning strategy and path planning are adjusted in real time.
[0029] Establish a cleaning task execution status evaluation mechanism to record the cleaning completion rate, time consumption, and energy consumption data of each area. Based on the evaluation results, optimize the parameter configuration and path planning of subsequent cleaning tasks to form an adaptive learning and continuous improvement closed-loop control system.
[0030] Optionally, the step of identifying residual stains on the ground based on the second image data and a preset cleaning effect recognition model includes:
[0031] The second image data is preprocessed, including image denoising, brightness equalization, and HSV color space conversion, to enhance the contrast between residual stains and the cleaned floor.
[0032] The preprocessed second image data is input into a cleaning effect recognition model based on YOLOv8, which integrates RGB features and HSV color space features and uses a multi-scale feature extraction network to identify different types of residual stains.
[0033] The identified residual stains were classified and labeled into four categories: dust residue, hair residue, liquid stain residue, and particulate residue. The pixel coordinates and contour information of each residual stain were extracted.
[0034] Based on the mapping relationship between pixel coordinates and the actual ground, the actual area of each residual stain is calculated, and the calculation result is compared with a preset first area threshold.
[0035] Based on the type and area of the residual stains, a corresponding re-cleaning instruction is generated, including the coordinates of the target cleaning area, the recommended cleaning mode, and the estimated cleaning time.
[0036] Optionally, the step of triggering a re-cleaning procedure if the area of the residual stain exceeds a preset first area threshold, and automatically switching the corresponding cleaning mode according to the type of residual stain, includes:
[0037] The calculated residual stain area is compared with a preset first area threshold. When the residual stain area is greater than the first area threshold, the system automatically triggers the re-cleaning program and generates a re-cleaning task queue.
[0038] Based on the identified residual stains, the corresponding cleaning parameter configuration is matched from the preset cleaning mode database. The cleaning parameter configuration includes suction level, brush head rotation speed, travel speed and cleaning time.
[0039] Different cleaning modes are switched according to different types of residual stains;
[0040] Based on the spatial coordinate information of residual stains, a precise re-cleaning path is planned, enabling the robot vacuum cleaner to directly navigate to the stained area and perform localized deep cleaning according to the optimized cleaning trajectory.
[0041] The cleaning effect is monitored in real time during the re-cleaning process. Real-time feedback data of the cleaned area is obtained through onboard sensors. The re-cleaning program is automatically ended when the stains are detected to be removed or the preset cleaning time is reached.
[0042] Optionally, the three-stream neural network fusion process employs a multi-scale adaptive weight allocation mechanism, and the formula for calculating the fused feature vector is as follows:
[0043]
[0044] in:
[0045] F fusion (x, y, z) represents the final multimodal fusion feature vector, where x represents the image modal input data, y represents the three-dimensional spatial modal input data, and z represents the audio modal input data;
[0046] This represents the summation over R different scale / resolution levels, where R represents the total number of layers in the multi-scale fusion, and r represents the current scale index.
[0047] μ r This represents the global weight coefficient at the r-th scale, used to balance the importance of features at different scales, satisfying...
[0048] F image(x, r) represents the feature vector of the image modality at the r-th scale. F represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (z,r) represents the feature vector of the audio modality at the r-th scale;
[0049] α image (r) represents the feature vector of the image modality at scale r. α represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (r) represents the attention weight of the audio modality at the r-th scale;
[0050] β is the regularization intensity coefficient, γ is the sensitivity adjustment parameter, tanh() is the hyperbolic tangent activation function used for numerical stability; ||·|| represents the vector norm (usually the L2 norm); F m Let m be the eigenvector of the m-th mode. Let be the mean of the feature vector of the m-th modality; i, s, and a represent the image, spatial, and sound modalities, respectively.
[0051] δ represents the confidence level moderating coefficient, σ() is the Sigmoid activation function, and A conf (x, y, z) is the multimodal confidence function. is the second-order gradient operator (Hessian matrix), and ||·‖‖2 is the L2 norm normalization.
[0052] Optionally, the steps for performing a cleaning task using an adaptive variational cleaning strategy include:
[0053] An intelligent parameter adjustment algorithm based on multi-dimensional pollution characteristics is adopted, and its cleaning intensity adjustment formula is as follows:
[0054]
[0055] in:
[0056] I clean (u, v, t) represents the sweeping intensity at position (u, v) at time t; I base Based on the basic cleaning intensity constant;
[0057] D(u, v, t) represents the pollution density distribution function at location (u, v) at time t;
[0058] and Let represent the second-order partial derivatives of the pollution density in the u and v directions, respectively;
[0059] ξ and η are gradient sensitivity coefficients used to control the degree of response of cleaning intensity to changes in contamination density gradient;
[0060] C effect (u, v, th) represents the cleaning effect evaluation value of position (u, v) at the previous time t-1;
[0061] λ is the historical effect decay coefficient, used to avoid over-cleaning areas that have already been thoroughly cleaned;
[0062] κ is the dirt type adjustment coefficient, used to control the weight of the influence of different dirt types on cleaning intensity;
[0063] H dirt (u, v, t) represents the dirt hardness characteristic function at position (u, v) at time t, reflecting the adhesion and difficulty of cleaning the dirt;
[0064] V robot (t) represents the robot's actual speed at time t;
[0065] V optimal (u, v) represents the optimal cleaning speed at position (u, v);
[0066] v is the velocity normalization constant, which prevents division by zero and adjusts the sensitivity of the velocity term;
[0067] Ψ(·) is the velocity adaptation function, defined as follows: Where ρ is the velocity sensitivity parameter;
[0068] exp(·) is an exponential function used to achieve non-linear adjustment of the cleaning effect;
[0069] Optionally, it also includes the step of establishing an intelligent learning system based on user habits; wherein the intelligent learning system adopts a dynamic preference evaluation algorithm based on multi-dimensional feature fusion, and its comprehensive preference evaluation formula is:
[0070]
[0071] in:
[0072] P total (s,a) represents the overall user preference score for performing action a in state s;
[0073] M represents the total number of users;
[0074] ω m Let the weight coefficient of the m-th user be denoted as , satisfying
[0075] N m This represents the number of historical interaction records for the m-th user;
[0076] P m(s, a, n) represents the preference rating of the m-th user for action a in state s during the n-th interaction;
[0077] φ n This represents the confidence coefficient of the nth interaction record;
[0078] t current Indicates the current time;
[0079] t n This indicates the time when the nth interaction occurs;
[0080] This represents the time decay variance of the m-th user's preference;
[0081] Θ(s, a, m) represents the personalized fitness function of user m for state-action pair (s, a), reflecting the consistency of user behavior;
[0082] ζ is the periodic adjustment coefficient, used to control the strength of the influence of time periodicity on preferences;
[0083] T cycle A periodic time window (such as 24 hours, 7 days, etc.) representing user behavior;
[0084] cos(·) is a cosine function used to model the periodic variation characteristics of user preferences;
[0085] mod represents the modulo operation;
[0086] Γ context (s, a, t) current ) represents the context adaptation function, defined as:
[0087]
[0088] Where L is the number of environmental feature dimensions, E l (t current () represents the value of the l-th environmental feature at the current moment. τ represents the historical mean of the l-th environmental feature. l and These are the weighting coefficient and standardization parameter for the l-th environmental feature, respectively;
[0089] exp(·) is an exponential function, and tanh(·) is a hyperbolic tangent function.
[0090] Another aspect of the present invention provides a control system for a sweeping robot, for executing a control method for a sweeping robot, comprising: a sweeping robot and a server;
[0091] The server is configured as follows:
[0092] Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner;
[0093] The first image data, the first three-dimensional spatial data, and the first sound data are processed using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment.
[0094] The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result;
[0095] Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task;
[0096] After the cleaning task is completed, a second image of the cleaned area is collected;
[0097] Based on the second image data and the preset cleaning effect recognition model, identify residual stains on the ground;
[0098] If the area of the residual stain exceeds a preset first area threshold, a re-cleaning procedure is triggered, and the corresponding cleaning mode is automatically switched according to the type of residual stain.
[0099] The control method for a robotic vacuum cleaner using the technical solution of this invention includes: acquiring first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robotic vacuum cleaner; processing the first image data, first three-dimensional spatial data, and first sound data using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment; adaptively fusing the object attribute features, spatial features, and sound attribute features through a cross-modal attention mechanism to obtain object recognition results; performing hierarchical scene adaptive planning and employing an adaptive variational cleaning strategy to perform the cleaning task based on the object recognition results; acquiring second image data of the cleaned area after the cleaning task is completed; identifying residual stains on the ground according to the second image data and a preset cleaning effect recognition model; and triggering a re-cleaning program if the area of the residual stains exceeds a preset first area threshold, and automatically switching the corresponding cleaning mode according to the type of residual stains. Through this invention, the robotic vacuum cleaner can intelligently take accurate countermeasures when encountering obstacles in the cleaning area and achieve excellent cleaning results. Attached Figure Description
[0100] Figure 1 This is a flowchart of a control method for a sweeping robot provided in one embodiment of the present invention;
[0101] Figure 2This is a schematic block diagram of a control system for a sweeping robot provided in one embodiment of the present invention. Detailed Implementation
[0102] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0103] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0104] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0105] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0106] The following reference Figures 1 to 2 This invention describes a control method and system for a sweeping robot according to some embodiments of the present invention.
[0107] like Figure 1 As shown, one embodiment of the present invention provides a control method for a sweeping robot, comprising:
[0108] Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner;
[0109] The first image data includes RGB image data acquired by an airborne camera, the first three-dimensional spatial data is acquired by an airborne depth sensor, and the first sound data is acquired by an airborne sound sensor.
[0110] A three-stream neural network with a fusion attention mechanism is used to process the first image data, the first three-dimensional spatial data, and the first sound data to extract object attribute features, spatial features, and sound attribute features in the environment. This includes: a first-stream network processing the image data to extract object attribute features such as texture, color, and shape; a second-stream network processing the three-dimensional spatial data to extract spatial features such as the three-dimensional contour and spatial location features of the object; and a third-stream network processing the sound data to extract sound attribute features such as sound source, sound content, and sound volume in the surrounding environment.
[0111] The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result;
[0112] In this step, the adaptively fused data is input into a pre-defined object recognition model. Based on a spatiotemporal sequence feature extraction algorithm, the model identifies the object type, ground material type, stain type, and spatial location of each object, distinguishing between static obstacles (such as furniture), dynamic movable objects (such as slippers and pet toys), and dynamic living targets (such as pets and children). The model constructs an object recognition result that includes obstacle location coordinates, object type, ground material type, stain type, motion trajectory prediction, and collision risk level. This step, through multimodal data fusion and dynamic target classification, enables accurate object classification and localization, improving the recognition accuracy of easily confused objects such as transparent, reflective, and low-contrast objects, thus providing a reliable environmental foundation for subsequent control strategies.
[0113] Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task;
[0114] The process includes implementing tiered scenario-adaptive planning, which involves: classifying the environment into high-risk, medium-risk, and low-risk areas based on the types of objects identified; generating detour paths and sending alarm information to the user terminal for high-risk areas (such as areas with spilled liquids, pet excrement, dense cables, or dynamic obstacles); adjusting the robot's cleaning parameters and travel speed for medium-risk areas (such as areas with carpets, thresholds, or scattered small objects); and optimizing the cleaning path for low-risk areas to reduce energy consumption and time consumption. The robot employs an adaptive variational cleaning strategy to perform cleaning tasks, including: automatically adjusting suction power, brush head speed, and brush head height based on the identified floor material type (e.g., hard floor, carpet, tile, etc.). This involves synchronously adjusting the output parameters of the drive motor, side brush motor, roller brush motor, and vacuum motor using a distributed control algorithm based on the dynamic task sequence; increasing the drive motor torque and adjusting the roller brush speed to 1800 rpm (1200 rpm for hard floors) based on feedback signals from pressure sensors when encountering carpet edges; automatically adjusting the cleaning mode and travel speed based on the identified stain type (e.g., dust, hair, liquid, etc.); dynamically adjusting the side brush contact force in 0.5N increments based on real-time distance data from ultrasonic sensors when pushing obstacles to ensure the object does not tip over or slip excessively during the pushing process; and improving the robot's success rate in overcoming obstacles across different floor materials (e.g., from tile to carpet) and increasing the accuracy of the side brush in pushing lightweight objects through multi-motor collaborative control and force-position closed-loop feedback.
[0115] After the cleaning task is completed, a second image of the cleaned area is collected;
[0116] Based on the second image data and the preset cleaning effect recognition model (which can be based on the YOLOv8 improved algorithm and add HSV color space feature input), the residual stains on the ground (such as residual dust, hair, liquid stains, etc.) are identified.
[0117] If the area of the residual stain exceeds a preset first area threshold (e.g., 20cm²), 2 If the stain is not properly cleaned, the cleaning process will be triggered, and the corresponding cleaning mode will be automatically switched according to the type of residual stain (e.g., the water absorption module will be activated for liquid stains, and the vacuuming power will be increased for particulate stains).
[0118] In this step, by introducing a post-cleaning visual verification mechanism, the extensive mode that relies on fixed cleaning time in the traditional solution is transformed into a results-oriented closed-loop control, which improves the cleaning compliance rate of key areas, shortens the average completion time of the overall cleaning task, and avoids repeated and ineffective cleaning.
[0119] Using the technical solution of this embodiment, when the robot vacuum cleaner encounters obstacles in the cleaning area, it can intelligently take accurate countermeasures and achieve excellent cleaning results.
[0120] In some possible embodiments of the present invention, it further includes:
[0121] The data processing and decision-making process is executed by leveraging edge computing and cloud computing collaborative processing mechanisms, including: timely recognition and obstacle avoidance decisions are completed in the robot's local edge computing unit; complex scene data and cleaning history records are uploaded to the cloud server for deep learning analysis; the cloud server periodically distributes optimized recognition models and decision-making strategies based on the historical data analysis results to achieve continuous improvement of robot performance;
[0122] Establish an intelligent learning system based on user habits, including: recording user manual interventions and adjustments, including user focus on cleaning areas and corrections to robot behavior; using reinforcement learning algorithms to continuously optimize cleaning strategies based on user feedback, automatically adjusting the cleaning frequency and intensity of different areas; conducting risk assessments for areas where users frequently move around, planning avoidance time in advance, and reducing interference with users' daily activities.
[0123] Perform autonomous maintenance management, including: monitoring the operating status and wear level of various robot components through sensors; automatically executing simple self-maintenance procedures when abnormal operating parameters or wear levels approaching critical values are detected; and generating detailed fault diagnosis reports and pushing maintenance suggestions to users for problems that cannot be resolved on their own.
[0124] In some possible embodiments of the present invention, the step of processing the first image data, the first three-dimensional spatial data, and the first sound data using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment includes:
[0125] The first image data is input into a first-stream convolutional neural network, the first three-dimensional spatial data is input into a second-stream point cloud processing network, and the first sound data is input into a third-stream temporal neural network, and the initial feature representations of each modality are extracted in parallel.
[0126] Among them, the first-level network (i.e., the first-level convolutional neural network) extracts object attribute features such as texture, color and shape of the image; the second-level network (i.e., the second-level point cloud processing network) extracts spatial features such as three-dimensional contour and spatial location; and the third-level network (i.e., the third-level temporal neural network) extracts sound attribute features such as sound source, spectrum and temporal sequence.
[0127] The initial feature representations output by each stream network are processed to unify the feature dimensions, mapping the feature vectors of different modalities to the same feature space dimension, ensuring the compatibility of subsequent fusion processing;
[0128] The importance weights of features within each modality are calculated separately using a self-attention mechanism, thereby enhancing the expression strength of key features within each modality and suppressing interference from redundant information.
[0129] A cross-modal attention mechanism is used to calculate the correlation matrix between different modalities, and the fusion weights of each modal feature are adaptively allocated according to the current environmental scene to achieve the integration of complementary information between modalities.
[0130] The weighted multimodal features are cascaded and fused, and a comprehensive environmental feature representation containing object attribute features, spatial features, and sound attribute features is output through a fully connected layer, which serves as the input data for subsequent object recognition and classification.
[0131] This embodiment compensates for the blind spots of a single sensor by fusing data from three sensors; a self-attention mechanism highlights key features and improves target recognition accuracy; cross-modal weights are adaptively adjusted to adapt to different lighting and noise conditions; combining 3D spatial data overcomes the difficulty of recognizing transparent objects from images; sound features assist in recognizing objects that are difficult to visually identify on reflective surfaces; multimodal fusion improves target detection capabilities in low-light environments; other modalities can compensate for the failure of a single sensor, improving system reliability; a cross-modal verification mechanism filters false detections and reduces the false recognition rate; attention weights are dynamically adjusted to adapt to environmental changes and interference; parallel computing using a three-stream network fully utilizes hardware resources and improves processing speed; a unified feature space reduces redundant calculations and lowers computational complexity; the attention mechanism reduces the processing of irrelevant information and improves overall efficiency; it maintains high recognition accuracy in indoor environments with dense furniture and complex lighting, while simultaneously recognizing static obstacles, dynamic objects, and living targets, meeting the response time requirements for real-time navigation and obstacle avoidance of the robotic vacuum cleaner.
[0132] In some possible embodiments of the present invention, the step of adaptively fusing object attribute features, spatial features, and sound attribute features through a cross-modal attention mechanism to obtain object recognition results includes:
[0133] A cross-modal correlation mapping table is constructed, and the correlation strength matrix between different modal features is established by calculating the semantic similarity between object attribute features, spatial features and sound attribute features;
[0134] Specifically, for the feature representations of the same object in different modalities, similarity scores between their feature vectors are calculated to form semantic correspondences between modalities;
[0135] Based on the correlation strength matrix, attention weight coefficients for each modality are dynamically generated, and the fusion ratio of object attribute features, spatial features, and sound attribute features is adaptively adjusted according to the reliability and importance of information of each modality in the current environment.
[0136] Among them, the weight of object attribute features is increased in well-lit environments, the weight of spatial features is increased in complex spatial environments, and the weight of sound attribute features is decreased in noisy environments.
[0137] The weighted multimodal features are fused at the feature level, and a unified feature representation containing multimodal information is generated through feature concatenation and dimensionality reduction.
[0138] Among these, maintaining the independence of each modality's features while achieving information complementarity avoids information conflict and feature redundancy between modalities;
[0139] The unified feature representation is input into a pre-trained object recognition classifier, which performs feature parsing and pattern matching through a deep neural network, and outputs the object category probability distribution and confidence score.
[0140] The classifier is trained based on a large amount of multimodal labeled data and can identify different types of objects such as static obstacles, dynamic movable objects, and dynamic living targets.
[0141] By combining the confidence score and the preset recognition threshold, the final object recognition result is determined, and structured recognition information containing object type, location coordinates, motion state and risk level is generated.
[0142] When the confidence level is lower than a preset threshold, a multimodal information re-fusion and secondary recognition process is triggered to ensure the accuracy and reliability of the recognition results.
[0143] This embodiment reduces single-modal misidentification through cross-modal information cross-validation; ensures logical consistency of recognition results across different modalities, providing reliable quantifiable metrics for recognition results; automatically adjusts the contribution of each modality according to environmental conditions, maintaining recognition performance under different lighting, noise, and occlusion conditions, and effectively compensating for the lack of information in a single modality with other modalities; distinguishes complex categories such as static obstacles, dynamic objects, and living targets, while simultaneously identifying multi-dimensional attributes such as object type, position, and motion state, assigning a corresponding risk level to each identified object; avoids simple feature superposition, reducing redundant calculations; reduces invalid recognition attempts through confidence thresholds; only performs secondary processing on low-confidence results, improving overall efficiency; and employs cross-modal consistency checks and confidence dual verification to identify and process inter-modal conflict information, ensuring the system can still function normally even when a single sensor fails. This cross-modal attention fusion mechanism provides the robotic vacuum cleaner with intelligent environmental understanding capabilities, representing a key technological breakthrough for achieving advanced autonomous cleaning.
[0144] In some possible embodiments of the present invention, the step of performing hierarchical scene adaptive planning and employing an adaptive variational cleaning strategy to perform the cleaning task based on the object recognition result includes:
[0145] Based on the object type, location coordinates, and risk level information in the object recognition results, the cleaning environment is divided into high-risk areas, medium-risk areas, and low-risk areas.
[0146] Areas containing spilled liquids, pet excrement, dense cables, or moving live targets are marked as high-risk areas; areas containing carpet edges, thresholds, or scattered small objects are marked as medium-risk areas; and the remaining open and flat areas are marked as low-risk areas.
[0147] Differentiated route planning strategies are generated for areas with different risk levels. For high-risk areas, detour routes are generated and alarm information is sent to user terminals. For medium-risk areas, deceleration routes are planned. For low-risk areas, efficient coverage routes are planned.
[0148] Among them, path planning takes into account the connectivity between areas and the optimization of the cleaning sequence to ensure the continuity and efficiency of the overall cleaning task.
[0149] Based on the identified floor material type and stain type, an adaptive cleaning parameter adjustment mechanism is built to dynamically adjust cleaning parameters including suction power, brush head speed and brush head height;
[0150] Specifically, a standard cleaning mode is used for hard floors, suction power is increased and travel speed is reduced for carpeted areas, a special cleaning module is used for liquid stains, and the number of repeated cleaning cycles is increased for particulate stains.
[0151] During the cleaning task, the cleaning effect and environmental changes are continuously monitored based on real-time sensor feedback information. When new obstacles are detected or the cleaning parameters deviate from expectations, the cleaning strategy and path planning are adjusted in real time.
[0152] Among them, pressure sensors monitor the ground contact status, and ultrasonic sensors monitor changes in obstacle distance to ensure the safety and effectiveness of the cleaning process;
[0153] Establish a cleaning task execution status evaluation mechanism to record the cleaning completion rate, time consumption, and energy consumption data of each area. Based on the evaluation results, optimize the parameter configuration and path planning of subsequent cleaning tasks to form an adaptive learning and continuous improvement closed-loop control system.
[0154] Successful cleaning strategy parameters are saved to a local database to provide a reference and optimization basis for subsequent cleaning tasks in similar scenarios.
[0155] This embodiment employs tiered control to avoid contact with hazardous substances and living targets, prevent equipment damage and malfunctions due to improper operation, and prevent damage to furniture and floors during cleaning. It adopts optimal cleaning strategies based on the characteristics of different areas; reduces repetitive cleaning and ineffective movement, shortening overall cleaning time; precisely adjusts cleaning intensity according to floor material and stain type; uses specialized cleaning modes for different stain types to ensure appropriate cleaning treatment for all types of areas; evaluates cleaning effectiveness in real time and adjusts strategy parameters promptly; automatically formulates cleaning strategies based on environmental perception results; continuously optimizes cleaning performance through historical data accumulation to cope with complex and ever-changing indoor environmental challenges; avoids impacting users' daily activities through risk grading; promptly reports to users in case of abnormal situations; and provides customized cleaning solutions based on the characteristics of the home environment. This tiered scene-adaptive planning and adaptive variational cleaning strategy represents a significant advancement in robotic vacuum cleaner technology, significantly improving cleaning safety, efficiency, and quality through intelligent environmental understanding and strategy adjustment.
[0156] In some possible embodiments of the present invention, the step of identifying residual stains on the ground based on the second image data and a preset cleaning effect recognition model includes:
[0157] The second image data is preprocessed, including image denoising, brightness equalization, and HSV color space conversion, to enhance the contrast between residual stains and the cleaned floor.
[0158] The preprocessed second image data is input into a cleaning effect recognition model based on YOLOv8, which integrates RGB features and HSV color space features and uses a multi-scale feature extraction network to identify different types of residual stains.
[0159] The identified residual stains were classified and labeled into four categories: dust residue, hair residue, liquid stain residue, and particulate residue. The pixel coordinates and contour information of each residual stain were extracted.
[0160] Based on the mapping relationship between pixel coordinates and the actual ground, the actual area of each residual stain is calculated, and the calculation result is compared with a preset first area threshold.
[0161] Based on the type and area of the residual stains, a corresponding re-cleaning instruction is generated, including the coordinates of the target cleaning area, the recommended cleaning mode, and the estimated cleaning time.
[0162] In this embodiment, by enhancing the features of the HSV color space, residual stains with low contrast in RGB images can be effectively identified, especially transparent liquid stains and fine dust, significantly improving the recognition accuracy. Residual stains are categorized and labeled by type, enabling subsequent cleaning strategies to employ the most suitable cleaning parameters for different stain types, avoiding a "one-size-fits-all" approach to cleaning. An area threshold judgment mechanism avoids repeated cleaning of small, meaningless stains while ensuring that truly necessary residual stains are effectively removed, improving overall cleaning efficiency. A closed-loop control system from visual detection to cleaning execution is established, giving the robot vacuum cleaner "inspection-remediation" capabilities similar to manual cleaning, improving user satisfaction. Objective visual verification replaces traditional time or path judgment, effectively reducing cleaning omission rates and ensuring the consistency and reliability of cleaning quality.
[0163] In some possible embodiments of the present invention, the step of triggering a re-cleaning procedure and automatically switching the corresponding cleaning mode according to the type of residual stain if the area of the residual stain exceeds a preset first area threshold includes:
[0164] The calculated residual stain area is compared with a preset first area threshold. When the residual stain area is greater than the first area threshold, the system automatically triggers the re-cleaning program and generates a re-cleaning task queue.
[0165] Based on the identified residual stains, the corresponding cleaning parameter configuration is matched from the preset cleaning mode database. The cleaning parameter configuration includes suction level, brush head rotation speed, travel speed and cleaning time.
[0166] Different cleaning modes are switched according to different types of residual stains, including: for liquid stains, the water absorption module is activated and the travel speed is reduced; for particulate matter, the suction power is increased to the maximum level; for hair, the frequency of the roller brush reversal is increased; and for dust, a reciprocating cleaning path is used.
[0167] Based on the spatial coordinate information of residual stains, a precise re-cleaning path is planned, enabling the robot vacuum cleaner to directly navigate to the stained area and perform localized deep cleaning according to the optimized cleaning trajectory.
[0168] The cleaning effect is monitored in real time during the re-cleaning process. Real-time feedback data of the cleaned area is obtained through onboard sensors. The re-cleaning program is automatically ended when the stains are detected to be removed or the preset cleaning time is reached.
[0169] In this embodiment, an automatic triggering mechanism based on area thresholds avoids manual intervention, enabling the robot vacuum to autonomously determine whether supplementary cleaning is needed, thus improving the device's intelligence. It matches specific cleaning parameters according to different stain types, avoiding the incomplete cleaning or energy waste that can occur with a uniform cleaning mode in traditional solutions, resulting in more precise cleaning. Precise path planning directly reaches the stained areas, avoiding repeated cleaning of the entire area, significantly shortening re-cleaning time and significantly improving overall cleaning efficiency. Localized, precise cleaning replaces large-area repetitive cleaning, effectively reducing battery consumption, extending the cleaning coverage area per charge, and improving device endurance. Closed-loop control ensures cleaning quality meets standards, reducing the frequency of manual supplementary cleaning by users, while a real-time monitoring mechanism ensures the controllability and reliability of the cleaning process, significantly improving user satisfaction.
[0170] In some possible embodiments of the present invention, the three-stream neural network fusion process employs a multi-scale adaptive weight allocation mechanism, and the formula for calculating the fused feature vector is as follows:
[0171]
[0172] in:
[0173] F fusion (x, y, z) represents the final multimodal fusion feature vector, where x represents the image modal input data, y represents the three-dimensional spatial modal input data, and z represents the audio modal input data;
[0174] This represents the summation over R different scale / resolution levels, where R represents the total number of layers in the multi-scale fusion, and r represents the current scale index.
[0175] μ r This represents the global weight coefficient at the r-th scale, used to balance the importance of features at different scales, satisfying...
[0176] F image (x, r) represents the feature vector of the image modality at the r-th scale. F represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (z, r) represents the feature vector of the audio modality at the r-th scale;
[0177] α image (r) represents the feature vector of the image modality at scale r. α represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (r) represents the attention weight of the audio modality at the r-th scale;
[0178] β is the regularization intensity coefficient, γ is the sensitivity adjustment parameter, tanh() is the hyperbolic tangent activation function used for numerical stability; ||·|| represents the vector norm (usually the L2 norm); F m Let m be the eigenvector of the m-th mode. Let be the mean of the feature vector of the m-th modality; i, s, and a represent the image, spatial, and sound modalities, respectively.
[0179] δ represents the confidence level moderating coefficient, σ() is the Sigmoid activation function, and A conf (x, y, z) is the multimodal confidence function. Let ||·||2 be the second-order gradient operator (Hessian matrix), and ||·||2 be the L2 norm normalization.
[0180] In this embodiment, by Achieve feature fusion across multiple resolution levels, with each scale using μ r By adjusting the weights, it is possible to capture feature information at different levels, from fine-grained to coarse-grained; α image (r), α audio (r) Implements dynamic modality importance allocation, adaptively adjusting the contribution of each modality based on the current scene; by calculating the deviation between each modality feature and its mean, and using the tanh() function for smoothing, it prevents any one modality from excessively dominating the fusion result; by analyzing the second derivative information of multimodal confidence, it dynamically adjusts the stability of the fusion process, improving robustness to noise and anomalies. This embodiment enhances the recognition capability of objects of different sizes through a multi-scale processing mechanism. The confidence-guided term can adaptively adjust the fusion strategy, enhancing the robustness of feature fusion in regions with drastic confidence changes, significantly improving the recognition accuracy and stability of diverse objects in complex scenes.
[0181] In some possible embodiments of the present invention, the steps of performing a cleaning task using an adaptive variational cleaning strategy include:
[0182] An intelligent parameter adjustment algorithm based on multi-dimensional pollution characteristics is adopted, and its cleaning intensity adjustment formula is as follows:
[0183]
[0184] in:
[0185] I clean (u, v, t) represents the sweeping intensity at position (u, v) at time t; I base Based on the sweeping intensity constant;
[0186] D(u,v,t) represents the pollution density distribution function at location (u,v) at time t;
[0187] and Let represent the second-order partial derivatives of the pollution density in the u and v directions, respectively;
[0188] ξ and η are gradient sensitivity coefficients used to control the degree of response of cleaning intensity to changes in contamination density gradient;
[0189] C effect (u, v, th) represents the cleaning effect evaluation value of position (u, v) at the previous time t-1;
[0190] λ is the historical effect decay coefficient, used to avoid over-cleaning areas that have already been thoroughly cleaned;
[0191] κ is the dirt type adjustment coefficient, used to control the weight of the influence of different dirt types on cleaning intensity;
[0192] H dirt (u, v, t) represents the dirt hardness characteristic function at position (u, v) at time t, reflecting the adhesion and difficulty of cleaning the dirt;
[0193] V robot (t) represents the robot's actual speed at time t;
[0194] V optimal (u, v) represents the optimal cleaning speed at position (u, v);
[0195] v is the velocity normalization constant, which prevents division by zero and adjusts the sensitivity of the velocity term;
[0196] Ψ(·) is the velocity adaptation function, defined as follows: Where ρ is the velocity sensitivity parameter;
[0197] exp(·) is an exponential function used to achieve non-linear adjustment of the cleaning effect;
[0198] In this embodiment, by introducing a dirt hardness characteristic term, the cleaning intensity can be adaptively adjusted according to the physical properties of different types of dirt, and the speed adaptation term ensures the optimal match between the cleaning intensity and the robot's moving speed, thereby achieving refined cleaning control for complex polluted environments and significantly improving cleaning quality and energy utilization efficiency.
[0199] In some possible embodiments of the present invention, the method further includes the step of establishing an intelligent learning system based on user habits; wherein the intelligent learning system employs a dynamic preference evaluation algorithm based on multi-dimensional feature fusion, and its comprehensive preference evaluation formula is:
[0200]
[0201] in:
[0202] P total (s, a) represents the overall user preference score for performing action a in state s;
[0203] M represents the total number of users;
[0204] ω m Let the weight coefficient of the m-th user be denoted as , satisfying
[0205] N m This represents the number of historical interaction records for the m-th user;
[0206] P m (s, a, n) represents the preference rating of the m-th user for action a in state s during the n-th interaction;
[0207] φ n This represents the confidence coefficient of the nth interaction record;
[0208] t current Indicates the current time;
[0209] t n This indicates the time when the nth interaction occurs;
[0210] This represents the time decay variance of the m-th user's preference;
[0211] Θ(s, a, m) represents the personalized fitness function of user m for state-action pair (s, a), reflecting the consistency of user behavior;
[0212] ζ is the periodic adjustment coefficient, used to control the strength of the influence of time periodicity on preferences;
[0213] T cycle A periodic time window (such as 24 hours, 7 days, etc.) representing user behavior;
[0214] cos(·) is a cosine function used to model the periodic variation characteristics of user preferences;
[0215] mod represents the modulo operation;
[0216] Γ context (s,a,t) current ) represents the context adaptation function, defined as:
[0217]
[0218] Where L is the number of environmental feature dimensions, E l (t current () represents the value of the l-th environmental feature at the current moment. τ represents the historical mean of the l-th environmental feature. l and These are the weighting coefficient and standardization parameter for the l-th environmental feature, respectively;
[0219] exp(·) is an exponential function, and tanh(·) is a hyperbolic tangent function;
[0220] This embodiment enhances the ability to distinguish different user behavior patterns through a personalized fitness function, captures the changing patterns of user preferences over different time periods through a periodic adjustment term, and considers the dynamic impact of environmental factors on user preferences through a contextual fitness function. This achieves accurate modeling of user preferences across multiple dimensions and time scales, significantly improving the system's adaptability to complex user needs and its predictive accuracy.
[0221] Please see Figure 2 Another aspect of the present invention provides a control system for a sweeping robot, for executing a control method for a sweeping robot, comprising: a sweeping robot and a server;
[0222] The server is configured as follows:
[0223] Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner;
[0224] The first image data, the first three-dimensional spatial data, and the first sound data are processed using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment.
[0225] The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result;
[0226] Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task;
[0227] After the cleaning task is completed, a second image of the cleaned area is collected;
[0228] Based on the second image data and the preset cleaning effect recognition model, identify residual stains on the ground;
[0229] If the area of the residual stain exceeds a preset first area threshold, a re-cleaning procedure is triggered, and the corresponding cleaning mode is automatically switched according to the type of residual stain.
[0230] It should be known that, Figure 2The block diagram of the control system for the robotic vacuum cleaner shown is for illustrative purposes only, and the number of modules shown does not limit the scope of protection of this invention. The control system for the robotic vacuum cleaner provided in this embodiment can be used to execute various embodiments of the corresponding control method for the robotic vacuum cleaner. For specific implementation details, please refer to the descriptions of the respective method embodiments, which will not be repeated here.
[0231] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0232] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0233] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0234] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0235] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0236] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0237] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0238] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0239] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions without departing from the spirit and scope of the present invention, and various modifications and alterations can be made, including combinations of the different functions and implementation steps described above, as well as software and hardware implementation methods, all of which are within the protection scope of the present invention.
Claims
1. A control method for a sweeping robot, characterized in that, include: Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner; The first image data, the first three-dimensional spatial data, and the first sound data are processed using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment. The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result; Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task; After the cleaning task is completed, a second image of the cleaned area is collected; Based on the second image data and the preset cleaning effect recognition model, identify residual stains on the ground; If the area of the residual stain exceeds a preset first area threshold, a re-cleaning procedure is triggered, and the corresponding cleaning mode is automatically switched according to the type of residual stain.
2. The control method for a sweeping robot according to claim 1, characterized in that, The step of processing the first image data, the first three-dimensional spatial data, and the first sound data using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment includes: The first image data is input into a first-stream convolutional neural network, the first three-dimensional spatial data is input into a second-stream point cloud processing network, and the first sound data is input into a third-stream temporal neural network, and the initial feature representations of each modality are extracted in parallel. The initial feature representations output by each stream network are processed to unify the feature dimensions, mapping the feature vectors of different modalities to the same feature space dimension, ensuring the compatibility of subsequent fusion processing; The importance weights of features within each modality are calculated separately using a self-attention mechanism, thereby enhancing the expression strength of key features within each modality and suppressing interference from redundant information. A cross-modal attention mechanism is used to calculate the correlation matrix between different modalities, and the fusion weights of each modal feature are adaptively allocated according to the current environmental scene to achieve the integration of complementary information between modalities. The weighted multimodal features are cascaded and fused, and a comprehensive environmental feature representation containing object attribute features, spatial features, and sound attribute features is output through a fully connected layer, which serves as the input data for subsequent object recognition and classification.
3. The control method for a sweeping robot according to claim 2, characterized in that, The step of adaptively fusing object attribute features, spatial features, and sound attribute features through a cross-modal attention mechanism to obtain object recognition results includes: A cross-modal correlation mapping table is constructed, and the correlation strength matrix between different modal features is established by calculating the semantic similarity between object attribute features, spatial features and sound attribute features; Based on the correlation strength matrix, attention weight coefficients for each modality are dynamically generated, and the fusion ratio of object attribute features, spatial features, and sound attribute features is adaptively adjusted according to the reliability and importance of information of each modality in the current environment. The weighted multimodal features are fused at the feature level, and a unified feature representation containing multimodal information is generated through feature concatenation and dimensionality reduction. The unified feature representation is input into a pre-trained object recognition classifier, and feature parsing and pattern matching are performed through a deep neural network to output the object category probability distribution and confidence score. By combining the confidence score and the preset recognition threshold, the final object recognition result is determined, and structured recognition information containing object type, location coordinates, motion state and risk level is generated.
4. The control method for the sweeping robot according to claim 3, characterized in that, The steps of performing hierarchical scene adaptive planning and adopting an adaptive variational cleaning strategy to perform the cleaning task based on the object recognition results include: Based on the object type, location coordinates, and risk level information in the object recognition results, the cleaning environment is divided into high-risk areas, medium-risk areas, and low-risk areas. Differentiated route planning strategies are generated for areas with different risk levels. For high-risk areas, detour routes are generated and alarm information is sent to user terminals. For medium-risk areas, deceleration routes are planned. For low-risk areas, efficient coverage routes are planned. Based on the identified floor material type and stain type, an adaptive cleaning parameter adjustment mechanism is built to dynamically adjust cleaning parameters including suction power, brush head speed and brush head height; During the cleaning task, the cleaning effect and environmental changes are continuously monitored based on real-time sensor feedback information. When new obstacles are detected or the cleaning parameters deviate from expectations, the cleaning strategy and path planning are adjusted in real time. Establish a cleaning task execution status evaluation mechanism to record the cleaning completion rate, time consumption, and energy consumption data of each area. Based on the evaluation results, optimize the parameter configuration and path planning of subsequent cleaning tasks to form an adaptive learning and continuous improvement closed-loop control system.
5. The control method for a sweeping robot according to claim 4, characterized in that, The step of identifying residual stains on the ground based on the second image data and a preset cleaning effect recognition model includes: The second image data is preprocessed, including image denoising, brightness equalization, and HSV color space conversion, to enhance the contrast between residual stains and the cleaned floor. The preprocessed second image data is input into a cleaning effect recognition model based on YOLOv8, which integrates RGB features and HSV color space features and uses a multi-scale feature extraction network to identify different types of residual stains. The identified residual stains were classified and labeled into four categories: dust residue, hair residue, liquid stain residue, and particulate residue. The pixel coordinates and contour information of each residual stain were extracted. Based on the mapping relationship between pixel coordinates and the actual ground, the actual area of each residual stain is calculated, and the calculation result is compared with a preset first area threshold. Based on the type and area of the residual stains, a corresponding re-cleaning instruction is generated, including the coordinates of the target cleaning area, the recommended cleaning mode, and the estimated cleaning time.
6. The control method for a sweeping robot according to claim 5, characterized in that, The step of triggering a re-cleaning procedure and automatically switching the corresponding cleaning mode according to the type of residual stain if the area of the residual stain exceeds a preset first area threshold includes: The calculated residual stain area is compared with a preset first area threshold. When the residual stain area is greater than the first area threshold, the system automatically triggers the re-cleaning program and generates a re-cleaning task queue. Based on the identified residual stains, the corresponding cleaning parameter configuration is matched from the preset cleaning mode database. The cleaning parameter configuration includes suction level, brush head rotation speed, travel speed and cleaning time. Different cleaning modes are switched according to different types of residual stains; Based on the spatial coordinate information of residual stains, a precise re-cleaning path is planned, enabling the robot vacuum cleaner to directly navigate to the stained area and perform localized deep cleaning according to the optimized cleaning trajectory. The cleaning effect is monitored in real time during the re-cleaning process. Real-time feedback data of the cleaned area is obtained through onboard sensors. The re-cleaning program is automatically ended when the stains are detected to be removed or the preset cleaning time is reached.
7. The control method for a sweeping robot according to claim 6, characterized in that, The three-stream neural network fusion process employs a multi-scale adaptive weight allocation mechanism, and the formula for calculating its fused feature vector is as follows: in: F fusion (x, y, z) represents the final multimodal fusion feature vector, where x represents the image modal input data, y represents the three-dimensional spatial modal input data, and z represents the audio modal input data; This represents the summation over R different scale / resolution levels, where R represents the total number of layers in the multi-scale fusion, and r represents the current scale index. μ r This represents the global weight coefficient at the r-th scale, used to balance the importance of features at different scales, satisfying... F image (x, r) represents the feature vector of the image modality at the r-th scale. F represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (z, r) represents the feature vector of the audio modality at the r-th scale; α image (r) represents the feature vector of the image modality at scale r. α represents the eigenvector of a three-dimensional spatial mode at the r-th scale. audio (r) represents the attention weight of the audio modality at the r-th scale; β is the regularization intensity coefficient, γ is the sensitivity adjustment parameter, and tanh() is the hyperbolic tangent activation function. Used for numerical stability; ||·|| denotes the vector norm (usually the L2 norm); F m Let m be the eigenvector of the m-th mode. Let be the mean of the feature vector of the m-th modality; i, s, and a represent the image, spatial, and sound modalities, respectively. δ represents the confidence level moderating coefficient, σ() is the Sigmoid activation function, and A conf (x, y, z) is the multimodal confidence function. Let ||·||2 be the second-order gradient operator (Hessian matrix), and ||·||2 be the L2 norm normalization.
8. The control method for a sweeping robot according to claim 7, characterized in that, The steps for performing a cleaning task using an adaptive variational cleaning strategy include: An intelligent parameter adjustment algorithm based on multi-dimensional pollution characteristics is adopted, and its cleaning intensity adjustment formula is as follows: in: I clean (u, v, t) represents the sweeping intensity at position (u, v) at time t; I base Based on the basic cleaning intensity constant; D(u, v, t) represents the pollution density distribution function at location (u, v) at time t; and Let represent the second-order partial derivatives of the pollution density in the u and v directions, respectively; ξ and η are gradient sensitivity coefficients used to control the degree of response of cleaning intensity to changes in contamination density gradient; C effect (u, v, th) represents the cleaning effect evaluation value of position (u, v) at the previous time t-1; λ is the historical effect decay coefficient, used to avoid over-cleaning areas that have already been thoroughly cleaned; κ is the dirt type adjustment coefficient, used to control the weight of the influence of different dirt types on cleaning intensity; H dirt (u, v, t) represents the dirt hardness characteristic function at position (u, v) at time t, reflecting the adhesion and difficulty of cleaning the dirt; V robot (t) represents the robot's actual speed at time t; V optimal (u, v) represents the optimal cleaning speed at position (u, v); v is the velocity normalization constant, which prevents division by zero and adjusts the sensitivity of the velocity term; Ψ(·) is the velocity adaptation function, defined as follows: Where ρ is the velocity sensitivity parameter; exp(·) is an exponential function used to achieve non-linear adjustment of the cleaning effect.
9. The control method for a sweeping robot according to claim 8, characterized in that, It also includes the step of establishing an intelligent learning system based on user habits; wherein, the intelligent learning system adopts a dynamic preference evaluation algorithm based on multi-dimensional feature fusion, and its comprehensive preference evaluation formula is: in: P total (s, a) represents the overall user preference score for performing action a in state s; M represents the total number of users; ω m Let the weight coefficient of the m-th user be denoted as , satisfying N m This represents the number of historical interaction records for the m-th user; P m (s, a, n) represents the preference rating of the m-th user for action a in state s during the n-th interaction; φ n This represents the confidence coefficient of the nth interaction record; t current Indicates the current time; t n This indicates the time when the nth interaction occurs; This represents the time decay variance of the m-th user's preference; Θ(s, a, m) represents the personalized fitness function of user m for state-action pair (s, a), reflecting the consistency of user behavior; ζ is the periodic adjustment coefficient, used to control the strength of the influence of time periodicity on preferences; T cycle A periodic time window (such as 24 hours, 7 days, etc.) representing user behavior; cos(·) is a cosine function used to model the periodic variation characteristics of user preferences; mod represents the modulo operation; Γ context (s, a, t) current ) represents the context adaptation function, defined as: Where L is the number of environmental feature dimensions, E l (t current () represents the value of the l-th environmental feature at the current moment. τ represents the historical mean of the l-th environmental feature. l and These are the weighting coefficient and standardization parameter for the l-th environmental feature, respectively; exp(·) is an exponential function, and tanh(·) is a hyperbolic tangent function.
10. A control system for a sweeping robot, used to execute the control method for a sweeping robot as described in any one of claims 1 to 9, characterized in that, include: Robotic vacuum cleaners and servers; The server is configured as follows: Acquire first image data, first three-dimensional spatial data, and first sound data of the environment surrounding the robot vacuum cleaner; The first image data, the first three-dimensional spatial data, and the first sound data are processed using a three-stream neural network with a fusion attention mechanism to extract object attribute features, spatial features, and sound attribute features in the environment. The object attribute features, spatial features, and sound attribute features are adaptively fused through a cross-modal attention mechanism to obtain the object recognition result; Based on the object recognition results, hierarchical scene adaptive planning is performed and an adaptive variational cleaning strategy is adopted to perform the cleaning task; After the cleaning task is completed, a second image of the cleaned area is collected; Based on the second image data and the preset cleaning effect recognition model, identify residual stains on the ground; If the area of the residual stain exceeds a preset first area threshold, a re-cleaning procedure is triggered, and the corresponding cleaning mode is automatically switched according to the type of residual stain.
Citation Information
Cited By
Autonomous cleaning method and device for ground stains and storage medium
CN121392503A
Hand-held scrubber control method and system based on time sequence behavior and vision
CN121890914A