Autonomous cleaning method and device for ground stains and storage medium

By generating multimodal data and determining the type of stain composition, and combining cleaning strategy database to match cleaning parameters, the cleaning equipment is controlled to solve the problem of poor performance of robot vacuum cleaners in handling liquid stains, achieving efficient and precise autonomous cleaning.

CN121392503AActive Publication Date: 2026-01-23SHENZHEN SHUNTER TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511950104.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-01-23
Estimated Expiration
2045-12-23

AI Technical Summary

Technical Problem

Existing robotic vacuum cleaners are not very effective at handling liquid stains when cleaning floors. The use of a uniform cleaning strategy leads to poor cleaning results, making it difficult to meet the needs of efficient and precise cleaning in household settings.

Method used

By collecting raw data from the area to be processed, multimodal data is generated according to the information processing rules corresponding to the data type. The correlation between material characteristics is analyzed to determine the composition category of the stain. Multimodal data is spliced ​​and aligned through intermodal pixel adaptation strategy to generate a segmentation mask to define the location information of the stain. The target cleaning parameters are matched with the cleaning strategy database, and the parameter adaptability is verified to form a target cleaning strategy and control the cleaning equipment to perform cleaning.

Benefits of technology

It improves the accuracy of stain recognition and the adaptability of cleaning parameters, enabling precise and intelligent cleaning of the area to be treated, solving the problem of poor cleaning effect, and improving the efficiency of cleaning operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392503A_ABST
    Figure CN121392503A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic cleaning method and equipment for ground stains and a storage medium, and relates to the technical field of cleaning robots. The method comprises the steps that after original data of a to-be-processed area is collected, processing is conducted according to an information processing rule corresponding to a data type, and multi-modal data is obtained; splicing and aligning the multi-modal data through an inter-modal pixel adaptation strategy to generate fused data, analyzing a substance characteristic association relationship in the fused data to judge the stain component category, extracting high-level characteristics representing stain distribution of each modal, performing weighted fusion, and generating a segmentation mask used for defining stain position information; processing the multi-modal data according to a decision-making level fusion rule, matching a target cleaning parameter in combination with a cleaning strategy database, verifying the suitability of the parameter with the stain component category and position information, and forming a target cleaning strategy; and according to the target cleaning strategy, the cleaning equipment is controlled to complete cleaning of the to-be-treated area, the problem that the cleaning effect is poor is solved, and the efficient and accurate ground stain cleaning effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cleaning robots, and particularly relates to a self-cleaning method for ground stains, a device and a storage medium. BACKGROUND

[0002] In the scenario of cleaning the ground, the core value of the sweeping robot as a convenient cleaning device lies in efficiently completing the ground cleaning task and reducing manual intervention.

[0003] In the related art, the sweeping robot performs indiscriminate sweeping according to a preset path, and can only take a simple processing mode of ignoring or single wet mopping for liquid stains, and a unified cleaning strategy is adopted for all types of stains, which leads to poor cleaning effect and is difficult to meet the efficient and accurate cleaning demand in the home scenario.

[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0005] The main purpose of the present application is to provide a self-cleaning method for ground stains, a device and a storage medium, which aims to solve the technical problem of poor cleaning effect.

[0006] To achieve the above purpose, the present application provides a self-cleaning method for ground stains, which comprises: Collecting original data of a to-be-processed area; According to an information processing rule corresponding to a data type of the original data, processing the original data to obtain multi-modal data; According to a modal inter-pixel adaptation strategy, splicing and aligning the multi-modal data to obtain fusion data, and analyzing the correlation of the characteristics of each substance in the fusion data to determine the classification of the stains; Extracting high-level features of each modality in the multi-modal data representing stain distribution and weighting fusion to generate a segmentation mask to define the position information of the stains through the segmentation mask; According to a decision-level fusion rule, processing the multi-modal data, combining a cleaning strategy database to match target cleaning parameters, and verifying the target cleaning parameters and the classification of the stains, the position information of the stains to obtain a target cleaning strategy; According to the target cleaning strategy, controlling a cleaning device to clean the to-be-processed area.

[0007] In an embodiment, the original data is divided according to the data type, and the corresponding rule is screened in the information processing rule library to determine the information processing rule corresponding to each data type; The visual data in the original data is denoised by an edge-preserving algorithm, the spectral data in the original data is baseline corrected, and the acoustic data in the original data is filtered, to obtain processed single-modality data; The processed single-modality data are associated and integrated according to a preset spatio-temporal alignment rule, to output multi-modality data.

[0008] In an embodiment, the aligned multi-modality data are subjected to resolution normalization processing, and pixel coordinates of each modality data are accurately mapped based on the inter-modality pixel adaptation strategy, to generate multi-modality pixel-level data; The multi-modality pixel-level data are subjected to channel splicing and fusion according to a preset channel combination rule, to obtain initial fusion data; The initial fusion data are subjected to format regularization processing and verification, to output the fusion data.

[0009] In an embodiment, features reflecting stain characteristic information are extracted from pixel information corresponding to the stain category, to form a multi-modality substance characteristic feature set; The complementary association relationship between different substance characteristics in the multi-modality substance characteristic feature set is determined through association analysis; The complementary association relationship is inferred according to a preset mapping rule of substance characteristics and categories, to determine the stain category.

[0010] In an embodiment, the high-level features representing stain outline, stain morphology and substance characteristics are extracted according to the modality of the multi-modality data, to obtain a high-level feature set independent of each modality; The high-level feature sets corresponding to each modality are weighted and fused according to the contribution of each modality to stain identification, to output fusion features; The fusion features are subjected to pixel-level classification using a semantic segmentation decoder, to label the stain attribution attribute of each pixel, to obtain the segmentation mask consistent with the size of the stain in the original image.

[0011] In an embodiment, the segmentation mask is analyzed, and the complete outline, boundary range, spatial coordinates and area parameters of the stain are extracted, to integrate to obtain stain position feature data; The stain position feature data and the stain category are associated and fused, and the attribute information corresponding to the stain is supplemented, to integrate to obtain stain comprehensive information data; According to the stain comprehensive information data, the position change trend of the stain is predicted, to generate position information describing the stain distribution.

[0012] In an embodiment, the position parameters, the category, the area data, and the characteristic confidence of the stains in the multi-modal data are parsed to form stain attribute data; According to a preset validity determination threshold, abnormal data and invalid information in the stain attribute data are eliminated, and key stain attribute information is screened; The stain key attribute information is matched with the database by calling a pre-stored cleaning strategy database to obtain a candidate cleaning parameter set; The candidate parameters are adaptively evaluated and prioritized in combination with the current ground material information and the current state of the equipment to obtain adaptive parameters, and the adaptive parameters are fused according to the decision-level fusion rule to generate the target cleaning parameter.

[0013] In an embodiment, the target cleaning parameter and the category of the stain and the position information are adaptively verified to obtain a parameter adaptive verification result; If there is an adaptive deviation in the parameter adaptive verification result, the cleaning parameter is finely tuned according to the category of the stain and the position information to generate a calibrated cleaning execution parameter; According to the calibrated cleaning execution parameter, the target cleaning strategy is generated in combination with the current position information of the equipment and the environmental map.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes an autonomous cleaning equipment, which comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the autonomous cleaning method of the ground stain as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the autonomous cleaning method of the ground stain as described above.

[0016] The application provides an autonomous cleaning method for ground stains, comprising collecting original data of a to-be-processed area, processing multiple modal data according to information processing rules corresponding to the data types of the original data, splicing and aligning the multiple modal data to generate fusion data through an inter-modal pixel adaptation strategy, determining stain categories according to the correlation of the material characteristics in the fusion data, extracting high-level features of each modality in the multiple modal data to generate a segmentation mask through weighted fusion to define stain position information, processing the multiple modal data according to a decision-level fusion rule, matching target cleaning parameters in combination with a cleaning strategy database, checking the adaptability of the parameters to the stain categories and the position information to form a target cleaning strategy, and finally controlling a cleaning device to complete cleaning of the to-be-processed area according to the target cleaning strategy, thereby solving the technical problems of poor cleaning effect, waste of resources caused by ambiguous stain category determination, inaccurate position definition and lack of pertinence of cleaning parameter matching in traditional cleaning schemes, improving the accuracy of stain recognition, the adaptability of cleaning parameters and the efficiency of cleaning operation, and realizing precise and intelligent cleaning of the to-be-processed area.

[0017] In summary, the application generates multiple modal data by synchronously collecting original data and preprocessing, determines stain categories through inter-modal pixel adaptation splicing to obtain fusion data, extracts high-level features of each modality to generate a segmentation mask through weighted fusion, and obtains a target cleaning strategy by matching and checking target cleaning parameters through decision-level fusion, thereby solving the technical problem of poor cleaning effect, improving the collaborative efficiency of stain recognition and cleaning execution, and realizing efficient and precise autonomous cleaning effect. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate an embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0020] Figure 1 The flowchart of the first embodiment of the autonomous cleaning method for ground stains of the present application; Figure 2 The flowchart of the seventh embodiment of the autonomous cleaning method for ground stains of the present application; Figure 3 The block diagram of the system composition of the present application; Figure 4 The structural diagram of the autonomous cleaning device of the present application.

[0021] The object, functional characteristics and advantages of the present application will be further illustrated in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0022] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.

[0023] In the related art, the sweeping robot performs indiscriminate cleaning according to a preset path, and can only take a simple processing mode of ignoring or single wet mopping for liquid stains, and a unified cleaning strategy is adopted for all types of stains, resulting in poor cleaning effect and being difficult to meet the efficient and accurate cleaning demand in a home scenario.

[0024] The present application provides a solution: first, collecting original data of a to-be-processed area, then processing the original data according to an information processing rule corresponding to a data type of the original data to obtain multi-modal data, then splicing and aligning the multi-modal data according to an inter-modal pixel adaptation strategy to obtain fusion data, and analyzing the correlation of the characteristics of each substance in the fusion data to determine the classification of the stain, then extracting the high-level features of each modality in the multi-modal data representing the distribution of the stain and weighting and fusing to generate a segmentation mask to define the position information of the stain through the segmentation mask, then processing the multi-modal data according to a decision-level fusion rule, matching target cleaning parameters in combination with a cleaning strategy database, and verifying the target cleaning parameters and the classification of the stain, the position information of the stain to obtain a target cleaning strategy, and finally controlling a cleaning device to clean the to-be-processed area according to the target cleaning strategy.

[0025] It should be noted that the execution subject of the present embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an autonomous cleaning device, etc. capable of realizing the above functions. The present embodiment and the following embodiments will be described below taking the autonomous cleaning device as an example.

[0026] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the drawings and specific embodiments of the present application.

[0027] The present embodiment provides an autonomous cleaning method for ground stains, which is described in detail with reference to Figure 1 , Figure 1 The present embodiment provides an autonomous cleaning method for ground stains, which is described in detail with reference to

[0028] In the present embodiment, the autonomous cleaning method for ground stains includes steps S10-S60: Step S10, collecting original data of a to-be-processed area.

[0029] In the embodiment, the to-be-processed region refers to a specific spatial range that needs to be cleaned. The raw data refers to a set of multi-modal basic information directly obtained from the to-be-processed region without any processing.

[0030] As an optional implementation, the to-be-processed region is spatially divided to determine the boundary range and collection priority of each sub-region. Then, data collection is carried out for each sub-region according to the preset multi-modal type one by one, ensuring that complete basic information of each modality is collected. Finally, all collected basic information is integrated according to the sub-region and modality type classification to form structured and comprehensive raw data. The data collected by this method is complete and comprehensive in modality information, and the basic quality is guaranteed.

[0031] As another optional implementation, the core collection range and key modality type of the to-be-processed region are determined first, focusing on the core information dimension that is strongly related to stain identification and cleaning decision. Fast collection is carried out according to the modality priority one by one, and only key basic information is collected for each core modality. Only basic effectiveness verification is performed during the collection process to confirm that the data is not severely distorted. All collected core basic information is integrated to form raw data focusing on key dimensions. The collection process of this method is simple, fast in operation, low in resource consumption, and can quickly complete data acquisition.

[0032] In step S20, the raw data is processed according to the information processing rule corresponding to the data type of the raw data to obtain multi-modal data.

[0033] In the embodiment, the data type refers to the category of the raw data divided according to the information carrier and characteristics. The information processing rule refers to the standardized processing specification preset for different data types to optimize data quality and extract effective information. The multi-modal data refers to a structured data set formed by classified processing and integration, containing multiple types of effective information.

[0034] As an optional implementation, all data types contained in the original data and the information characteristics of each type are comprehensively analyzed. Then, according to the preset mapping relationship between the data type and the information processing rule, a dedicated processing rule is matched for each type of data. Subsequently, according to the matched rule, each type of original data is processed one by one. The visual data is optimized in picture quality and the detail features are strengthened through a noise reduction algorithm. The spectral data is purified through a calibration process to eliminate background interference and extract effective band signals. The acoustic data is processed through filtering to separate effective signals and noise. During the processing, the data quality is checked in real time, and the processing result is verified to see whether it meets the rule standard. Deviation data is corrected in real time. Finally, all the processed effective data is classified and integrated according to the modal type, forming a multi-modal data with complete structure and reliable quality. The data processing of this method is highly targeted and deep, and the generated multi-modal data is high in purity and good in information integrity.

[0035] As another optional implementation, the core data type and the key information segment in the original data are extracted first. Then, according to the simplified information processing rule mapping relationship, a suitable quick processing rule is matched for the core data type. The core data type is processed in batches according to the rule. The visual data is only subjected to basic noise reduction processing. The spectral data is quickly calibrated for the main effective band. The acoustic data is simplified in the filtering process to retain the core signal. Then, the processed core data is quickly integrated, and necessary modal identification information is supplemented to form multi-modal data focusing on key information. The processing flow of this method is simple, fast in operation, and low in resource consumption, and can quickly complete data processing.

[0036] In step S30, the multi-modal data after alignment is spliced according to the inter-modal pixel adaptation strategy to obtain fusion data, and the correlation of the characteristics of each substance in the fusion data is analyzed to determine the composition category of the stain.

[0037] In this embodiment, the inter-modal pixel adaptation strategy refers to a standardized rule preset for realizing spatial alignment and feature matching of different modal data at the pixel level. The fusion data refers to a unified data body formed by integrating the advantages of each modal after accurate pixel-level alignment according to the adaptation strategy. The substance characteristic feature refers to quantitative or qualitative information representing the inherent properties of a substance. The composition category of the stain refers to the determination result of the attribution of the substance composition of the stain.

[0038] As an optional implementation, the core requirements of the inter-modal pixel adaptation strategy are determined, including the pixel space coordinate calibration standard, feature matching threshold, and alignment priority of different modal data. Then, the multi-modal data after alignment are subjected to pixel-by-pixel precise alignment operation according to the strategy, the feature consistency of the pixels at the same spatial position of each modality is verified, and the coordinate offset and feature misalignment problem is corrected, so that the pixel information of different modalities is accurately corresponding in the spatial dimension and the feature dimension. Subsequently, all the aligned modal data are integrated to generate fusion data. Then, various material characteristic features in the fusion data are extracted, including the spectral response features related to the composition and the texture features related to the state, the internal logical correlation between the features is analyzed, and the reliability of the correlation results is strengthened through multiple rounds of feature cross-validation. Finally, according to the correlation analysis conclusion and the preset composition classification determination standard, the composition classification of the stain is comprehensively determined based on the feature matching degree. This method is accurate in pixel alignment, comprehensive in feature correlation analysis, and has high accuracy and reliability in composition classification determination.

[0039] As another optional implementation, the core alignment indicators of the inter-modal pixel adaptation strategy are extracted first, focusing on the key spatial regions and core feature dimensions, and the multi-modal data after alignment are subjected to batch alignment operation according to the core indicators, the corresponding relationship of the modal data in the main spatial region is verified, and the obvious coordinate offset problem is corrected. Subsequently, the batch-aligned modal data are integrated to form fusion data focusing on core information. Then, only the core material characteristic features strongly related to the composition classification in the fusion data are extracted, and the feature dimension is simplified. By analyzing the direct correlation between the core features, without multiple rounds of cross-validation, the composition classification of the stain is directly output according to the correlation results and the simplified composition classification determination standard. The processing flow of this method is simple and efficient, and the batch operation greatly shortens the processing period and reduces resource consumption.

[0040] In step S40, the high-level features of each modality representing the distribution of the stain in the multi-modal data are extracted and weighted fused to generate a segmentation mask, so as to define the position information of the stain through the segmentation mask.

[0041] In this embodiment, the high-level features of each modality representing the distribution of the stain refer to the core data features in each modality that can deeply reflect the spatial distribution rule of the stain. The segmentation mask refers to the feature data body whose stain attribution, size, and stain in the original data are consistent. Defining the position information of the stain refers to the process of clearly defining the position-related data of the spatial coordinates, boundary range, and contour shape of the stain through the segmentation mask.

[0042] As an optional implementation, the core components of each modality in the multi-modal data are analyzed, the high-level features of each modality representing the stain distribution are extracted, the contribution of each feature to the representation of the stain distribution is analyzed, and the adaptive weight is allocated according to the contribution, so as to ensure that the key features have a higher weight ratio. Then, the high-level features of each modality are aligned one by one according to the weight ratio, the features of the same spatial position are superimposed and fused, and the unified features after fusion are generated. Based on the unified features after fusion, the pixels are classified and labeled according to the pixel attribution, the stain pixels and non-stain pixels are distinguished, and a complete segmentation mask is generated. By analyzing the pixel distribution rule of the segmentation mask, the related data of the outline coordinates, boundary range and core region position of the stain are extracted, and the position information of the stain is accurately defined. This method has comprehensive feature extraction, accurate weight allocation, high completeness and accuracy of the segmentation mask, and strong position definition accuracy.

[0043] As another optional implementation, the core high-level features of each modality representing the stain distribution in the multi-modal data are first extracted, and the basic weight is uniformly allocated according to the modality type. The core high-level features of each modality are batch superimposed and fused, and only the spatial range is roughly matched to ensure the fusion direction is accurate. Based on the features after batch fusion, the core region and main boundary of the stain pixels are labeled, and a segmentation mask focusing on the core distribution is generated. By analyzing the core pixel distribution of the segmentation mask, the key position data such as the approximate spatial coordinates and main boundary range of the stain are extracted, and the position information of the stain is defined. This method has a simple feature extraction and fusion process, fast processing speed, low resource consumption, and can quickly output the position information.

[0044] In step S50, the multi-modal data is processed according to the decision-level fusion rule, the target cleaning parameters are matched by combining the cleaning strategy database, and the target cleaning parameters are verified with the classification of the stain and the position information of the stain, to obtain the target cleaning strategy.

[0045] In this embodiment, the decision-level fusion rule refers to a standardized rule system for high-order logical integration and decision output of multi-modal data. The cleaning strategy database refers to a standardized data set pre-stored containing the correspondence between stain features and cleaning parameters. The target cleaning parameter refers to a cleaning parameter combination preliminarily determined after matching and screening, which is suitable for the stain features. The target cleaning strategy refers to a complete scheme that can directly guide the cleaning operation after parameter verification and optimization.

[0046] As an optional implementation, the core logic and application standard of the decision-level fusion rule are comprehensively analyzed, high-order feature extraction and logical integration of the multi-modal data are performed according to the decision-level fusion rule, and core information strongly related to the cleaning decision is refined. Then, the cleaning strategy database is called, a multi-dimensional mapping relationship between the core information and the cleaning parameters in the database is established, a plurality of candidate cleaning parameters are screened out through feature-by-feature comparison and adaptability scoring, and the initial target cleaning parameter is determined in combination with the processing requirements of the stain categories and the spatial features of the position information. Then, comprehensive verification is carried out from three dimensions of the adaptability of the initial target cleaning parameter and the composition, the matching degree of the parameter coverage range and the position distribution, and the compatibility of the parameter intensity and the stain characteristics, the adaptive deviation is identified through multiple rounds of cross verification, and a correction scheme is developed to optimize the parameter. Finally, the parameters after verification and optimization are integrated to form a target cleaning strategy with rigorous logic and strong adaptability. The decision logic of this method is rigorous, the parameter matching is accurate, and the verification dimension is comprehensive.

[0047] As another optional implementation, the core application elements of the decision-level fusion rule are extracted first, and only the core features of the multi-modal data are refined, focusing on the key information directly related to the cleaning parameter matching. After calling the cleaning strategy database, the target cleaning parameter that meets the basic adaptation requirements is directly matched according to the core information. Then, basic adaptability verification is carried out to check whether the parameter and the category of the parameter can cover the main position area of the stain. If there is no obvious contradiction in the verification, the basic flow information of the cleaning operation is directly supplemented based on the parameter to form the target cleaning strategy. This method is simple and efficient in process, fast in data processing and parameter matching, and low in resource consumption.

[0048] Step S60, controlling the cleaning equipment to clean the to-be-processed area according to the target cleaning strategy.

[0049] In this embodiment, the cleaning equipment refers to a functional carrier that carries a cleaning function and can complete a stain removal operation according to instructions.

[0050] As an optional implementation, the core instructions in the target cleaning strategy are comprehensively analyzed to determine the key nodes of the cleaning path, the cleaning parameters corresponding to each node, the execution sequence, and the triggering conditions, and the target cleaning strategy is disassembled into step-by-step operation instructions that can be directly executed. According to the disassembled instruction sequence, the cleaning equipment is sequentially dispatched to start the corresponding function modules, and the cleaning operation is gradually advanced along the preset path, with accurate output of matching cleaning parameters at each key node to ensure that the parameters are accurately adapted to the characteristics of the stains in the area. The operation state and the area cleaning effect are monitored in real time during the cleaning process, and the cleaning standard is verified by comparing with the preset standard. If it is found that the local is not up to standard or the parameter adaptation deviation, the equipment operation parameters and the advance rhythm are immediately adjusted according to the correction rules preset in the strategy. After completing the cleaning of the entire path, the entire area to be processed is comprehensively checked, and the cleaning operation is terminated after confirming that there is no missed stain area. The cleaning operation of this method is highly refined, has strong parameter and path adaptability, and has complete and no-missing cleaning effect, thereby ensuring the cleaning quality and operation standardization.

[0051] Exemplarily, in the scenario of floor cleaning, multi-modal data is synchronously collected: the system triggers multiple sensors simultaneously, collects different types of data, and stamps a uniform timestamp on the data, which is the basis for subsequent fusion. Modality 1: visible light vision information: sensor: high-definition RGB camera. Information value: provides rich color, texture, and contour information. It is the main basis for identifying colored stains such as coffee and soy sauce. Modality 2: depth / thermal image information: sensor: RGB-D depth camera or infrared thermal imager. Information value: depth information: determines whether a liquid has "diffuse reflection" characteristics (smooth liquid surface, unique reflection pattern for structured light, different from dry floor), and can calculate the approximate volume / thickness of the stain. Thermal image information: if there is a temperature difference between the stain and the floor (such as freshly poured warm water or iced drinks), thermal imaging can provide a very strong identification signal. This is a complement of physical characteristics. Modality 3: floor contact / proximity information: sensor: humidity sensor, optical refractive index sensor (used by some industrial robots). Information value: provides direct, local physical evidence. For example, when a humidity sensor comes into contact with a liquid, the resistance changes, giving a definite signal that "there is definitely liquid here", but the coverage is small. Data preprocessing and spatio-temporal alignment: data preprocessing: denoising, correction (such as camera distortion correction), and standardization are performed on each modality data. Spatio-temporal alignment: this is a key prerequisite for fusion. Through sensor calibration, the data in different sensor coordinate systems is unified into the robot coordinate system. Spatial alignment: the point cloud of the depth camera, the image pixels of the thermal imager, and the image pixels of the RGB camera are one-to-one corresponding. That is, for a pixel point in the RGB image, the depth value and temperature value of the same physical point can be found in the depth map and thermal image. Temporal alignment: using the timestamp stamped during collection, it is ensured that the data frames used for fusion are collected at the same time. Multi-modal fusion strategy: fusion can be carried out at three levels, and one or a combination of the following innovative methods can be emphasized in the patent: First, data-level fusion, which splices the aligned multi-source raw data at the pixel level. Example: splice the three channels (R, G, B) of the RGB image with the depth (D) channel and the thermal image (T) channel into a 6-channel "multi-modal image" [R, G, B, D, T]. Advantages and challenges: theoretically, the information is retained most completely, but the sensor alignment requirements are extremely high, and the data volume is large, with high computational overhead. Second, feature-level fusion (most commonly used and effective), which lets different deep learning branch networks extract high-level features from different modalities of data, then fuses these features, and then performs final recognition. Branch 1 (RGB branch): a convolutional neural network extracts color and texture features F rgb from the RGB image. Branch 2 (depth branch): another CNN extracts geometric shape and height change features F depth from the depth map. Feature fusion: fuse features F rgb and F depthThe splicing or weighted addition is performed in the middle layer to form a fused feature F fused . Common decision: F fused is sent to a subsequent decoder network to generate a final segmentation mask. Advantage: high flexibility, complex network structure can be designed to mine the complementary relationship between different modalities, which is the current research hotspot. Third, decision-level fusion, each modality uses an independent model for identification to obtain its own preliminary identification result (such as an RGB model outputting a segmentation map and a depth model outputting another segmentation map), and finally merging through rules (such as voting, logical and / or). For example: when both the RGB model and the depth model consider a certain area to be a stain, it is finally determined to be a stain. This can greatly reduce the false positives of a single modality (such as misjudging ground shadows as stains). Advantage: good system fault tolerance, the system can still work when a sensor fails. Implementation is relatively simple. Final decision based on fusion information: the accurate segmentation map output by the fusion model is combined with other information to form the final decision: stain positioning: combining the fusion recognition result and the robot's own positioning to determine the precise coordinates of the stain on the map. Stain attribute judgment (optional, but innovative): using fusion information, the stain type can be further judged. For example: the RGB feature (color) suggests "dark liquid", and the thermal imaging feature suggests "temperature higher than the environment", which can be inferred as "hot coffee". This can provide a basis for subsequent cleaning strategies (such as whether to use suction first or mop). Cleaning strategy generation: according to the size (calculated from the image), type (if it can be judged), and thickness (inferred from the depth information) of the stain, the decision engine can generate more targeted cleaning instructions (such as: large-area water area only needs to be sucked dry, and small and viscous syrup needs to be sprayed and brushed).

[0052] Due to the accurate integration and hierarchical processing of multi-modal data, the problems of poor data collaboration and blind cleaning parameter matching in component construction are solved, and the data interaction efficiency and cleaning parameter adaptation accuracy between components are improved.

[0053] Based on any of the above embodiments, in the second embodiment of the present application, the step S20 includes steps A11-A13: Step A11, according to the data type, divide the original data, and filter the corresponding rules in the information processing rule library to determine the information processing rule corresponding to each data type.

[0054] In this embodiment, the information processing rule library refers to a collection of pre-constructed processing specifications corresponding to various data types.

[0055] As an optional implementation, the information carrier form, feature form and data structure of the original data are analyzed, and the original data is compared and classified one by one according to the preset refined classification system, and each data is labeled with a special class identifier, and the consistency of the classification result is checked to avoid cross-class confusion. Then, the complete information processing rule library is called, and the core feature requirements of the corresponding data type are associated according to the class identifier, the noise type and effective information extraction key of each class of data are analyzed, and the special information processing rule containing processing steps, parameter standards and quality check requirements is selected from the rule library. Through multi-dimensional adaptive verification of classification and rules, it is confirmed that the rule can accurately match the processing requirements of the data type, correct the matching deviation, and finally form a precise mapping relationship between the classification result and the corresponding information processing rule. This method is classified finely and the rule matching is highly targeted, which can adapt to the individual processing needs of various data.

[0056] Step A12, denoising the visual data in the original data by edge-preserving algorithm, baseline correction of the spectral data in the original data, and filtering processing of the acoustic data in the original data to obtain processed single-modal data.

[0057] In this embodiment, edge-preserving algorithm denoising refers to a processing method that removes noise interference in visual data while preserving image edge details and feature information. Visual data refers to modal data in the original data that is presented in the form of an image, contains spatial pixels and visual features. Baseline correction refers to a processing operation that corrects baseline drift and background interference in spectral data to make the spectral signal reference stable. Spectral data refers to modal data in the original data that is presented in the form of a spectral curve and contains substance band response characteristics. Filtering processing refers to a processing method that separates effective signals and noise signals in acoustic data to purify target signals. Acoustic data refers to modal data in the original data that is presented in the form of a sound wave signal and contains reflection or transmission acoustic characteristics. Single-modal data refers to structured data that contains only single-modal effective information after targeted processing.

[0058] As an optional implementation, edge feature recognition is first performed on visual data in the original data to mark key edge regions such as contours and boundaries in the image. Then, an edge-preserving algorithm is applied to perform hierarchical noise reduction processing on non-edge regions, and the processing intensity is adjusted according to the noise intensity, while the integrity of the edge region is monitored in real time to avoid edge blurring. For spectral data, the baseline drift trend and background interference source are first analyzed, and baseline fitting and correction are performed according to the drift amplitude. After each segment is corrected, the smoothness of the spectral signal and the integrity of the characteristic peak are verified, and the correction deviation is corrected. For acoustic data, the signal frequency distribution characteristics are first identified to distinguish the frequency interval of effective signals and noise, and different types of noise are separated by targeted filtering methods to gradually purify the effective signals. During this period, the filtering effect is verified through signal intensity and frequency stability. Finally, the processing results of the three types of data are integrated to form single-mode data that retains core characteristics and has minimal noise interference. This method has strong processing pertinence and can accurately extract and integrate core effective information in each modality, ensuring that the data has high purity and complete detail retention.

[0059] Step A13, the processed single-mode data is associated and integrated according to the preset spatiotemporal alignment rule, and multi-modal data is output.

[0060] In this embodiment, the preset spatiotemporal alignment rule refers to a standardized matching specification set in advance for realizing the time dimension synchronization and spatial dimension correspondence of different modal data.

[0061] As an optional implementation, the preset spatiotemporal alignment rule is analyzed to determine the reference standard for time synchronization and the spatial coordinate mapping relationship. Then, the time dimension of each single-mode data is calibrated one by one, the time stamp format of all data is unified according to the rule, the collection period deviation is corrected, and the different modal data is completely synchronized on the time axis. Then, according to the spatial coordinate mapping relationship, the single-mode data after time synchronization is matched point by point in space to establish a one-to-one correspondence between the pixel points of visual data, the detection points of spectral data, and the signal points of acoustic data. During the matching process, the consistency of spatiotemporal alignment is checked in real time, and the coordinate offset and time lag are corrected through cross-validation. Finally, the single-mode data after spatiotemporal precise alignment is deeply associated and integrated according to the preset structure, the core effective information of all modalities is retained, and the multi-modal data with regular structure and strong spatiotemporal consistency is output. This method has high spatiotemporal alignment accuracy and tight data association, and the output multi-modal data has strong consistency and integrity, effectively avoiding analysis errors caused by spatiotemporal misalignment.

[0062] Exemplarily, in the ground cleaning scene, the RGB camera 1920x1080 pixel image data, 16-band infrared sensor spectral image data, and 40kHz ultrasonic distance data are synchronously collected, and a 2024-06-10 09:15:30.500 timestamp is uniformly stamped on the three types of modal original data; the corresponding data purification processing of median filtering denoising on the RGB image data, mean filtering denoising on the infrared data, and wavelet threshold denoising on the ultrasonic data is performed to obtain the denoised modal data. According to the RGB image perspective transformation calibration rule, the infrared data radiation calibration rule, and the ultrasonic data distance compensation calibration rule, the denoised original data is calibrated respectively, and the calibrated original data is output. The corrected original data is subjected to data scale normalization processing to the range of [0, 1] through the Min-Max scaling algorithm, and is subjected to format standardization processing in the JSON format, and the field naming and data storage structure are unified, and finally the multi-modal data containing image, spectral, and distance information is obtained, and the entire process is completed by OpenCV for image processing and Scikit-learn for normalization operation.

[0063] Due to the multi-modal data synchronous calibration and standardization integration, the problems of multi-source data asynchronization, noise interference, and format heterogeneity in ground cleaning are solved, the data consistency and reliability are improved, and high-quality data support is provided for subsequent stain identification.

[0064] Based on any of the above embodiments, in the third embodiment of the present application, the step S30 includes steps B11-B13: Step B11, the aligned multi-modal data is subjected to resolution normalization processing, and the pixel coordinates of each modal data are accurately mapped based on the inter-modal pixel adaptation strategy to generate multi-modal pixel-level data.

[0065] In this embodiment, the resolution normalization processing refers to the operation of adjusting the aligned multi-modal data of different resolutions to a uniform resolution. The accurate mapping of the pixel coordinates of each modal data refers to the operation of one-to-one correspondence of the pixels of different modal data in spatial position. The multi-modal pixel-level data refers to the composite data of accurate matching of each modal pixel coordinate and uniform resolution.

[0066] As an optional implementation, after alignment, the multi-modal data is split by modality type, and for the original resolution characteristics of each modality data, a targeted resolution adjustment operation is performed to ensure that the data does not lose core features during the resolution conversion process through step-by-step scaling and detail compensation, and all modality data is unified to a preset resolution standard. Subsequently, a pixel-by-pixel coordinate mapping process is started, and a one-to-one correspondence between the pixel coordinates of each modality data is established according to the inter-modality pixel adaptation strategy. Through multiple rounds of verification such as adjacent pixel correlation checking and edge feature alignment verification during the mapping process, mapping deviations are corrected to ensure that the same spatial position of different modalities corresponds to the same pixel coordinates, and finally multi-modal pixel-level data is integrated. This method has complete resolution adjustment and detail retention, high pixel coordinate mapping accuracy, and strong data spatial consistency.

[0067] Step B12, according to the preset channel combination rule, the multi-modal pixel-level data is channel spliced and fused to obtain initial fusion data.

[0068] In this embodiment, the preset channel combination rule refers to a rule that specifies the arrangement order and fusion method of the channels corresponding to each modality data. Channel splicing and fusion refer to the operation of orderly integrating multi-modal pixel-level data by channel dimension. Initial fusion data refers to a comprehensive data body containing all modality channel information after channel splicing.

[0069] As an optional implementation, first analyze the channel attributes, information types and data priorities of each modality in the multi-modal pixel-level data, and determine the channel arrangement position and splicing order corresponding to each modality by comparing with the preset channel combination rule. Then perform the splicing operation according to the order, one modality at a time, and one channel at a time. After completing the channel splicing of each modality, immediately check the compatibility of the modality channel and the spliced channel, including data format consistency and information overlap verification, to avoid channel conflict or information redundancy. After all modality channels are spliced, perform overall channel integrity verification to confirm that there is no channel omission or out-of-order situation, and real-time correction is performed on the existing minor deviations, and finally the initial fusion data is obtained. This method has regular channel arrangement, accurate modality information matching, no redundancy or conflict, high data integrity, and effectively improves the efficiency and accuracy of subsequent processing.

[0070] Step B13, performing format regularity processing and verification on the initial fusion data, and outputting the fusion data.

[0071] In this embodiment, format regularity processing refers to the operation of unifying the storage structure, encoding specification and field definition of the initial fusion data. Fusion data refers to the final comprehensive data body with unified format and complete information after format regularity processing and verification.

[0072] As an optional implementation, the existing storage structure, encoding type and field distribution of the initial fusion data are first analyzed, and the storage mode, encoding format of the data is adjusted field by field, the field naming and arrangement order are unified, and it is ensured that the format of each data unit conforms to the specification. Then, a multi-level verification process is started, the correctness of the format of a single field is first checked, then the integrity of the overall structure of the data is checked, and finally the correlation and consistency of the modal information after the regularization are verified. The format deviation and field missing problems found in the verification are corrected one by one. After correction, the data is rechecked to determine that there is no missing problem, and finally the fusion data is output. This method is format regular and detailed, and the verification is comprehensive, which can maximize the standardization and integrity of the data and improve the smoothness of subsequent processing.

[0073] Exemplarily, in the ground cleaning scene, the aligned multi-modal data includes 1920x1080 pixel RGB image data (3 channels), 640x480 pixel 16-band infrared spectrum image data (16 channels), and 320x240 pixel ultrasonic depth map data (1 channel). The resolution normalization processing is first performed on the three types of data, all data is uniformly upsampled to the highest resolution of 1920x1080 pixels based on the bilinear interpolation algorithm, and the pixel coordinates of each modal data are accurately mapped according to the inter-modal pixel adaptation strategy, thereby establishing a one-to-one correspondence relationship between the spatial positions, and generating multi-modal pixel-level data. According to the preset channel combination rule of "RGB channel, infrared channel, ultrasonic channel" order, the multi-modal pixel-level data is sequentially spliced and fused according to the channel dimension, and the initial fusion data of 20 channels is obtained. The HDF5 format is used to perform format regularization processing on the initial fusion data, and the field naming and data storage structure are unified. The data integrity, channel order correctness and format standardization are checked by using the OpenCV tool, and the identified channel dislocation abnormality is corrected, and finally the fusion data is output. The whole process is completed by using the Scikit-image library to complete the resolution conversion.

[0074] Due to the multi-modal data resolution unification, pixel accurate mapping and format regularization verification, the spatial dislocation and format heterogeneity problems of multi-modal data in ground cleaning are solved, and the consistency and accuracy of data fusion are improved.

[0075] Based on any of the above embodiments, in the fourth embodiment of the present application, the step S30 includes steps C11-C13: Step C11, extracting features reflecting stain characteristic information from pixel information corresponding to different categories of stains to form a multi-modal material characteristic feature set.

[0076] In this embodiment, the corresponding pixel information refers to the associated information in the data associated with the pixel position and the attribute of the pixel. The features reflecting the material characteristic information refer to the core data features that can reflect the essential attributes of the material. The multi-modal material characteristic feature set refers to a unified feature set formed by integrating the material characteristic features extracted in different modalities.

[0077] As an optional implementation, the pixel region labeled as a stain is selected from the stain category, and all modal pixel information corresponding to the region is located. After splitting the modal pixel information according to the modal type, the features in each modality that can reflect the essential attributes of the material composition, state, etc. are extracted pixel by pixel, covering multi-dimensional features such as texture, spectral response, and spatial distribution. During the extraction process, each feature is annotated for effectiveness, and through cross-modal feature cross-validation, contradictory or invalid features are removed to ensure the authenticity of individual features. Finally, all valid features are sorted according to the modal classification, and the association mapping of the features, modalities, and pixel positions is established to integrate and form a structured multi-modal material characteristic feature set. This method is fine-grained in feature extraction, covers a comprehensive dimension, and has strong feature authenticity and relevance, providing comprehensive feature support for subsequent stain analysis.

[0078] Step C12, by analyzing the correlation between different material characteristics in the multi-modal material characteristic feature set, the complementary correlation between the features is determined.

[0079] In this embodiment, feature correlation analysis refers to the analysis process of mining the correlation rules of complementary features by analyzing the internal logic of different material characteristics. Different material characteristics refer to independent features belonging to different modalities or different attribute dimensions in the set. Complementary correlation refers to the correlation form in which different material characteristics complement each other's shortcomings in representing material attributes and strengthen the overall recognition. Complementary correlation refers to the stable logical correspondence or mutual support relationship between different material characteristics after analysis.

[0080] As an optional implementation, the composition dimensions of the multi-modal material characteristic feature set are first analyzed, and the material characteristics under each modality are extracted one by one to determine the representation direction and coverage of each feature. Subsequently, two-by-two complementary analysis is carried out for each two features to mine the complementary points between the features in representing material attributes, record the complementary strength and logical correlation form. Through multiple rounds of cross-validation, the analysis results are strengthened, false correlations are excluded, and the interlocking complementary relationships between multiple features are integrated to form a multi-dimensional feature correlation network. Finally, the core logic of the correlation network is sorted out to determine the priority and stability of the feature correlation in different scenarios, and a structured and logical feature correlation relationship is determined. This method has comprehensive feature analysis and in-depth correlation mining, and can capture subtle complementary correlations, with strong accuracy and completeness of the correlation relationship.

[0081] Step C13, according to the preset mapping rule of the material properties and categories, the complementary association relationship is inferred, and the category of the stain is determined.

[0082] In the embodiment, the preset mapping rule of the material properties and categories refers to a standardized corresponding standard between the features representing the inherent properties of the material and the categories of the stains.

[0083] As an optional implementation, the preset mapping rule of the material properties and categories is analyzed, the corresponding core material property feature combination, feature threshold range and complementary association requirement of each category are determined, and a multi-dimensional corresponding index of the rule and the features is established. Then, the core features and the complementary logic in the feature association relationship are disassembled, the key feature combination strongly related to the mapping rule is extracted, and the combination is compared with the feature requirements of each category in the mapping rule one by one. Through multiple rounds of cross-validation to strengthen the inference process, the consistency of the feature combination and the category corresponding relationship is checked, the feature conflict or rule matching deviation is excluded, the fuzzy matching condition is further refined and deduced combined with the feature complementarity strength, and the inference result is obtained. Finally, according to the adaptation degree of the inference result and the mapping rule, the category of the stain is comprehensively determined. The method has rigorous inference logic, comprehensive rule matching, and the accuracy and reliability of the category determination are extremely high, which provides a high credibility category basis for subsequent cleaning parameter matching.

[0084] Exemplarily, in a ground cleaning scene, the determination of the category of the stain is processed based on multi-modal fusion data and pixel-level segmentation information. The input data includes: 1280x720 pixel multi-modal fusion data (containing RGB, infrared and ultrasonic channels), and pixel attribution information identifying the stain area. First, high-level features representing the characteristics of the material are extracted from each modality data from the stain pixel area (for example, a rectangular area with a pixel coordinate range of x axis 350 to 680 and y axis 220 to 510) of the data: texture and color features of the RGB modality are extracted through a pre-trained ResNet model; absorption peak and waveband response features of the infrared modality are extracted through a spectral feature extractor; reflection intensity and signal attenuation features of the ultrasonic modality are extracted through a statistical method. These features together constitute a multi-modal material characteristic feature set. Subsequently, Pearson correlation analysis is used to mine the complementary and correlation relationship between different modal features, such as the correlation between texture roughness and ultrasonic reflection intensity. According to the rules of mapping the preset "spectral peak interval", "texture type", "reflection intensity threshold" and other features to the categories such as "oil stain", "water stain" and "sauce stain", the preliminary category determination result and its confidence (for example, determined as "oil stain" with a confidence of 88%) are generated through random forest classification reasoning. Finally, the preliminary determination result is accurately associated and integrated with the original fusion data and the stain pixel area to generate a structured determination record for each stain area, which contains its pixel coordinate range, the extracted multi-modal feature vector, the final determined category label and the classification confidence. In this embodiment, feature processing and correlation analysis can be realized through the Scikit-learn library, and feature extraction and classification reasoning based on deep learning can be completed with the help of the TensorFlow framework. Due to multi-modal feature extraction, correlation reasoning and data integration, the problem of insufficient basis for stain category determination and data fragmentation in ground cleaning is solved, and the accuracy of stain classification and the integrity of data are improved.

[0085] Based on any of the above embodiments, in the fifth embodiment of the present application, the step S40 includes steps D11-D13: Step D11, according to the modality of the multi-modal data, the high-level features representing the stain outline, the stain morphology and the material characteristics are extracted respectively, and the independent high-level feature set of each modality is obtained.

[0086] In this embodiment, the high-level features representing the stain profile refer to deep data features that can reflect the boundary range and profile shape of the stain. The high-level features representing the stain shape refer to deep data features that can reflect the external structure and shape features of the substance. The high-level features representing the substance characteristics refer to deep data features that can reflect the essential attributes of the substance. The high-level feature set of each modality refers to a dedicated feature set formed by separately integrating the three types of features representing the corresponding modality according to different modalities.

[0087] As an optional implementation, the multi-modal data is first completely split according to the modality type, and the core data dimension and feature extraction priority of each modality are determined. For each modality, the stain profile feature extraction process is started in turn, deep profile features are mined through edge feature enhancement and profile fitting, shape features that can reflect the external structure of the substance are extracted, including distribution rules and geometric features, and finally the characteristic features that reflect the essential attributes of the substance are deeply mined. During the extraction process, each feature is effectively screened to eliminate redundant and interfering features, and the feature authenticity is ensured through intra-modality feature consistency verification. Finally, the three types of features are integrated according to the modality to form a structured and complete high-level feature set of each modality. This method has comprehensive feature extraction dimensions and strong pertinence, high feature purity and effectiveness, and clear independence of features of each modality.

[0088] Step D12, according to the contribution of each modality to stain identification, an adaptive weight is assigned, the high-level feature set corresponding to each modality is weighted and fused, and a fused feature is output.

[0089] In this embodiment, the contribution of each modality to stain identification refers to the degree of the role played by different modalities in supporting the stain identification process. The adaptive weight refers to a distribution value that matches the contribution and is used to adjust the influence of the features of each modality. The fused feature refers to a unified feature that combines the advantages of each modality after weighted integration.

[0090] As an optional implementation, the high-level feature set of each modality is subjected to effectiveness analysis, the correlation between the features and the stain identification target, and the dimension of the feature recognition degree are combined to quantitatively evaluate the contribution of each modality to stain identification, and an evaluation result is obtained. According to the evaluation result, an adaptive weight is assigned to each modality, and the higher the contribution, the greater the weight proportion. At the same time, through inter-modality weight balance verification, feature bias caused by excessive weight of a single modality is avoided. Subsequently, corresponding features in the high-level feature set of each modality are subjected to fusion operation one by one according to the weight proportion, and the rationality of feature fusion is checked in real time during the operation process to correct the weight distribution deviation. Finally, all operation results are integrated to output a fused feature that combines the advantages of each modality. This method has accurate weight distribution and rigorous fusion logic, can fully play the core role of high-contribution modalities, has high-quality fused features, and effectively improves the segmentation accuracy.

[0091] As another optional implementation, real-time effectiveness monitoring is performed on the high-level feature set of each modality, dynamic change indicators such as feature recognition degree and noise interference degree are analyzed, and the real-time contribution degree of each modality is dynamically quantitatively evaluated in combination with the real-time correlation closeness of the features and the stain identification target, to ensure that the contribution degree evaluation can accurately reflect the real-time performance of the features. According to the dynamic contribution degree, the adaptive weight is dynamically allocated, and the weight is adjusted up when the contribution degree rises, and the weight is adjusted down when the contribution degree falls, and at the same time, through the weight balancing mechanism among the modalities, the feature imbalance caused by the excessive fluctuation of the weight of a single modality is avoided. Then, the corresponding features in the high-level feature set of each modality are aligned one by one according to the dynamic weight ratio and fusion operation is performed, the weight change and feature fusion effect are tracked in real time during the operation process, the fusion deviation is dynamically corrected, and finally all dynamic operation results are integrated to output the fusion features. The weight allocation of this method can accurately adapt to the real-time state of the features, the dynamic adaptability of the fusion features is strong, the recognition degree is high, and a high-dynamic-adaptability feature basis is provided for subsequent segmentation mask generation.

[0092] In step D13, the semantic segmentation decoder is used to perform pixel-level classification on the fusion features, to label the stain attribution attribute of each pixel, to obtain the segmentation mask with the same size as the stain in the original image.

[0093] In this embodiment, the semantic segmentation decoder refers to a feature processing module for performing pixel-level classification on the fusion features and assigning pixel attribution attributes. The pixel-level classification processing refers to a fine processing operation of classifying each pixel one by one. The stain attribution attribute refers to an attribute identifier indicating whether a pixel belongs to a stain. The size of the stain in the original image refers to the actual size and range of the stain itself in the original image.

[0094] As an optional implementation, the fusion features are first subjected to feature enhancement processing to strengthen the feature difference between the stain and the background, and to provide a clear basis for pixel-level classification. Then, the enhanced fusion features are processed pixel by pixel by the semantic segmentation decoder, combined with feature similarity comparison and neighborhood pixel correlation analysis, to accurately determine the stain attribution attribute of each pixel and label whether it belongs to a stain. During the classification process, the size parameters of the stain in the original image are referred to, and through coordinate mapping and scale calibration, it is ensured that the size of the pixel set after classification and labeling is completely consistent with the original stain. After the classification is completed, the edge pixel correction and integrity verification are performed on the labeling result to fill the classification gaps and correct the misjudgment pixels, and finally the segmentation mask is formed. This method is accurate in pixel classification, accurate in attribution attribute labeling, and high in size matching degree, and can completely restore the real shape of the stain.

[0095] Exemplarily, in the ground cleaning scene, the input multi-modal data includes a 1920x1080 pixel RGB image, a 16-band infrared spectrum image, and 40 kHz ultrasonic wave depth data. First, high-level features are extracted by modality: edge gradient features representing stain outline, geometric structure features representing stain morphology in the RGB modality are extracted by a pre-trained ResNet18; band absorption peak features representing material characteristics in the infrared modality are extracted by a spectrum feature extractor; spatial distribution features representing stain morphology and reflection attenuation features representing material characteristics in the ultrasonic wave modality are extracted by a PointNet, thereby obtaining a high-level feature set independent of each modality. According to the contribution of each modality to stain identification (which can be determined by historical data analysis or attention mechanism), a fusion weight is assigned to each feature set (for example, RGB modality weight 0.4, infrared modality weight 0.35, ultrasonic wave modality weight 0.25), and the high-level feature sets corresponding to each modality are fused by weighted summation operation to generate unified fusion features. Next, a U-Net semantic segmentation decoder is used for pixel-level classification processing of the fusion features, and the decoder output is processed by a Softmax function to label the stain attribution attribute of each pixel (for example, stain pixels are set to a threshold such as confidence ≥ 0.85, labeled as "stain", and the rest is "background"), to determine the pixel attribution, thereby generating a preliminary segmentation result. Finally, through the above sampling and coordinate mapping operations, the size and coordinate system of the preliminary segmentation result are calibrated to the standard of the original input image (i.e. 1920x1080 pixels), and finally a segmentation mask is generated which is completely consistent with the size and position of the stain area in the original image. The mask can accurately define the stain area, for example, its boundary coordinate range is x axis 420 to 780, y axis 230 to 560.

[0096] Due to the precise fusion of multi-modal high-level features and pixel-level semantic segmentation, the problem of unclear stain outline and inaccurate positioning in ground cleaning is solved, and the accuracy of stain segmentation and the reliability of the mask are improved.

[0097] Based on any of the above embodiments, in Embodiment Six of the present application, the step S40 includes steps E11-E13: Step E11, analyze the segmentation mask to extract the complete outline, boundary range, spatial coordinates, and area parameters of the stain, and integrate to obtain stain location feature data.

[0098] In this embodiment, the complete outline of the stain refers to the morphological feature of the continuous closed edge of the stain. The boundary range refers to the limit interval of the periphery of the stain. The spatial coordinates refer to the position identification information of the stain in the data coordinate system. The area parameter refers to the quantitative data representing the size of the stain coverage. The stain location feature data refers to the comprehensive location information set formed by integrating the complete outline, boundary range, spatial coordinates, and area parameters of the stain.

[0099] As an optional implementation, the pixel attribution in the segmentation mask is first parsed pixel by pixel, and the continuous closed form of the stain is constructed by edge pixel correlation tracking to form a complete contour. Based on the distribution rule of the contour edge pixels, the peripheral limit interval of the stain is defined, and the boundary range is determined. The position identification information of the contour key inflection point and vertex is located, and the spatial coordinates in the unified coordinate system are formed through coordinate correlation calibration. The total amount of pixels belonging to the stain in the contour is counted, and the quantitative data representing the size of the coverage range are obtained by combining data scale conversion. In the analysis process, the abnormal data points are removed through multiple rounds of correction operations such as contour continuity verification, boundary range consistency verification, and coordinate accuracy calibration. Finally, the complete contour, boundary range, spatial coordinates, and area parameters are integrated according to the preset structure to obtain the stain position feature data. This method has high analysis accuracy, and the extracted parameters are complete and accurate, and the data consistency is good.

[0100] As another optional implementation, the edge pixels of the stain in the segmentation mask are detected point by point, and the discontinuous points, jagged protrusions and recess defects of the edge are identified to establish an edge defect position mapping table. Subsequently, based on the distribution rule and curvature characteristics of the neighborhood edge pixels, the jagged defects are passivated by using an edge smoothing algorithm, and the discontinuous gaps are filled by using a missing pixel interpolation algorithm to ensure that the edge forms a continuous closed form. After optimization, the boundary range of the stain is defined based on the optimized complete contour, the spatial coordinates of the contour key inflection point and vertex are located, the total amount of stain pixels in the contour is counted, and the area parameter is obtained by combining scale conversion. Through multiple rounds of verification such as edge continuity verification, contour and original image stain size comparison, etc., the optimization deviation is corrected, and finally all the optimized parameters are integrated to obtain the stain position feature data. This method optimizes the edge accurately, and the extracted contour, boundary and other parameters have high detail restoration degree and strong data accuracy.

[0101] Step E12, the stain position feature data and the stain are associated and fused into categories, and the corresponding attribute information of the stain is supplemented to obtain the stain comprehensive information data.

[0102] In this embodiment, the attribute information corresponding to the stain refers to the supplementary information representing the inherent additional characteristics of the stain. Association and fusion refer to the operation of establishing a corresponding relationship between the two types of data and integrating them into a unified body. The stain comprehensive information data refers to a comprehensive stain information set formed by integrating the position feature, multi-modal determination result and supplementary attribute information.

[0103] As an optional implementation, the core dimension of the stain position feature data is first analyzed, and the category label, confidence and other key information of the stain category are determined. A dimension-by-dimension correlation mapping is established to ensure that the position parameters and the determination result are accurately corresponding to the same stain unit. Then, based on the material characteristic information in the stain category, additional attribute information such as the state and viscosity of the stain is supplemented, and each attribute information is associated with the position feature and the determination result. During the integration process, multi-level verification is performed, including verifying the consistency of the two types of data, checking the matching rationality of the supplemented attributes and the core information, eliminating contradictory data and correcting deviations, and finally sorting all information according to the preset unified structure to form comprehensive stain information data with complete fields and close correlations. This method has accurate correlation mapping, comprehensive attribute supplementation, strong data consistency and completeness, and provides high-precision and comprehensive data support for subsequent cleaning strategy formulation.

[0104] Step E13, predicting the position change trend of the stain based on the comprehensive stain information data, and generating position information for describing the stain distribution.

[0105] In this embodiment, the position change trend refers to predicting the change law of the stain in spatial dimensions such as diffusion, movement and contraction based on the stain's own characteristics and environmental influences. The position information of the stain refers to structured data including the predicted spatial coordinates, boundary range, contour shape and change area of the stain.

[0106] As an optional implementation, the core influencing factors in the comprehensive stain information data are analyzed, including the diffusion characteristics of the stain category, the viscosity and adhesion in the material characteristics, and the spatial environmental correlation information of the initial position. The action weight and logical relationship of each factor on the position change are determined. Then, the historical position related data of the stain is sorted out, and a change law analysis model is established combining the current state data to deduce the possible position offset amplitude, diffusion direction and boundary change range of the stain at different time nodes. Through multiple rounds of cross-validation, the prediction bias is corrected, and finally the complete stain position information including static position and dynamic trend is generated by integrating the predicted spatial coordinates, dynamic boundary range, contour change characteristics and key change node information. This method has comprehensive prediction dimensions, rigorous logic, strong accuracy and completeness of position information, and can accurately capture subtle position changes to ensure the accuracy of cleaning operations.

[0107] Exemplarily, in the ground cleaning scene, the segmentation mask is 1920x1080 pixels (the stain pixels are marked as "1" and the background is "0"), the complete contour (continuous closed edge curve), boundary range (x-axis: 310-760 pixels, y-axis: 220-580 pixels), spatial coordinates (upper left corner (310, 220), lower right corner (760, 580)) and area parameter (212500 pixels) of the stain are extracted by OpenCV analysis, and the stain location feature data is integrated. The data is associated and fused with the stain category information according to the corresponding pixel area (for example, containing the "oil stain" category label, 90% characteristic confidence and multi-modal feature), and other attribute information of the stain (for example, medium viscosity, 35% water content) is supplemented, so as to integrate the preliminary stain comprehensive information data. Finally, the data set is standardized: using the Min-Max scaler of Scikit-learn, the numerical features such as area, confidence and water content are normalized to the interval [0, 1], the field format is regularized according to the preset "location parameter", "category information", "attribute data" and "confidence", and redundant information is removed, to generate standardized stain comprehensive information (containing the stain category, the location information of the stain) that can be directly called by the subsequent decision-making process.

[0108] Due to the position feature extraction, multi-dimensional data fusion and quantization and regularization, the problem of stain information fragmentation and incomplete attributes in ground cleaning is solved, and the data integration efficiency and the availability of core data are improved.

[0109] Based on any one of the above embodiments, in the seventh embodiment of the present application, the step S50 described in the flowchart of the seventh embodiment of the autonomous cleaning method for ground stains of the present application comprises steps F11-F14: Figure 2 , Figure 2 The step F11 comprises the following steps: Step F11, analyzing the location parameter, category, area data and characteristic confidence of the stain in the multi-modal data to form the stain attribute data.

[0110] In this embodiment, the location parameter refers to quantitative data representing the spatial position of the stain. The category refers to the material composition attribution identification of the stain. The area data refers to quantitative information representing the size of the stain coverage. The characteristic confidence refers to a quantitative index representing the reliability of the category determination. The multi-dimensional cross verification refers to a process of mutual verification and logical checking of the location parameter, category, area data and characteristic confidence. The stain attribute data refers to a standardized data set formed after multi-dimensional cross verification, which fully reflects the core attributes of the stain.

[0111] As an optional implementation, the four-dimensional information of the location parameter, the category, the area data and the characteristic confidence in the position information of the stain category is first analyzed, the logical association and the verification standard of each dimension data are determined, and then the multi-dimensional cross verification process is started according to the preset decision level fusion rule, the matching rationality of the location parameter and the area data, the corresponding reliability of the category and the characteristic confidence are verified one by one, and the contradiction between different dimension data is checked. In the verification process, the logical deviation and the data anomaly found are corrected, the consistency and accuracy of each dimension data are ensured through multiple iterations, and finally all the dimension information after verification is integrated in a preset format to generate stain attribute data with complete structure and strict logic. The method has comprehensive verification dimension, in-depth logical checking, strong data accuracy and reliability.

[0112] Step F12, according to the preset validity determination threshold, the abnormal data and invalid information in the stain attribute data are removed, and the key attribute information of the stain is screened.

[0113] In this embodiment, the preset validity determination threshold refers to a standard limit preset for determining whether the data meets the requirements. The abnormal data refers to invalid data that is beyond the reasonable range and does not conform to the logic. The invalid information refers to redundant information that has no actual decision value and no reference significance. The key attribute information of the stain refers to the stain attribute data that is retained after screening and has a core supporting role for decision-making.

[0114] As an optional implementation, all the dimension information of the stain attribute data is analyzed, the preset validity determination threshold corresponding to each attribute dimension is determined, and the accurate correspondence between the dimension and the threshold is established. Then the data is checked dimension by dimension, the abnormal data beyond the threshold range is marked and removed, and the invalid information with no actual decision value is identified and cleaned. In the screening process, the logical consistency of the remaining data is verified to ensure that the removal operation does not affect the integrity of the core information. Finally, according to the attribute importance, all the core data that meets the threshold requirements and is logically self-consistent are integrated to form the key attribute information of the stain. The method has comprehensive screening dimension, accurate determination, complete removal and no damage to valid information, and high data reliability.

[0115] Step F13, calling the pre-stored cleaning strategy database, matching the key attribute information of the stain with the database, and obtaining a candidate cleaning parameter set.

[0116] In this embodiment, the pre-stored cleaning strategy database refers to a standardized data set containing the correspondence between the cleaning parameter combination and the stain attribute, which is stored in advance. The cleaning parameter combination refers to a set of various associated parameters required for cleaning operation. The candidate cleaning parameter set refers to a set of multiple potential applicable cleaning parameter combinations obtained after matching.

[0117] As an optional implementation, all core dimensions of the stain key attribute information are first parsed, the attribute characteristics and logical association of each dimension are determined, and then the pre-stored cleaning strategy database is called to establish a precise mapping relationship between the attribute dimension hierarchy and the cleaning parameter combination. Through dimension-by-dimension feature comparison and cross-dimension adaptability verification, the cleaning parameter combination that fits all attribute dimensions is selected, and the logical consistency within the parameter combination is checked to eliminate combinations with conflicts or insufficient adaptability. Finally, according to the adaptability priority of the parameter combination, all cleaning parameter combinations that meet the requirements are integrated to form a candidate cleaning parameter set. This method has comprehensive matching dimensions, precise mapping, strong adaptability and reliability of the candidate set.

[0118] Step F14, in combination with the current ground material information and the current state of the equipment, the adaptability of the candidate parameters is evaluated and prioritized to obtain the adapted parameters, and the adapted parameters are fused according to the decision-level fusion rule to generate the target cleaning parameters.

[0119] In this embodiment, the current ground material information refers to the inherent attribute related information of the current cleaning area ground. The current state of the equipment refers to the running related state information of the equipment at the moment of performing cleaning operation. The candidate parameters refer to various cleaning parameter combinations in the candidate cleaning parameter set. Adaptability evaluation refers to the process of checking the matching degree of the candidate parameters with the ground material and the equipment state. Priority sorting refers to the operation of arranging the candidate parameters in order according to the adaptability and other core indicators.

[0120] As an optional implementation, the core attribute dimensions of the current ground material information and the running characteristics of the current equipment state are integrated to establish an adaptability evaluation system containing multiple dimensions such as material compatibility, equipment load adaptation degree, and cleaning effect expectation. Then, the candidate parameters are checked for adaptability one by one, the compatibility of each parameter combination with the ground material is analyzed, it is judged whether it exceeds the upper limit of the current running load of the equipment, and the cleaning effect achievement rate is estimated. Based on the evaluation results, each candidate parameter is given a quantitative score, the comprehensive score is calculated by combining the preset weight of each evaluation dimension, and the priority is sorted from high to low according to the comprehensive score. After sorting, the actual adaptability feasibility of the candidate parameters is confirmed through cross verification, the sorting deviation is corrected, and the optimal cleaning parameter is finally determined. This method has comprehensive evaluation dimensions, rigorous sorting logic, strong adaptability and reliability of the optimal parameter, and improves the scientificity and safety of the cleaning operation.

[0121] Exemplarily, in the ground cleaning scene, the algorithm flow includes stain detection and identification: image collection: when the robot travels along the path, the vision sensing module collects images of the ground in front at a certain frequency (such as 5 frames per second). Real-time inference: the collected images are sent to a lightweight convolutional neural network (CNN) model deployed on the robot side for real-time inference. Model input: ground RGB image (multiple frames can be spliced to provide a larger field of view). Model output: segmentation map: pixel-level classification, labeling which pixels in the image belong to “liquid stain”. (For example: using a lightweight DeepLabV3+ or UNet architecture). Classification label: identify the type of stain (for example: water, coffee, wine, oil stain, sauce). Severity assessment: estimate the area and approximate volume of the stain (calculated by fusing segmentation area and depth information). Cleaning strategy decision: the main control unit matches the optimal strategy from the pre-defined “cleaning strategy database” according to the algorithm’s identification results. The database pre-stores cleaning parameters corresponding to different stain types (for example), as shown in Table 1.

[0122] Table 1: Example of cleaning strategy database table

[0123] Autonomous execution and verification: accurate positioning and movement: the robot accurately moves to the center point of the stain through vision SLAM technology combined with the position of the stain in the image. Execute cleaning strategy: first, spray pre-treated water or cleaning liquid and wait for a certain time (such as 10 seconds) for soaking and softening. Then, control the mop module to descend and perform scrubbing according to the strategy set mode (vibration frequency, rotation speed, movement trajectory). Repeat the “spray-scrub” process for a set number of times. Cleaning effect verification: after completing the scheduled process, the robot retreats and collects images of the original area again. Use the same lightweight model for secondary inference. If the identification result shows that the stain has disappeared or the area has decreased below the threshold, it is determined that the cleaning is successful, and the global cleaning continues. If the stain is still obvious, trigger the intensive cleaning mode (such as increasing the cleaning agent concentration, extending the scrubbing time) or mark the area, and notify the user through the APP that “stubborn stains need manual processing”.

[0124] Due to the multi-dimensional verification screening and scenario-based adaptive evaluation, the problem of ground cleaning parameter matching being divorced from stain properties and actual scene is solved, and the pertinence of cleaning parameters, operation safety and cleaning efficiency are improved.

[0125] Based on any of the above embodiments, in the eighth embodiment of the present application, the step S50 includes steps G11-G13: Step G11, adaptively verify the target cleaning parameter and the classification and position information of the stain to obtain a parameter adaptation verification result.

[0126] In this embodiment, the stain category refers to the identification of the material composition of the stain. The stain area size refers to quantitative information representing the coverage range of the stain. The stain position distribution refers to the distribution characteristics of the stain in the spatial coordinate system. The adaptability refers to the matching degree of the optimal cleaning parameters and the relevant characteristics of the stain category and the position information of the stain. The parameter adaptation verification result refers to the conclusive information representing the matching state of the parameters and the stain characteristics after adaptability verification. The verification dimensions of the adaptability verification include the stain category, the stain area size, and the stain position distribution.

[0127] As an optional implementation, the stain category, the stain area size, and the position distribution in the position information of the stain are first analyzed to determine the key attributes and verification standards of each dimension. Then, the key indicators in the optimal cleaning parameters that are adapted to each characteristic dimension are correspondingly disassembled to establish a precise correspondence between the characteristic dimensions and the parameter indicators. Depth adaptability verification is carried out dimension by dimension to verify the processing pertinence of the cleaning parameters for the stain category, the intensity / time adaptation degree for the area size, and the range / path coverage degree for the position distribution. At the same time, cross-dimension cross-verification is performed to verify the logical consistency of the adaptation results of different dimensions, avoiding single-dimension adaptation bias. Finally, the verification results of each dimension and the cross-verification conclusions are integrated to form the parameter adaptation verification result containing the adaptation degree, potential bias, and correction suggestions. This method has comprehensive verification dimensions and in-depth verification, and the result has strong accuracy and integrity, effectively avoiding the problem of poor cleaning effect caused by improper adaptation.

[0128] Step G12, if there is an adaptation bias in the parameter adaptation verification result, the cleaning parameters are fine-tuned according to the stain category and position information of the stain to generate calibrated cleaning execution parameters.

[0129] In this embodiment, the adaptation bias refers to the difference between the parameters and the core characteristics of the stain in the result. The fine-tuning of the cleaning parameters refers to the precise and meticulous adjustment of the cleaning parameters in response to the adaptation bias. The calibrated cleaning execution parameters refer to the directly executable cleaning parameters that are adapted to the core characteristics of the stain after fine-tuning and optimization.

[0130] As an optional implementation, the adaptation deviation dimension in the parameter adaptation verification result is first analyzed to determine the corresponding stain category and position information characteristics of the deviation, and a precise association between the deviation and the stain characteristics is established. Then, the adjustable indicators of the cleaning parameters are disassembled, and a fine-tuning scheme is developed for each deviation dimension: for category deviation, the composition processing targeted indicators of the parameters are adjusted; for area deviation, the parameter strength or duration is optimized; for position deviation, the coverage range indicators of the parameters are corrected. During the fine-tuning process, each adjustment of a parameter is verified for adaptation with the corresponding stain core characteristics, and through multiple rounds of iterative correction, all deviations are eliminated, and cross-indicator logical verification is simultaneously performed to ensure the consistency between parameters. Finally, all fine-tuned parameters are integrated to generate calibrated cleaning execution parameters. This method has strong fine-tuning pertinence and high precision, can completely eliminate all types of adaptation deviations, has good parameter adaptation, and avoids the problem of incomplete cleaning caused by parameter deviation.

[0131] As another optional implementation, the receiving operation of the parameter adaptation verification result is first performed, and the integrity of the received result information is verified to confirm the key contents including the adaptation judgment conclusion and the core dimension matching state. Then, the verification pass identifier in the result is analyzed, and the preset adaptation standard is simultaneously called to recheck the reasonableness of the verification conclusion and verify the matching consistency of the parameters with the core characteristics of the stain category, area size, and position distribution. After confirming that the rechecking has no objection and no hidden deviation is found, the parameter reuse mechanism is started, all configuration indicators of the optimal cleaning parameter are completely retained, and the optimal cleaning parameter is directly used as the basis parameter for subsequent cleaning execution. This method has a rigorous process, avoids misjudgment through rechecking, and has high reliability in parameter application; the disadvantage is that the rechecking link is added, the overall processing time is slightly longer, and the system information processing efficiency has certain requirements; the effect is to ensure that the reused optimal cleaning parameter completely meets the adaptation standard, provides stable and reliable parameter support for cleaning operation, and guarantees the consistency of cleaning effect.

[0132] As another optional implementation, the adaptation deviation in the parameter adaptation verification result is analyzed, and according to the severity of the impact on the cleaning effect, the deviation is divided into three levels: core deviation, important deviation, and secondary deviation, and an association mapping between the deviation level and the stain category and position information characteristics of the stain is established. For core deviation, the key adjustable indicators of the cleaning parameters are deeply disassembled, a special fine-tuning scheme is developed, and the corresponding parameters are greatly adjusted. For important deviation, the core indicators of the parameters are moderately adjusted. For secondary deviation, the auxiliary indicators of the parameters are fine-tuned. During the fine-tuning process, after completing the correction of a level of deviation, the corresponding stain core characteristics are verified for adaptation, all deviations are eliminated through multiple rounds of iterative correction, and the logical consistency between parameters is verified, and finally the fine-tuned parameters are integrated to generate calibrated cleaning execution parameters. This method has strong deviation handling pertinence and high fine-tuning precision, can completely eliminate deviations of different levels, has good parameter adaptation, and can maximize the guarantee of cleaning effect.

[0133] Step G13, according to the calibrated cleaning execution parameters, combined with the current position information of the device and the environment map, to generate the target cleaning strategy.

[0134] In this embodiment, the current position information of the device refers to the spatial positioning information of the device at the moment. The environment map refers to the map data representing the spatial layout and obstacle distribution of the cleaning area.

[0135] As an optional implementation, the core indicators such as intensity, duration, coverage range in the calibrated cleaning execution parameters are integrated, the cleaning starting point is determined in combination with the current position information of the device, and the spatial layout and obstacle distribution characteristics in the environment map are analyzed synchronously. Based on the category of stains and the position distribution in the position information of stains, the optimal cleaning path covering all stain areas and avoiding obstacles is planned, the corresponding cleaning execution parameters are assigned according to the path nodes, and the parameter adaptability of different areas is ensured. The phased execution process is formulated, the path order, the cleaning duration of each node and the intensity switching logic are determined, and the real-time feedback mechanism in the cleaning process is set. The path planning, parameter allocation and execution process are integrated to generate a target cleaning strategy with complete structure and rigorous logic, ensuring that the device can autonomously traverse the cleaning area according to the strategy and accurately perform various cleaning operations. This method has fine strategy planning, high path and parameter matching degree, and strong cleaning pertinence and efficiency.

[0136] Exemplarily, refer to Figure 3 , Figure 3The system block diagram of the present application. Each subsystem and module is described in detail. Perception subsystem: responsible for collecting all environmental raw data. Navigation and positioning sensors: function: realize the autonomous movement of the robot, obstacle avoidance and map construction. Typical hardware: laser radar, inertial measurement unit, ultrasonic sensor, wheel odometer, etc. Stain detection visual sensor: function: used for high-definition collection of ground images, is the data source for stain identification. Typical hardware: high-definition RGB camera, RGB-D depth camera (can provide 3D information), special light source (such as polarized light, used to enhance the liquid reflection characteristics). Environmental perception module: function: preprocess, fuse, and coordinate the raw data of the above sensors, providing standardized data for subsequent algorithms. Central processing and control subsystem: contains the main algorithm and logic control unit. Master control unit: function: the computing core of the system, responsible for scheduling all software modules, managing task flow, and communicating with the outside. It is usually a high-performance embedded computing platform (such as NVIDIA Jetson, Huawei Atlas, etc.). Stain identification algorithm module (core AI algorithm): function: runs a pre-trained deep learning model (such as U-Net, DeepLabV3+, etc. semantic segmentation network), performs pixel-level analysis on the input image, and outputs the precise outline and position information of the stain. Key data: input is image, output is "stain segmentation mask" and "stain bounding box". Decision and path planning module: function: makes intelligent decisions based on the identified stain information. Task decision: judges the severity of the stain and determines the cleaning priority. Path planning: generates the optimal cleaning path that covers the entire stain area with the highest efficiency (such as reciprocating, spiral). Motion control module: function: decomposes the planned path into specific motor control instructions (speed, direction), and accurately controls the mobile chassis to move to the target position. Cleaning control module: function: generates control instructions for the cleaning mechanism. For example, control the switch and water volume of the water pump, control the start and speed of the roller brush, control the suction force, etc. Execution subsystem: responsible for executing specific physical actions. Mobile chassis mechanism: function: carries the entire system, and realizes forward, backward, turning, etc. actions according to the motion control instructions. Typical hardware: drive wheels, motors, reducers, etc. Cleaning execution module: function: executes specific cleaning actions. Typical hardware: miniature constant water pump, liquid tank (clean water / sewage), cleaning roller brush, vacuum moisture removal device, scraper, etc. Human-computer interaction and communication subsystem (optional but important): function: provides an interface for users to interact with the system, and the ability for the system to communicate with external networks. Typical components: local UI interface: power on / pause button, status indicator light, touch screen (used to display identification results, cleaning map, system status). Cloud / remote monitoring platform: upload cleaning reports, stain images, and device status to the cloud through Wi-Fi / 4G / 5G, support remote monitoring, management and algorithm model OTA upgrade.

[0137] Due to the parameter adaptation verification and fine tuning, and the scene-based strategy generation, the problem of mismatch between parameters and stain characteristics, and the problem of lack of targeted strategy in ground cleaning are solved, and the accuracy of cleaning operation and the self-cleaning ability of the equipment are improved.

[0138] The application provides an autonomous cleaning device, which comprises at least one processor, and a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the autonomous cleaning method for ground stains in the embodiment I.

[0139] Reference will be made to the accompanying drawings Figure 4 which shows a structural schematic diagram of an autonomous cleaning device suitable for implementing the embodiments of the application. The autonomous cleaning device in the embodiments of the application can include but is not limited to mobile terminals such as mobile phones, notebook computers, smart sweeping robots, personal digital assistants (PDA), tablet computers (PAD), portable multimedia players (PMP), professional autonomous cleaning devices, and the like, and fixed terminals such as special cleaning robots, desktop computers, and the like. Figure 4 The illustrated autonomous cleaning device is only an example and should not impose any limitation on the functions and use range of the embodiments of the application.

[0140] As Figure 4As shown, the autonomous cleaning device can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the autonomous cleaning device are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. In general, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the autonomous cleaning device to communicate wirelessly or wired with other devices to exchange data. Although the autonomous cleaning device having various systems is shown in the figure, it should be understood that all of the shown systems are not required to be implemented or possessed. More or less systems can be alternatively implemented or possessed.

[0141] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0142] The autonomous cleaning device provided by the present disclosure adopts the autonomous cleaning method for ground stains in the above embodiments, and can solve the technical problem of poor cleaning effect. Compared with the prior art, the autonomous cleaning device provided by the present disclosure has the same beneficial effects as the autonomous cleaning method for ground stains provided by the above embodiments, and other technical features in the autonomous cleaning device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0143] It should be understood that various aspects of the disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0144] The above description is merely illustrative of the application and is not intended to limit the scope of the application. Any variations and modifications that can be made by any person skilled in the art within the spirit and scope of the application are intended to be encompassed by the application. The scope of the application is defined by the appended claims.

[0145] The application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e., a computer program) for performing the autonomous cleaning method of ground stains in the above-described embodiments.

[0146] The computer readable storage medium provided by the application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any appropriate medium, including but not limited to an electrical wire, an optical cable, a radio frequency (RF), or the like, or any appropriate combination thereof.

[0147] The above computer readable storage medium can be contained in the autonomous cleaning device; or can exist separately without being assembled into the autonomous cleaning device.

[0148] The computer readable storage medium described above carries one or more programs, when the one or more programs are executed by the autonomous cleaning device, the autonomous cleaning device: collects original data of a to-be-processed area; processes the original data according to an information processing rule corresponding to a data type of the original data to obtain multi-modal data; splices the aligned multi-modal data according to an inter-modal pixel adaptation strategy to obtain fusion data, and analyzes a correlation relationship of each material characteristic feature in the fusion data to determine a category of the stain; extracts a high-level feature of each modality representing a stain distribution in the multi-modal data and performs weighted fusion to generate a segmentation mask to define position information of the stain through the segmentation mask; processes the multi-modal data according to a decision-level fusion rule, matches a target cleaning parameter in combination with a cleaning strategy database, and verifies the target cleaning parameter and the category of the stain and the position information of the stain to obtain a target cleaning strategy; and controls the cleaning device to clean the to-be-processed area according to the target cleaning strategy.

[0149] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0150] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

[0151] The modules involved in the embodiments of the present application can be implemented in a software manner or in a hardware manner. In some cases, the names of the modules do not constitute a limitation on the modules themselves.

[0152] The readable storage medium provided by the application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the autonomous cleaning method of ground stains described above, and can solve the technical problem of poor cleaning effect. Compared with the prior art, the beneficial effects of the computer readable storage medium provided by the application are the same as those of the autonomous cleaning method of ground stains provided by the above-mentioned embodiments, and are not described here.

[0153] The above only describes some embodiments of the application, and does not limit the patent scope of the application. Any equivalent structural transformation made by using the content of the specification and drawings, or direct / indirect application in other related technical fields under the technical concept of the application is included in the patent protection scope of the application.

Claims

1. A method for self-cleaning floor stains, characterized in that, The method includes: Collect raw data from the area to be processed; The original data is processed according to the information processing rules corresponding to the data type of the original data to obtain multimodal data. Based on the intermodal pixel adaptation strategy, the spliced ​​and aligned multimodal data is spliced ​​to obtain fused data, and the correlation relationship of the material characteristics in the fused data is analyzed to determine the component category of the stain. High-level features representing the distribution of stains in each modality are extracted from the multimodal data and weighted and fused to generate a segmentation mask, which is used to define the location information of the stains. The multimodal data is processed according to decision-level fusion rules, and the target cleaning parameters are matched with the cleaning strategy database. The target cleaning parameters are then verified to match the composition category of the stain and the location information of the stain to obtain the target cleaning strategy. The cleaning equipment is controlled to clean the area to be treated according to the target cleaning strategy.

2. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of processing the original data according to the information processing rules corresponding to the data type of the original data to obtain multimodal data includes: The raw data is divided according to data type, and corresponding rules are selected from the information processing rule base to determine the information processing rules corresponding to each data type. The visual data in the original data is denoised by edge-preserving algorithm, the spectral data in the original data is baseline-corrected, and the acoustic data in the original data is filtered to obtain the processed single-modal data. The processed single-modal data are associated and integrated according to preset spatiotemporal alignment rules to output multimodal data.

3. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of stitching and aligning the multimodal data according to the intermodal pixel adaptation strategy to obtain the fused data includes: The aligned multimodal data is subjected to resolution normalization processing, and based on the intermodal pixel adaptation strategy, the pixel coordinates of each modal data are accurately mapped to generate multimodal pixel-level data. According to the preset channel combination rules, the multimodal pixel-level data is spliced ​​and fused to obtain the initial fused data; The initial fused data is formatted and validated, and then the fused data is output.

4. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of analyzing the correlation between the characteristics of various substances in the fused data and determining the component category of the stain includes: Features reflecting stain characteristics are extracted from pixel information corresponding to the component categories of the stain, forming a multimodal material characteristic feature set; By conducting correlation analysis on the different material characteristics in the multimodal material characteristic feature set, the complementary correlation relationships between the characteristics are determined; The complementary correlation is inferred according to the preset mapping rules between material properties and categories, and the component category of the stain is determined.

5. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of extracting high-level features representing the stain distribution of each modality in the multimodal data and weighting and fusing them to generate a segmentation mask includes: According to the modalities of the multimodal data, high-level features characterizing stain outline, stain morphology and material properties are extracted respectively to obtain an independent set of high-level features for each modality; Based on the contribution of each modality to stain recognition, an adaptation weight is assigned, and the high-level feature sets corresponding to each modality are weighted and fused to output the fused features. The fused features are classified at the pixel level using a semantic segmentation decoder, and the stain attribution attribute of each pixel is labeled to obtain the segmentation mask that is consistent with the stain size in the original image.

6. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of defining the location information of the stain using the segmentation mask includes: The segmentation mask is parsed to extract the complete outline, boundary range, spatial coordinates, and area parameters of the stain, and then integrated to obtain the stain location feature data. The stain location feature data and the stain composition category are correlated and fused, and the corresponding attribute information of the stain is supplemented to obtain comprehensive stain information data. Based on the comprehensive stain information data, the location change trend of the stain is predicted, and location information describing the stain distribution is generated.

7. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of processing the multimodal data according to decision-level fusion rules and matching target cleaning parameters with the cleaning strategy database includes: The location parameters, composition categories, area data, and characteristic confidence scores of the stains in the multimodal data are analyzed to form stain attribute data; Based on a preset validity threshold, abnormal data and invalid information in the stain attribute data are removed, and key stain attribute information is obtained by filtering. The pre-stored cleaning strategy database is invoked, and the key attribute information of the stain is matched with the database to obtain a set of candidate cleaning parameters; By combining the current ground material information and the current status of the equipment, the candidate parameters are evaluated for suitability and prioritized to obtain suitable parameters. The suitable parameters are then fused according to the decision-level fusion rules to generate the target cleaning parameters.

8. The self-cleaning method for floor stains as described in claim 1, characterized in that, The step of verifying the target cleaning parameters with the component category and location information of the stain to obtain the target cleaning strategy includes: The compatibility of the target cleaning parameters with the composition, type, and location information of the stain is verified to obtain the parameter compatibility verification result. If there is an adaptation deviation in the parameter adaptation verification result, the cleaning parameters are finely adjusted according to the composition and location information of the stain to generate calibrated cleaning execution parameters; Based on the calibrated cleaning execution parameters, combined with the current location information of the equipment and the environmental map, the target cleaning strategy is generated.

9. A self-cleaning device, characterized in that, The self-cleaning device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the self-cleaning method for floor stains as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the autonomous cleaning method for floor stains as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent cleaning method of sweeping robot based on image recognition

    CN120477632A

  • Intelligent washing parameter optimization decision-making method and system based on multi-modal perception fusion

    CN120491496A

  • Road surface cleaning decision-making method and device based on image recognition, equipment and medium

    CN120853127A

  • Sweeping robot control method and system

    CN120918529A

  • Sweeping robot cleaning control system with self-adaptive mode adjustment

    CN120959623A