Surface dirty image recognition method applied to window cleaning robot and related device

Through multi-spectral imaging and safety signal fusion technology, window cleaning robots can accurately identify window dirt under different lighting conditions, avoid safety hazards, optimize cleaning paths, and achieve efficient and safe cleaning effects.

CN120472414AInactive Publication Date: 2025-08-12SHENZHEN YIJIE INTELLIGENT TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510955027.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional window cleaning robots cannot adjust cleaning strategies based on the actual dirt distribution of window surfaces, resulting in poor cleaning results or waste of resources, and it is difficult to accurately identify transparent or translucent dirt under different lighting conditions and ensure safety.

Method used

Multispectral imaging technology is used to collect images on the window surface, and a window boundary safety map is constructed by combining side wheel collision signals and drop detection signals, dynamic weighted mixed attention processing and convolution multi-scale strip pooling treatment are carried out to generate dirty segmentation results and cleaning strategy mapping tables, and cleaning paths are planned through dirty density heat maps and improved RRT* algorithm.

Benefits of technology

It realizes accurate identification of multiple dirt types under different lighting conditions, avoids collision and fall accidents, formulates personalized cleaning parameters, optimizes cleaning paths, and improves cleaning efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472414A_ABST
    Figure CN120472414A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses a surface dirty image recognition method applied to a window-cleaning robot and a related device, and the method comprises the steps: carrying out the preprocessing of a multispectral image collected by the window-cleaning robot, and obtaining a preprocessed image, meanwhile, integrating a side wheel collision signal and a falling detection signal of the window cleaning robot to obtain a window boundary safety map; performing dynamic weighted mixed attention processing and convolution multi-scale stripe pooling processing under edge constraint to obtain an enhanced feature map; edge security perception smudginess segmentation is carried out on the enhanced feature map to obtain a smudginess segmentation result and an edge security region identifier, and a cleaning strategy mapping table for smudginess at different positions is established; according to the method, a dirt density thermodynamic diagram is generated, a path planning algorithm is executed, and a safe cleaning track is generated, so that stable spectral response can be obtained under various illumination conditions, and the stability and reliability of dirt recognition in different environments are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a surface dirt image recognition method and related devices used in a window cleaning robot. Background Art

[0002] Traditional window-cleaning robots often use fixed paths (such as zigzag or spiral patterns) for cleaning operations. These strategies are unable to adapt to the actual distribution of dirt on the window surface, resulting in poor cleaning results and wasted resources. Furthermore, the diverse types of dirt on window surfaces (water stains, fingerprints, oil stains, dust, etc.) require different cleaning parameters, but existing technologies often use a single cleaning method, making it difficult to cope with complex and changing dirt scenarios.

[0003] The unique physical properties of window surfaces pose significant challenges to stain image recognition. Window materials like glass are highly reflective, producing strong light reflections and shadow interference under varying lighting conditions, resulting in blurred stain edges and unclear features. This is particularly true for transparent or translucent stains like water stains and fingerprints, which have low contrast with the background and are difficult to distinguish effectively using traditional RGB image processing methods. Furthermore, window-cleaning robots operate at high altitudes, posing safety risks to window edges and special structural areas. Existing stain recognition technologies lack effective integration with safety mechanisms, making it impossible to achieve precise cleaning while ensuring safety. Summary of the Invention

[0004] The present application provides a surface dirt image recognition method and related devices used in a window cleaning robot, which can obtain a stable spectral response under various lighting conditions, ensuring the stability and reliability of dirt recognition in different environments.

[0005] In a first aspect, the present application provides a surface dirt image recognition method for a window cleaning robot, the surface dirt image recognition method for a window cleaning robot comprising: The multispectral images collected by the window cleaning robot are preprocessed to obtain a preprocessed image. At the same time, the window cleaning robot's wheel collision signal and fall detection signal are integrated to obtain a window boundary safety map. Performing edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing on the preprocessed image and the window boundary safety map to obtain an enhanced feature map; Performing edge safety perception dirt segmentation on the enhanced feature map to obtain dirt segmentation results and edge safety area identification, and establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot to obtain a cleaning strategy mapping table for dirt at different locations; A dirt density heat map is generated based on the dirt segmentation result and the cleaning strategy mapping table, and a path planning algorithm is executed to generate a safe cleaning trajectory.

[0006] A second aspect of the present application provides a surface dirt image recognition device for a window cleaning robot, the surface dirt image recognition device for the window cleaning robot comprising: The preprocessing module is used to preprocess the multispectral images collected by the window cleaning robot to obtain a preprocessed image, and at the same time integrate the window cleaning robot's side wheel collision signal and fall detection signal to obtain a window boundary safety map; a convolution processing module, configured to perform edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing on the pre-processed image and the window boundary safety map to obtain an enhanced feature map; Establish a module for performing edge safety perception dirt segmentation on the enhanced feature map, obtaining dirt segmentation results and edge safety area identification, and establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot to obtain a cleaning strategy mapping table for dirt at different locations; A generation module is used to generate a dirt density heat map based on the dirt segmentation result and the cleaning strategy mapping table, and execute a path planning algorithm to generate a safe cleaning trajectory.

[0007] The third aspect of the present application provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned surface dirt image recognition method for a window cleaning robot.

[0008] A fourth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the above-mentioned surface dirt image recognition method for a window cleaning robot.

[0009] Compared with existing technologies, the present application has the following advantages: Through multispectral imaging and dynamic weighted hybrid attention processing, the method of the present invention can accurately distinguish various types of dirt, such as water stains, fingerprints, oil stains, dust, and mixed stains. It is particularly effective in identifying transparent or translucent dirt, overcoming the limitations of traditional RGB image processing methods. By integrating edge wheel collision signals and fall detection signals to construct a window boundary safety map, the system achieves a deep integration of dirt identification and safety assurance, effectively preventing collisions or falls during window cleaning robot operations. A cleaning strategy mapping table established based on dirt segmentation results and edge safety area identification can customize cleaning parameters for different types and locations of dirt, achieving precise cleaning. Through a dirt density heat map and an improved RRT* algorithm, the method of the present invention achieves intelligent cleaning path planning under safety constraints, dynamically adjusting path density and cleaning parameters to avoid dangerous areas while prioritizing key dirty areas. The use of multispectral imaging technology and professional image preprocessing enables the system to obtain stable spectral responses under various lighting conditions, ensuring the stability and reliability of dirt identification in different environments. Based on precise identification and path optimization of dirt distribution characteristics, the window cleaning robot can concentrate its limited battery capacity and cleaning resources on areas that are truly needed, reducing unnecessary cleaning operations and extending single operation time. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which the present invention can be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size should still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and objectives that can be achieved by the present invention.

[0012] Figure 1 1 is a flow chart of a surface dirt image recognition method for a window cleaning robot provided by an embodiment of the present invention; Figure 2 This is a schematic block diagram of the structure of a surface dirt image recognition device for a window cleaning robot provided by an embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiment is only one embodiment of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0014] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or merged, so the actual execution order may vary depending on the actual situation.

[0015] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0016] It should be further understood that the term "and / or" used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations. Figure 1 In one embodiment of the present application, a surface dirt image recognition method for a window cleaning robot includes: Step 100: Preprocess the multispectral image collected by the window cleaning robot to obtain a preprocessed image, and integrate the side wheel collision signal and fall detection signal of the window cleaning robot to obtain a window boundary safety map; It is understandable that the execution subject of this application can be a surface dirt image recognition device used in a window cleaning robot, or a terminal or a server, and the specific implementation is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0017] Specifically, a multispectral imaging system integrated into the window-cleaning robot captures images of the window surface. The imaging system covers the visible, near-infrared, and ultraviolet spectrum, generating a multispectral image. The multispectral image is then denoised using an improved non-local means filtering algorithm or an adaptive noise suppression algorithm. The filter window size and intensity are dynamically adjusted based on the actual noise level, preserving image edges and detail features as much as possible while suppressing the influence of high-frequency noise. This results in noise-suppressed image data. The denoised multispectral image is then radiometrically corrected using a lookup table correction method to normalize the imaging response of each band to eliminate systematic deviations caused by uneven sensor response and differences in incident energy across bands. This results in a multispectral corrected image. A phase correlation registration algorithm is then used to perform subpixel geometric alignment of the bands in the corrected multispectral image. By calculating the phase correlation characteristics between the reference band and the remaining bands and inferring the optimal transformation parameters, an affine or rigid transformation is then applied to all bands to achieve subpixel spatial consistency, ensuring that image pixels at the same spatial location correspond to each other, resulting in a preprocessed image. At the same time, during actual operation, the window-cleaning robot continuously acquires information related to the window boundary and safety status through various safety sensors on its main structure. For example, a touch button set on the side wheel socket can generate a side wheel collision signal in real time when the robot approaches the window edge and makes physical contact, while the drop bar and photoelectric switch output a fall detection signal when the robot is at risk of falling. The system synchronously collects these safety signals and marks them in the robot's current two-dimensional or three-dimensional coordinate system according to the spatial position corresponding to these signals. It gradually establishes a set of marker points for the window boundary and potential fall risk areas, and integrates these spatial points into the global coordinate system through interpolation, topological mapping, and other methods, forming a window boundary safety map that is highly aligned with the preprocessed image data.

[0018] Step 200: Perform edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing on the pre-processed image and the window boundary safety map to obtain an enhanced feature map; Specifically, the preprocessed image is fed into the spatial attention processing module. Using a multi-scale atrous convolutional architecture and a parallel receptive field strategy, it captures diverse spatial contextual information on the window surface across different spatial ranges, effectively distinguishing local details from global distribution features. It also makes preliminary spatial distinctions between dirty and non-dirty areas, generating a spatial feature map of the window surface that reflects spatial dependencies at different scales. Taking advantage of the multi-channel nature of the preprocessed image, a channel-wise attention analysis is performed. Based on global average pooling, fully connected weight assignment, and adaptive activation function adjustment, the response strengths of different channels are modeled. This extracts channel features that are more sensitive to dirt signatures, generating a channel-wise feature map of the window surface. Furthermore, considering that multispectral imaging covers multiple bands, including visible, near-infrared, and ultraviolet, a spectral attention mechanism is introduced for each band. Using a self-attention structure, the correlation matrix between spectra is dynamically modeled. This architecture enhances the useful complementary information between bands while suppressing the influence of redundant or noisy bands, resulting in a more discriminative spectral feature map of the window surface. Edge region feature extraction and edge safety mapping analysis are performed on the window boundary safety map. Using depthwise separable convolution, gradient operators, and spatial weight mapping, the system effectively extracts spatial structural features and fall risk distribution near window edges, organizing these features into a window edge safety constraint data map. The preprocessed image and window boundary safety map are fed into a conditional gating unit, built on a lightweight convolutional neural network. This unit comprehensively analyzes the input data's feature representations under different scenarios and boundary conditions and dynamically generates a set of normalized weight coefficients reflecting the combined importance of spatial, channel, spectral, and safety constraint features within the current window environment. Based on these normalized weight coefficients, the window surface spatial feature map, window surface channel feature map, window surface spectral feature map, and window edge safety constraint map are weightedly fused to produce an edge safety constraint feature map. The edge safety constraint feature map is then subjected to safety-aware convolutional multi-scale stripe pooling. Using parallel stripe convolution and an adaptive pooling structure, the system performs deep mining of horizontal, vertical, diagonal, and point features. Safety masks are then incorporated to ensure that feature processing in high-risk areas does not lead to system misjudgments or operational safety hazards. The multi-scale features output by all branches are integrated through the security enhancement feature fusion network to form an enhanced feature map.

[0019] In this embodiment, the edge safety constraint feature map is fed into multiple dedicated convolution branches. The system includes horizontal stripe convolution branches, vertical stripe convolution branches, diagonal stripe convolution branches, and point convolution branches. The horizontal stripe convolution branch uses a 1×7 convolution kernel to enhance horizontal structure and stripe information while maintaining image resolution, effectively capturing directional dirt areas such as horizontal water stains and scratches. The vertical stripe convolution branch uses a 7×1 convolution kernel to focus on vertical stripes or stains, improving the robot's perception resolution of vertical stripe-shaped stains on windows. Meanwhile, the diagonal stripe convolution branch uses two 7×7 convolution kernels at a 45° angle to each other, enhancing the system's ability to extract complex stain patterns such as tilt, oblique drag marks, and special fingerprints in high-dimensional space. The point convolution branch uses a 1×1 convolution kernel to focus on extracting sparse but important point-shaped stains or subtle dirt signals on the window surface, achieving all-round, multi-scale, and multi-directional feature extraction. The signals collected in real time by the side wheel collision device and the fall detection device during the operation of the integrated window cleaning robot are used to automatically generate a safety mask corresponding to the spatial distribution based on this safety information. The mask calibrates the current physical boundaries of the window surface, potential fall risk areas and other key safety areas in a pixel-level manner. In order to ensure that no safety hazards are caused by the processing of high-risk areas during feature enhancement, all directional feature maps obtained through the convolution branch are convolved and superimposed with the safety mask, and only valid feature responses within the safe area or controlled boundary range are retained. After the multi-directional safety constraint feature data undergoes convolution operation, the system implements attenuated pooling processing for feature distributions of different scales and directions, and uses learnable parameters to adjust the maximum pooling and average pooling weights to achieve a dynamic balance between information intensity and noise interference, which can not only ensure the integrity of significant features, but also suppress isolated pulses and invalid high response areas. The multi-directional feature data after security mask constraint and attenuation pooling processing is input into the security enhancement feature fusion network. The network fully integrates all key features in the horizontal, vertical, diagonal and point directions in high-dimensional space, and adaptively adjusts the fusion strategy according to the weight of the security area, dynamically enhancing the feature contribution of high-value areas, while suppressing the interference of high-risk areas or information noise to obtain an enhanced feature map.

[0020] Step 300: Perform edge safety perception dirt segmentation on the enhanced feature map to obtain dirt segmentation results and edge safety area identification, establish a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot, and obtain a cleaning strategy mapping table for dirt at different locations; It is important to note that the enhanced feature map is input into the encoder portion of the dirt segmentation network. This encoder utilizes a multi-layer convolution and pooling architecture to extract multi-scale spatial features across different receptive fields, compressing the feature map's resolution layer by layer while continuously increasing its level of abstraction. This results in a multi-scale encoded feature map that covers both global and detailed information. The multi-scale features output by the encoder are decoded layer by layer and their resolution restored using transposed convolution or upsampling modules. Leveraging cross-layer skip connections and batch normalization, a hierarchical, multi-scale decoded feature map is generated. Edge response analysis is performed on the multi-scale encoded feature map using wheel collision and fall detection signals collected in real time by the window-cleaning robot. An edge response map is generated using computational methods similar to Sobel or other gradient operators, effectively highlighting the spatial characteristics of the window's physical boundaries, corners, and potential risk areas. This edge response feature serves as an important spatial prior, directly impacting the segmentation accuracy and safety constraints of boundary regions during feature fusion and post-processing. Simultaneously, the multi-scale encoded feature map is combined with safety sensor signals to perform a safety risk assessment. Using mechanisms such as safety mask generation and risk weight assignment, edge safety constraint features are extracted to describe the risk level and safety status of each spatial region. These features are then used to modulate the network's response to high-risk areas, improving the operational safety of the segmentation results. During the decoding phase, the generated edge response features and edge safety constraint features are adaptively fused with the decoded feature map. This fusion process utilizes methods such as spatial adaptive normalization, conditional gated fusion, or an attention mechanism to dynamically adjust the influence of each feature on the final decoding result based on boundary signals and safety weights. The fused decoding output achieves high levels of pixel-level classification accuracy for dirt segmentation, boundary detail characterization, and safety risk zone identification. Two core outputs are directly generated: a multi-class dirt segmentation result that distinguishes different types of dirt, such as water stains, oil stains, fingerprints, and dust, and their accurate spatial distribution; and a pixel-level spatial identification of edge safety zones, providing a high-precision safety reference for the robot during cleaning operations. Based on the segmentation results and safety zone identification, the system extracts information such as the spatial geometric characteristics of each stain instance, stain type, severity, and risk distance from the safety boundary. Combined with the actual operational capabilities and cleaning parameters of the window cleaning robot, the system uses fuzzy rules, decision trees, or deep neural mapping to establish a correspondence between stain characteristics and the robot's cleaning parameters. The system outputs a cleaning strategy mapping table, indexed by spatial region, stain type, and risk level. This table specifies the cleaning force, speed, water volume, operation mode, and even whether to avoid or perform special treatment for each stain location under different safety constraints, achieving refined and differentiated cleaning operations and ensuring safety throughout the entire process.

[0021] In this embodiment, connected domain analysis is performed on the stain segmentation results. By determining pixel connectivity within the segmented image, spatially contiguous stain regions of the same type are clustered into a set of independent stain instances, providing a foundation for modeling each real stain region and subsequent parameter configuration. For each stain instance in the set, morphological analysis is used to characterize its spatial shape and geometric characteristics in multiple dimensions, including geometric parameters such as the stain's area, perimeter, principal direction, aspect ratio, compactness, and centroid location. Material features are also extracted for each stain instance. Based on multispectral imaging data or high-resolution images, statistical analysis of pixel values—such as mean, variance, skewness, and kurtosis—is performed. Combined with features such as multispectral reflectance and band normalized difference, a feature vector is generated that reflects the stain's material, thickness, uniformity, and spectral response. This improves the robot's ability to automatically distinguish complex stains (such as water stains, oil stains, fingerprints, and dust) and adapt to cleaning modes. Combined with previously acquired edge safety zone identification data, each contamination instance is precisely spatially mapped to the actual physical boundaries of the window and potential fall risk areas. Spatial calculations are used to determine the shortest distance from each contamination instance to the window edge and the closest distance to the fall risk zone. Edge risk values are calculated based on spatial distribution, geometric location, and known risk zone weights, generating safety feature data that reflects the safety requirements of cleaning operations. Deep feature fusion is performed on basic geometric, material, and safety features. Using a multi-head attention mechanism, feature concatenation weighting, or deep neural network-based feature compression and embedding mapping, multimodal features are integrated into a target contamination feature vector with consistent length and high information density. This target contamination feature vector is then input into an edge safety perception fuzzy decision tree for multi-level reasoning and analysis. The decision tree, incorporating a rule base, fuzzy logic, and a multi-level safety constraint mechanism, automatically infers the most appropriate cleaning mode, cleaning intensity, speed, water volume, and special treatment strategies based on the type of contamination, severity, morphological characteristics, material properties, and edge risk level. For example, for oil stains near edges or high-risk areas, the decision tree will proactively reduce pressure, slow down the speed, or even switch to safe mode or skip cleaning. For large, non-risky water stains in the center, the efficient standard mode is used for rapid cleaning. All analysis results are ultimately organized into a cleaning strategy mapping table, covering the optimal cleaning parameter allocation for different spatial locations, stain types, and risk levels.

[0022] Step 400: Generate a dirt density heat map based on the dirt segmentation result and the cleaning strategy mapping table, and execute a path planning algorithm to generate a safe cleaning trajectory.

[0023] Specifically, by performing regional statistics and weighted fusion processing on the dirt segmentation image, the segmentation results are weighted and superimposed according to the distribution of different types of dirt, the severity of each category of stains, and the importance of cleaning each area, to obtain a dirt density heat map that can reflect the spatial distribution characteristics. Among them, different types of stains such as water stains, oil stains, fingerprints, and dust are assigned different priority weights. The severity of the stains is assigned using indicators such as pixel density and reflective characteristics, and the regional importance is quantified based on the visual or functional impact of different spaces such as the center, edge, and corner of the window. In the heat map generation process, the side wheel collision areas and fall risk areas detected by the window cleaning robot are treated as special safety constraint areas and assigned extremely low or even zero cleaning priority to ensure that these areas are fully avoided or given key safety protection in subsequent path planning and task allocation. Based on the dirt density heat map, the improved RRT* path planning algorithm is used as the main means of generating intelligent paths. When initializing the RRT search space, a heatmap-guided sampling strategy is used to assign high-probability sampling points to areas with high dirt density. Specifically, the sampling probability is proportional to the current pixel's dirt density. This increases the sampling frequency in high-density areas, ensuring that the main path trunk preferentially traverses key cleaning areas and improving overall efficiency. During RRT tree node expansion, the system dynamically adjusts the expansion step size based on the current node's dirt density as determined by the heatmap and its distance from the nearest safety constraint. Smaller step sizes are used in high-density dirt areas and areas far from the safety zone to achieve path segmentation and increased coverage density, while larger step sizes are applied in low-density areas or near safety boundaries to improve path generation efficiency and safety. This process enables the node expansion sequence to adapt to the actual dirt distribution on the window surface and environmental safety constraints. Based on heatmap-guided high-frequency sampling and dynamic step size expansion, the initially generated path is optimized to accommodate complex environments. Local replanning is performed on each path segment, including smoothing, backoff for obstacle avoidance, and insertion of detour nodes. The result is an initial cleaning path that covers all key dirt areas and avoids high-risk areas. Based on the established cleaning strategy mapping table, the most appropriate cleaning parameters, including cleaning pressure, speed, water volume, and operating mode, are assigned to each path node on the initial path. If significant differences are found between the cleaning parameters assigned to adjacent nodes on the path and exceed safety or performance thresholds, transition nodes are automatically inserted. Through smooth parameter transformation, the continuity, stability, and efficiency of the robot's motion and operation are guaranteed, generating a fully parameterized cleaning trajectory. To ensure that the generated path not only meets the operational requirements of intelligent cleaning but also adapts to the mechanical structure and dynamic limits of the window cleaning robot, the parameterized cleaning trajectory is matched and simulated with the robot's physical constraint parameters one by one to check whether the trajectory meets the physical conditions such as maximum motion speed, acceleration, battery life, and real-time response.At the same time, for trajectory points in the side wheel collision risk area, high-risk falling area or other inaccessible areas, the system actively inserts safety avoidance nodes, adjusts the cleaning path to keep it away from or bypass the dangerous area, or triggers path adjustment in real time when temporary risk signals are detected, and finally outputs a safe and clean trajectory.

[0024] In the embodiment of the present application, through multispectral imaging and dynamic weighted hybrid attention processing, the method of the present invention can accurately distinguish various types of dirt such as water stains, fingerprints, oil stains, dust and mixed stains, especially the recognition effect of transparent or translucent dirt is significant, overcoming the limitations of traditional RGB image processing methods. By integrating the edge wheel collision signal and the fall detection signal, a window boundary safety map is constructed to achieve a deep integration of dirt recognition and safety assurance, effectively avoiding collision or fall accidents of the window cleaning robot during operation. The cleaning strategy mapping table established based on the dirt segmentation results and the edge safety area identification can formulate personalized cleaning parameters for dirt of different types and locations to achieve precise cleaning. Through the dirt density heat map and the improved RRT* algorithm, the method of the present invention realizes the intelligent planning of cleaning paths under safety constraints, dynamically adjusts the path density and cleaning parameters, avoids dangerous areas and gives priority to key dirty areas. The use of multispectral imaging technology and professional image preprocessing enables the system to obtain stable spectral responses under various lighting conditions, ensuring the stability and reliability of dirt recognition in different environments. Based on precise identification and path optimization of dirt distribution characteristics, the window cleaning robot can concentrate its limited battery capacity and cleaning resources on areas that are truly needed, reducing unnecessary cleaning operations and extending single operation time.

[0025] In a specific embodiment, the process of executing step 100 may specifically include the following steps: The multispectral imaging system installed on the window cleaning robot collects window surface images including visible spectrum, near infrared spectrum and ultraviolet spectrum to obtain a multispectral image; Denoising the multispectral image to obtain noise-suppressed image data, and performing radiation correction on the noise-suppressed image data to obtain a corrected multispectral image; The phase correlation registration algorithm is used to perform sub-pixel geometric alignment on each band in the corrected multispectral image to obtain a preprocessed image. The side wheel collision signal is obtained from the touch button on the side wheel socket of the window cleaning robot, and the drop detection signal is obtained from the drop rod and photoelectric switch. A set of marking points of the window boundary and the fall risk position is established to obtain a window boundary safety map.

[0026] Specifically, a multispectral imaging system, mounted on the robot, collects multi-channel information from the target window surface. The system integrates a high-resolution imaging unit for the visible spectrum, along with specialized imaging devices for the near-infrared and ultraviolet spectrums. The visible spectrum primarily operates in the 400-700 nanometer range, the near-infrared spectrum covers the 700-1000 nanometer range, and the ultraviolet spectrum spans the shortwave range of 320-400 nanometers. Each imaging unit utilizes a high-sensitivity, low-noise imaging chip. Through a rational filter design and spectroscopic structure, the system simultaneously captures signals reflected or transmitted from the window surface across different wavelengths. To enhance the system's anti-interference and adaptive illumination capabilities, the imaging system is equipped with multiple LED arrays for each wavelength. These light sources utilize pulse-width modulation technology to dynamically adjust their intensity and illumination cycle to accommodate uneven illumination caused by complex weather conditions, time of day, and variations in glass material. This ensures that the acquired multispectral images exhibit a high signal-to-noise ratio, a wide dynamic range, and uniform brightness distribution. De-noising is then performed on the raw multispectral image data. Image denoising algorithms based on non-local mean filtering, bilateral filtering, or adaptive denoising convolutional networks are used. During processing, statistical analysis is performed on local pixel blocks of the image to identify and separate high-frequency noise from structural and texture information. For areas of high noise density, the algorithm adaptively enlarges the filter window and applies weighted averaging or feature autoregressive noise reduction to the pixel values. For image edges with detailed details, a smaller window and low-weight intervention is used to maximize the preservation of effective boundary features. Throughout the denoising process, dynamic noise estimation, edge preservation, and structural prior mechanisms are employed to produce a set of clear, multi-band images with excellent noise suppression, rich detail, and no noticeable artifacts. Radiometric correction is performed on the multispectral images using a lookup table method. The response characteristics of each band and pixel group are pre-calibrated using radiometric benchmarks such as standard whiteboards and gray cards. Dynamic compensation is then applied based on factors such as actual ambient brightness, exposure parameters, and sensor temperature drift. Each image undergoes pixel-level gain correction, brightness normalization, and color restoration, ensuring that the corrected images across different bands maintain a high degree of consistency in reflecting the actual physical properties of the window surface. The corrected multispectral images are geometrically aligned. Using a phase correlation registration algorithm, the reference band (such as the visible light green band) and the remaining bands are subjected to fast Fourier transforms to calculate their phase correlation in the frequency domain. By locating the peak position of the cross-correlation, the pixel-level translation and sub-pixel offset parameters are obtained, and further expanded to complex transformation models such as affine and perspective based on the actual image structure. After this registration process, images of all multispectral bands are aligned at the sub-pixel level in the spatial coordinate system. At the same time, when performing collection and cleaning tasks, the window cleaning robot uses a variety of safety sensors integrated into its body structure to synchronously perceive its relationship with the physical boundaries of the window and high-risk areas.A touch button on the wheel hub quickly generates a trigger signal when the robot reaches the window edge and the wheels make physical contact. This signal, transmitted by the microprocessor, instantly locates the current spatial point as the window's structural boundary marker. The drop bar, a passive boundary sensing device, triggers a photoelectric switch when the robot tilts, shifts its center of gravity, or is about to exceed the glass support surface, generating a fall risk signal. This signal then identifies the location as a high-risk, dangerous point requiring avoidance or special action. After each signal trigger, the system automatically records these spatial coordinates into a global map of the robot's workspace, creating a set of two-dimensional or three-dimensional spatial marker points encompassing all wheel collision points and high-risk fall points. Using interpolation, boundary connection, and topology optimization algorithms, the system infers the actual window boundary curve and all potential danger zones. All marker points and boundary information are aligned and fused with multispectral image data through unified coordinate mapping, forming a window boundary safety map compatible with visual perception and safety control.

[0027] In a specific embodiment, the process of executing step 200 may specifically include the following steps: The preprocessed image is input into the spatial attention processing module to perform spatial context feature extraction of different receptive fields to obtain the spatial feature data map of the window surface; Perform channel attention analysis on the preprocessed image to obtain the channel feature data map of the window surface, and perform spectral attention processing on the multispectral bands of the preprocessed image to obtain the spectral feature data map of the window surface; Perform edge area feature extraction and edge safety mapping analysis on the window boundary safety map to obtain the window edge safety constraint data map; The pre-processed image and the window boundary safety map are input into the conditional gating unit for feature fusion analysis to obtain the normalized weight coefficient; A weighted fusion calculation is performed on the window surface spatial feature data map, the window surface channel feature data map, the window surface spectral feature data map and the window edge safety constraint data map according to the normalized weight coefficient to obtain an edge safety constraint feature map; The edge safety constraint feature map is subjected to security-aware convolutional multi-scale strip pooling processing to obtain an enhanced feature map.

[0028] Specifically, the preprocessed image is input into the spatial attention processing module, which extracts spatial contextual features at different receptive fields. Based on a deep neural network, this module designs a multi-scale dilated convolutional architecture and parallel receptive field branches, each employing a different combination of convolution kernel sizes and dilation ratios. By merging and integrating the feature responses of the same input image at different spatial scales, a spatial feature map of the window surface is generated, which contains rich spatial structure, edge morphology, and the distribution of fine and fine dirt areas. Simultaneously, channel attention analysis is performed on the channel characteristics of the preprocessed image. By introducing a compression-excitation network structure, the algorithm uses a global average pooling operation to compress the information of each spatial channel into a global feature vector. Subsequently, the importance of each channel is adaptively modeled through multiple layers of fully connected neurons and nonlinear activation functions. The algorithm outputs a set of normalized channel weight coefficients that enhance the response to specific dirt-sensitive channels (e.g., bands with high reflectivity for oil or channels with low reflectivity for dust). This generates a channel feature map of the window surface that highlights discriminative information between different bands and emphasizes key physical properties. Taking into account the band redundancy and cross-complementation introduced by multispectral imaging systems, the algorithm performs spectral attention processing on the preprocessed image. A self-attention mechanism is used to model the feature associations of the data stream for each spectral band. By calculating the correlation matrix between different bands, it is possible to identify which band combinations have a synergistic amplification effect on certain types of stains and which bands introduce noise or redundancy. Through the allocation and dynamic weighting of spectral attention, a set of spectral feature data maps for the window surface is generated. Simultaneously, the window boundary safety map is input into the edge region feature extraction and safety mapping analysis module, focusing on window boundaries and safety constraints. This module uses a deep separable convolutional network to extract features such as the spatial structure, boundary gradient, and collision risk distribution of the edge region. It then combines fall detection signals with high-risk points marked by light touch signals to generate a set of edge safety response maps. During the safety mapping analysis phase, the algorithm automatically constructs a window edge safety constraint data map based on spatial topology, risk level, and window structural patterns. The preprocessed image and window boundary safety map are input into a conditional gating unit for feature fusion analysis. The conditional gating unit, composed of a lightweight convolutional neural network and an adaptive normalization layer, intelligently analyzes the importance of features in each dimension under multiple scene inputs and varying risk conditions. Through end-to-end training, it automatically learns the optimal contribution ratio of spatial, channel, spectral, and safety features in complex window scenes. Its core output is a set of normalized weight coefficients, reflecting the weighted relationship between the four feature types in the final fusion process under the current environment. The normalized weight coefficients are used to perform a weighted fusion of the spatial, channel, spectral, and edge safety constraint data maps. The four feature maps are multiplied by their corresponding weight coefficients and then summed to produce the edge safety constraint feature map. The edge safety constraint feature map is then subjected to security-aware convolutional multi-scale strip pooling.This processing flow incorporates multiple convolutional branches, including horizontal, vertical, diagonal, and point-wise convolutions, and incorporates security masks to constrain features in high-risk areas. After each branch outputs features, strategies such as maximum pooling, average pooling, and learnable attenuated pooling are employed to achieve comprehensive adaptive adjustment of feature scale, orientation, and intensity. The pooled features from all branches are integrated through a security-enhancing feature fusion network, resulting in an enhanced feature map with high directional sensitivity, scale robustness, and security adaptability.

[0029] In a specific embodiment, the step of performing security-aware convolutional multi-scale strip pooling processing on the edge safety constraint feature map to obtain an enhanced feature map may specifically include the following steps: The edge safety constraint feature map is input into the horizontal strip convolution branch for 1×7 convolution kernel processing to obtain a horizontal feature map; the edge safety constraint feature map is input into the vertical strip convolution branch for 7×1 convolution kernel processing to obtain a vertical feature map; the edge safety constraint feature map is input into the diagonal strip convolution branch for diagonal 7×7 convolution kernel processing in two directions to obtain a diagonal feature map; the edge safety constraint feature map is input into the point convolution branch for 1×1 convolution kernel processing to obtain a point feature map; A safety mask is generated based on the window cleaning robot's wheel collision signal and fall detection signal. The safety mask is then convolved with the horizontal feature map, vertical feature map, diagonal feature map, and point feature map to obtain multi-directional feature data under safety constraints. The multi-directional feature data under security constraints are subjected to attenuated pooling processing and integrated through a security enhancement feature fusion network to obtain an enhanced feature map.

[0030] Specifically, the edge safety constraint feature map is input into the multi-branch convolution structure to accurately characterize the spatial characteristics of various directional, strip-shaped, and point-shaped stains on the window surface, while simultaneously suppressing and responding to operational risks at the boundaries and safety areas in real time. There are horizontal strip convolution branches, vertical strip convolution branches, diagonal strip convolution branches, and point convolution branches. The horizontal strip convolution branch uses a 1×7 convolution kernel to perform one-dimensional horizontal weight sliding and feature extraction on the feature map. This structure is suitable for capturing horizontally distributed water stains, drag marks, or wipe marks on the glass, and enhances the robot's sensitivity and response accuracy to horizontal strip stains. The vertical strip convolution branch uses a 7×1 convolution kernel to specifically extract texture and stain changes in the vertical direction. Its goal is to accurately describe the dirt characteristics distributed along the height of the glass, such as rain marks and gravity drag residues, so that the robot can take more targeted and efficient cleaning actions for stains in these directions. The diagonal strip convolution branch, centered around two orthogonal 7×7 diagonal convolution kernels, extracts strip features in the +45° and -45° diagonal directions, respectively. This helps detect and enhance contamination in non-orthogonal directions, such as oblique scratches, tilted fingerprints, and residue from special operations. This enables the robot to clean along its normal main axis and precisely identify and address stains on irregular strips and complex glass structures. The point convolution branch, using a 1×1 convolution kernel, completes point-by-point mapping of the feature map. Its primary task is to preserve and enhance information about isolated local stains, tiny stains, or areas of spatial mutation, preventing loss of detail or "blurred" suppression of isolated, important stains during large-scale strip convolution. After feature extraction in all directions is complete, a safety mask is dynamically generated based on safety signals collected by the touch button on the side wheel socket and the photoelectric switch on the drop bar during actual operation. The safety mask is essentially a binary or probability distribution map of the same spatial size as the feature map. Areas corresponding to side-wheel collision signals and high-fall-risk boundaries are assigned low or even zero weights, while central, risk-free safety zones retain high weights. The safety mask is fused with the extracted horizontal, vertical, diagonal, and point-wise feature maps through channel-by-channel convolution or element-wise multiplication. This applies a safety weight to the directional feature values of each pixel. Feature expressions in high-risk areas are suppressed or masked, while features in safe areas are fully preserved or even enhanced, improving the environmental adaptability and safety perception capabilities of the enhanced feature map. Attenuated pooling is performed on multi-directional feature data under safety constraints. In addition to traditional maximum and average pooling, learnable attenuation weights are added to enable adaptive weighting based on local feature strength, directional distribution, and risk regions. For example, the maximum response is amplified in salient feature areas and safe regions, while the average response is adaptively attenuated in high-risk regions to suppress noise and redundant responses. The pooling results are fed into the safety-enhancing feature fusion network.This network integrates pooled features from various directions and safety mask weights. Through fully connected or multi-branch fusion within the network structure, it achieves deep reconstruction and weighted combination of features across different directions, spatial scales, and safety levels. The network's fusion module incorporates a conditional gating unit that automatically adjusts the fusion path and weight coefficients based on the input safety mask and spatial distribution. This enables the system to intelligently enhance the representation of features in critical areas (such as large oil and water stains) in various operational scenarios, including window centers, edges, and corners, while simultaneously suppressing and avoiding invalid or dangerous responses at high-risk boundaries. The safety-enhancing feature fusion network outputs an enhanced feature map.

[0031] In a specific embodiment, the process of executing step 300 may specifically include the following steps: The enhanced feature map is input into the encoder of the dirt segmentation network for multi-layer feature extraction to obtain a multi-scale encoding feature map, and the multi-scale encoding feature map is subjected to transposed convolution upsampling processing to obtain a multi-level decoding feature map; Based on the wheel collision signal and fall detection signal of the window cleaning robot, the edge response map is generated from the multi-scale encoding feature map to obtain the edge response feature. Based on the edge wheel collision signal and fall detection signal, the multi-scale encoding feature map is used to conduct safety risk assessment and obtain the edge safety constraint feature; Adaptively fuse the edge response features and edge safety constraint features with the multi-level decoding feature map to obtain the decoding results, and generate the dirt segmentation results and edge safety area identification based on the decoding results; Based on the dirt segmentation results and edge safety area identification, the correspondence between dirt characteristics and cleaning parameters of the window cleaning robot is established, and a cleaning strategy mapping table for dirt in different locations is obtained.

[0032] Specifically, the enhanced feature map is input into the encoder of the dirt segmentation network. The encoder utilizes a deep convolutional neural network. Internally, it uses multi-level convolution, pooling, and batch normalization operations to spatially compress and semantically abstract the input feature map, extracting a multi-scale encoded feature map that covers both global context and detailed features. The multi-scale encoded feature map is then upsampled using transposed convolutions, and the encoded feature map is upscaled layer by layer, gradually restoring its spatial resolution to the original input size. Each upsampling stage is accompanied by skip connections, fusing the high-resolution shallow features from the encoding stage with the deep features from the decoding stage. Through feature concatenation or weighted summation, cross-level multimodal information integration is achieved, resulting in a multi-level decoded feature map. Based on this, the window cleaning robot utilizes its integrated safety perception system—including the touch button on the side wheel socket and the photoelectric switch on the drop bar—to perform a specialized edge response map generation operation on the multi-scale encoded feature map. By applying Sobel, Laplacian, or learning-based edge detection operators to the encoded feature map, the spatial gradient responses of key locations, such as the window's physical boundaries, glass corners, support surfaces, and overhanging edges, are extracted. These edge response features significantly improve the segmentation network's accuracy in complex backgrounds and high-risk areas, reducing cross-boundary missegments and omissions. Simultaneously, based on the same safety sensor signals, a safety risk assessment is performed on the multi-scale encoded feature map. By spatially mapping collision and fall signals to the feature map coordinate system, an edge safety constraint feature map is automatically constructed, clearly defining high-risk operating areas requiring avoidance or special handling. These safety constraint features serve as spatial prior inputs in the segmentation stage, playing a global role in subsequent parameter decision-making and trajectory optimization, improving the robot's safety level in high-rise and complex window scenes. The generated edge response and edge safety constraint features are adaptively fused with the multi-level decoded feature map obtained through decoding. Using spatially adaptive normalization, a conditional gated attention mechanism, or a multi-head fusion network, the feature fusion weights of the segmentation response for each spatial location and category are dynamically adjusted based on the actual boundary signal and safety level. In areas near window edges or where risk signals are detected, the fused features automatically increase the contribution of safety constraint features, while in large, risk-free areas, spatial and channel features are emphasized, achieving an optimal dynamic balance between segmentation accuracy and safety robustness for the entire image. The decoded output of the fusion generates a contamination classification probability for each pixel and simultaneously outputs multi-dimensional attributes such as the spatial safety risk level and boundary response strength. Based on the decoded contamination segmentation results and edge safety zone identification, a mapping process is constructed to map contamination characteristics to the cleaning parameters of the window cleaning robot.Connected domain analysis and instance labeling are performed on the segmented image to extract basic morphological features of each stain instance, such as spatial location, geometric dimensions, main orientation, and compactness. Multispectral data is then combined to calculate statistical features of the stain's material, such as the mean, standard deviation, and normalized difference of reflectance across each band. This information is then deeply integrated into a target stain feature vector based on its spatial distance from the edge safety zone, risk level, and dynamic safety status. This feature vector is then fed into the edge safety perception system's fuzzy decision tree, neural network, or parameter mapping rule base. Combining knowledge rules and model reasoning capabilities, the system automatically infers the optimal cleaning strategy based on multi-dimensional conditions such as stain type, location, and risk level. This involves assigning an appropriate cleaning mode, pressure, speed, water volume, and path adjustment method to each stain instance. For stains near the edge or risk zone, the system proactively assigns low speed, low pressure, or avoidance strategies to prevent robot falls or collisions due to misoperation. For common stains in the central safety zone, efficient cleaning parameters are used to improve efficiency. All inference and assignment results are dynamically organized into a cleaning strategy mapping table.

[0033] In a specific embodiment, the step of establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot based on the dirt segmentation result and the edge safety area identification to obtain a cleaning strategy mapping table for dirt at different locations may specifically include the following steps: Performing connected domain analysis on the dirt segmentation results to obtain a dirt instance set, and performing morphological analysis on each dirt instance in the dirt instance set to obtain basic geometric feature data; Performing statistical feature extraction on each dirt instance in the dirt instance set to obtain material feature data; Based on the spatial relationship between each dirt instance in the dirt instance set and the edge safety area identifier, the distance to the window edge, the distance to the fall risk area, and the edge risk value are calculated to obtain safety feature data; Perform deep feature fusion on basic geometric feature data, material feature data and safety feature data to obtain the target dirt feature vector; The target dirt feature vector is input into the edge safety perception fuzzy decision tree for multi-level reasoning analysis to obtain a cleaning strategy mapping table for dirt at different locations.

[0034] Specifically, a pixel-level connected domain algorithm (such as four-connected, eight-connected, or flood-fill-based region growing) is used to group all spatially adjacent pixels predicted to be of the same category in the segmentation results into independent stain instances, thereby constructing a set of stain instances. Each stain instance corresponds to an actual contaminated area on the window surface. For each stain instance, basic geometric feature data is extracted using morphological analysis methods. The total number of pixels is counted to calculate the actual area of the stain instance. Edge detection and contour tracking algorithms are also used to obtain geometric information such as perimeter, principal direction, minimum bounding rectangle, and minimum bounding ellipse. Metrics such as aspect ratio, compactness, roundness, and aspect ratio are calculated to further analyze the stain's morphology, including its point, strip, diffuse, or sheet-like structure. For highly directional stains, parameters such as principal axis angle, endpoint coordinates, and center of mass are also recorded. This information helps the system determine whether the stain is easily covered by conventional cleaning patterns or requires special path adjustments or repeated cleaning. In addition to morphological analysis, statistical feature extraction is performed on each stain instance to obtain rich data reflecting its physical properties and material characteristics. By statistically analyzing the multispectral pixel values within the instance mask area, the mean, standard deviation, maximum and minimum values, skewness, and kurtosis are calculated for each band (e.g., visible light, near-infrared, and ultraviolet), reflecting the stain's color, thickness, optical density, and other properties. For multi-band data, inter-band ratio features, normalized differences, principal component analysis features, or color space transformation features are extracted to determine the stain's material category, such as high-reflectivity and low-absorption (e.g., water stains), high-absorption and low-reflectivity (e.g., oil film), dispersed particles (e.g., dust), or high-light diffuse reflectivity (e.g., fingerprints). Furthermore, considering the actual operational safety and operational risks of the window cleaning robot, each stain instance is spatially associated with an edge safety zone marker. Using spatial distance and risk assessment algorithms, the shortest distance from each stain to the window's physical edge, the closest distance to the fall risk zone, and a comprehensive edge risk value are calculated. Typically, a spatial distance map is created for the entire image using an edge safety zone mask and a distance transform (such as Euclidean distance transform or Chebyshev distance). The system then iterates through the centroid, boundary points, or principal axis endpoints of each instance, calculating the minimum distance from these points to the safety boundary and risk zone as a safety feature. Stains located near the boundary and adjacent to the risk zone are assigned a higher risk weight. By weightedly fusing distance values and spatial distribution characteristics, the system outputs comprehensive safety feature data reflecting the difficulty of stain detection, potential risk level, and operational safety constraints. Deep feature fusion is performed on basic geometric, material, and safety feature data to generate a target stain feature vector. During the decision-making and inference phase, the target stain feature vector is input into an edge safety-aware fuzzy decision tree. This decision tree incorporates conventional decision rules based on conditions such as morphology and material, and introduces multi-level safety constraint nodes such as spatial risk and boundary distance to enable reasoning under multiple conditions and weights.The decision tree uses a layered approach, first determining the soil type (stripes, spots, diffuse, or special), then assigning a basic cleaning mode based on material properties. It then adjusts key parameters such as cleaning force, pressure, and speed based on spatial location, boundary distance, and risk level. In extreme cases, it automatically inserts obstacle avoidance maneuvers, detours, or temporarily skips high-risk areas. Fuzzy logic allows the system to implement soft decisions and interval responses for special cases such as fuzzy features, unclear boundaries, and complex stains, avoiding rigid misjudgments caused by parameter limits or sudden environmental changes. All inference and allocation results are dynamically organized into a cleaning strategy mapping table. This mapping table specifies the spatial location, morphological category, material properties, safety level, and optimal cleaning parameter combination for each soiling instance, including the cleaning mode (standard / slow / high-frequency / low-pressure), the operation path (conventional / boundary avoidance / multiple repetitions / point-based focus), cleaning force and speed, water distribution, and special action requirements (such as automatic shutdown or alarm when operating in high-risk areas).

[0035] In a specific embodiment, the process of executing step 400 may specifically include the following steps: A weighted fusion calculation is performed based on the different dirt types, dirt severity, and regional importance in the dirt segmentation results. The window cleaning robot's side wheel collision area and fall risk area are set as safety constraint areas to obtain a dirt density heat map. Initialize the RRT* algorithm based on the dirt density heat map, and set a heat map-guided sampling strategy where the dirt density is proportional to the sampling probability. The expansion step size of the RRT* algorithm is dynamically adjusted according to the current node's dirt density value and the distance to the safety constraint area to obtain the node expansion sequence; Perform local replanning optimization based on heat map guided sampling strategy and node expansion sequence to obtain the initial clean path; According to the cleaning strategy mapping table, cleaning parameters are assigned to each node on the initial cleaning path, and transition nodes are inserted when the difference in parameters of adjacent nodes exceeds a set threshold to obtain a parameterized cleaning trajectory; The parameterized cleaning trajectory is matched and verified with the physical constraint parameters of the window cleaning robot, and safety avoidance nodes are added in the collision risk area and the fall risk area to obtain a safe cleaning trajectory.

[0036] Specifically, based on the stain segmentation results obtained by the front-end segmentation network, a comprehensive assessment and weighted fusion of the importance of various stain types, severity levels, and spatial regions on the window surface is performed. Based on each category's segmentation mask, preset stain type weights are used to assign differentiated priorities to stains such as water stains, oil stains, fingerprints, and dust, reflecting their actual impact on glass transparency, visual aesthetics, and cleaning difficulty. The cleaning urgency of each region is calculated by combining stain severity scores (e.g., area threshold, pixel density, reflective intensity, or color difference indicators). This is further weighted based on regional importance. For example, the central field of view, critical lighting areas, or high-touch areas are given higher priority, while edges and corners are given lower priority. This multi-weighted fusion generates a continuous stain density heat map in the two-dimensional window space. Each pixel value in the heat map comprehensively reflects the cleaning priority, stain complexity, and spatial impact weight of the region. Simultaneously, the side wheel collision signals and fall detection signals collected in real time by the window cleaning robot are spatially mapped into safety constraint areas. These areas are assigned extremely low or even zero priorities on the dirt density heat map, and are explicitly considered high-risk areas that are inaccessible or require special actions to avoid during the actual sampling and path generation stages. The RRT* algorithm is used to initialize path planning based on the heat map. RRT* is an efficient random sampling tree path search algorithm that is suitable for high-dimensional unstructured path planning in continuous space. In order to achieve efficient sampling in intelligent cleaning scenarios, a heat map-guided sampling strategy is introduced in the RRT* initialization stage, and the sampling probability of each spatial point is directly linked to its corresponding dirt density value, which greatly increases the probability of high-priority and high-density dirty areas being sampled, allowing the path trunk to actively pass through key cleaning areas to achieve optimal allocation and coverage of work resources. Compared with uniform sampling or random sampling, this guided sampling method improves the path convergence speed and spatial coverage uniformity, while also taking into account the cleaning effect of high-value work areas and significantly reducing sampling redundancy in low-priority or dangerous areas. During the node expansion phase of the RRT* tree structure, the system adjusts sampling probabilities based on heatmap weights and dynamically adjusts the expansion step size based on the current node's dirt density and its spatial distance to the nearest safety constraint zone. When a node is located in a high-density dirt zone or far from the safety boundary, the expansion step size is set smaller to ensure that the path precisely covers difficult areas with concentrated dirt, achieving high-density grid cleaning. When a node approaches a safety constraint zone or a low-priority area, the expansion step size is automatically increased, accelerating the path through irrelevant or low-value areas, shortening the overall operation path, and improving the robot's energy efficiency. As the RRT* path continues to grow, the system periodically performs local replanning and optimal optimization on the current path sequence. Using algorithms such as neighborhood reconnection, path smoothing, and redundant branch pruning, it continuously eliminates unnecessary loops and dead ends, forming an initial high-priority coverage cleaning path. This path avoids safety risk areas to the greatest extent possible, covers all major dirt distributions, and maintains optimal length and execution efficiency in space.The initial cleaning path is parameterized and associated with a cleaning strategy mapping table. Each node along the path is traversed, and the optimal cleaning parameters, including cleaning force, speed, water volume, and operation mode, are searched or inferred from the mapping table based on characteristics such as the type and severity of dirt in the space where the node is located. If cleaning parameter differences between adjacent nodes on the path exceed a set threshold, such as when switching from high-pressure cleaning mode to low-pressure and slow speed, or from wet wiping to dry wiping, the system automatically inserts transition nodes into the path to create smooth parameter transition sections. This ensures the continuity of the robot's motion and physical execution during operation, as well as the safety of its mechanical structure, and avoids mechanical shock, operational vibration, and even accidents caused by sudden parameter changes. After path parameterization, the final cleaning trajectory is matched against the window cleaning robot's physical constraints (such as maximum speed, maximum acceleration, operating radius, motion limits, and remaining battery power) and verified through simulation. This ensures that in actual operation, there will be no interruptions or uncontrolled risks due to trajectory violations, mechanical overload, or abnormal energy consumption. The system automatically checks the spatial relationship between all path nodes and the safety constraint area one by one, and inserts safety avoidance nodes into the trajectory segments that pass through or approach collision risk areas or high-risk areas for falling. These nodes can instruct the robot to slow down, limit pressure, switch to obstacle avoidance mode, or temporarily suspend the operation, and continue the cleaning action after manual intervention or risk elimination. After all the above optimizations and safety checks are completed, a highly safe, parameter-adaptive, fully covers key stains, and meets all the physical constraints of the window cleaning robot. The trajectory is sent to the robot execution unit in the form of a spatiotemporal sequence and parameter sequence, and is dynamically re-planned according to environmental changes (such as new collision signals and real-time updates of dirt) during the actual operation to form a closed-loop feedback.

[0037] The surface dirt image recognition method used in the window cleaning robot in the embodiment of the present application is described above. The surface dirt image recognition device 10 used in the window cleaning robot in the embodiment of the present application is described below. Figure 2 In one embodiment of the present application, a surface dirt image recognition device 10 used in a window cleaning robot includes: A preprocessing module 11 is used to preprocess the multispectral images collected by the window cleaning robot to obtain a preprocessed image, and simultaneously integrate the side wheel collision signal and fall detection signal of the window cleaning robot to obtain a window boundary safety map; A convolution processing module 12 is used to perform edge-constrained dynamic weighted hybrid attention processing and convolution multi-scale strip pooling processing on the pre-processed image and the window boundary safety map to obtain an enhanced feature map; Establishing module 13, for performing edge safety perception dirt segmentation on the enhanced feature map, obtaining dirt segmentation results and edge safety area identification, and establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot to obtain a cleaning strategy mapping table for dirt at different locations; The generation module 14 is used to generate a dirt density heat map based on the dirt segmentation result and the cleaning strategy mapping table, and execute a path planning algorithm to generate a safe cleaning trajectory.

[0038] By synergizing the aforementioned components, using multispectral imaging and dynamic weighted hybrid attention processing, the proposed method can accurately distinguish various soil types, including water stains, fingerprints, oil stains, dust, and mixed stains. It is particularly effective in identifying transparent or translucent stains, overcoming the limitations of traditional RGB image processing methods. By integrating edge wheel collision signals and fall detection signals to construct a window boundary safety map, the system achieves a deep integration of soil identification and safety assurance, effectively preventing collisions and falls during window cleaning robot operations. A cleaning strategy mapping table, based on soil segmentation results and edge safety zone identification, enables customized cleaning parameters to be customized for different soil types and locations, achieving precise cleaning. Using soil density heat maps and an improved RRT* algorithm, the proposed method achieves intelligent cleaning path planning under safety constraints, dynamically adjusting path density and cleaning parameters to avoid hazardous areas while prioritizing key soiled areas. The use of multispectral imaging technology and specialized image preprocessing enables the system to obtain stable spectral responses under various lighting conditions, ensuring the stability and reliability of soil identification in diverse environments. Based on precise identification and path optimization of dirt distribution characteristics, the window cleaning robot can concentrate its limited battery capacity and cleaning resources on areas that are truly needed, reducing unnecessary cleaning operations and extending single operation time.

[0039] See also Figure 3 , Figure 3 This is a schematic block diagram of the structure of an electronic device 300 provided in an embodiment of the present application. The electronic device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected via a device bus 303, wherein the memory 302 may include a non-volatile storage medium and an internal memory.

[0040] The non-volatile storage medium may store a computer program, which includes program instructions. When the program instructions are executed by the processor 301, the processor 301 may execute any of the above-mentioned surface dirt image recognition methods for the window cleaning robot.

[0041] The processor 301 is used to provide computing and control capabilities to support the operation of the entire electronic device 300 .

[0042] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can execute any of the above-mentioned surface dirt image recognition methods used in the window cleaning robot.

[0043] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of the structure related to the solution of the present application and does not constitute a limitation on the electronic device 300 involved in the solution of the present application. The specific electronic device 300 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0044] It should be understood that the processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0045] It should be noted that, those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the electronic device 300 described above can refer to the corresponding process of the surface dirt image recognition method used in the window cleaning robot, and will not be repeated here.

[0046] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the one or more processors implement the surface dirt image recognition method for a window cleaning robot as provided in an embodiment of the present application.

[0047] The computer-readable storage medium may be an internal storage unit of the electronic device 300 in the aforementioned embodiment, such as a hard disk or memory of the electronic device 300. The computer-readable storage medium may also be an external storage device of the electronic device 300, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc., equipped with the electronic device 300.

[0048] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0049] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or contributes to the existing technology or the entire technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0050] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A surface dirt image recognition method used in a window cleaning robot, characterized in that: include: The multispectral images collected by the window cleaning robot are preprocessed to obtain a preprocessed image. At the same time, the window cleaning robot's wheel collision signal and fall detection signal are integrated to obtain a window boundary safety map. Performing edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing on the preprocessed image and the window boundary safety map to obtain an enhanced feature map; Performing edge safety perception dirt segmentation on the enhanced feature map to obtain dirt segmentation results and edge safety area identification, and establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot to obtain a cleaning strategy mapping table for dirt at different locations; A dirt density heat map is generated based on the dirt segmentation result and the cleaning strategy mapping table, and a path planning algorithm is executed to generate a safe cleaning trajectory.

2. The surface dirt image recognition method for a window cleaning robot according to claim 1, characterized in that: The multispectral image collected by the window cleaning robot is preprocessed to obtain a preprocessed image, and the side wheel collision signal and fall detection signal of the window cleaning robot are integrated to obtain a window boundary safety map, including: The multispectral imaging system installed on the window cleaning robot collects window surface images including visible spectrum, near infrared spectrum and ultraviolet spectrum to obtain a multispectral image; Performing denoising on the multispectral image to obtain noise-suppressed image data, and performing radiation correction on the noise-suppressed image data to obtain a corrected multispectral image; Using a phase correlation registration algorithm to perform sub-pixel geometric alignment on each band in the corrected multispectral image to obtain a preprocessed image; The side wheel collision signal is obtained from the touch button on the side wheel socket of the window cleaning robot, and the drop detection signal is obtained from the drop rod and photoelectric switch. A set of marking points of the window boundary and the fall risk position is established to obtain a window boundary safety map.

3. The surface dirt image recognition method for a window cleaning robot according to claim 1, characterized in that: The pre-processed image and the window boundary safety map are subjected to edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing to obtain an enhanced feature map, including: Inputting the preprocessed image into the spatial attention processing module to perform spatial context feature extraction of different receptive fields to obtain a window surface spatial feature data map; Performing channel attention analysis on the preprocessed image to obtain a window surface channel feature data map, and performing spectral attention processing on the multispectral bands of the preprocessed image to obtain a window surface spectral feature data map; Performing edge area feature extraction and edge safety mapping analysis on the window boundary safety map to obtain a window edge safety constraint data map; Inputting the pre-processed image and the window boundary safety map into a conditional gating unit for feature fusion analysis to obtain a normalized weight coefficient; Performing a weighted fusion calculation on the window surface spatial feature data graph, the window surface channel feature data graph, the window surface spectral feature data graph, and the window edge safety constraint data graph according to the normalized weight coefficient to obtain an edge safety constraint feature graph; A security-aware convolutional multi-scale strip pooling process is performed on the edge security constraint feature map to obtain an enhanced feature map.

4. The surface dirt image recognition method for a window cleaning robot according to claim 3, characterized in that: The performing security-aware convolutional multi-scale strip pooling processing on the edge security constraint feature map to obtain an enhanced feature map includes: Input the edge safety constraint feature map into the horizontal strip convolution branch and perform 1×7 convolution kernel processing to obtain a horizontal feature map; input the edge safety constraint feature map into the vertical strip convolution branch and perform 7×1 convolution kernel processing to obtain a vertical feature map; input the edge safety constraint feature map into the diagonal strip convolution branch and perform 7×7 convolution kernel processing in two directions of diagonal shape to obtain a diagonal feature map; input the edge safety constraint feature map into the point convolution branch and perform 1×1 convolution kernel processing to obtain a point feature map; A safety mask is generated based on the side wheel collision signal and the fall detection signal of the window cleaning robot, and a convolution operation is performed on the safety mask with the horizontal feature map, the vertical feature map, the diagonal feature map, and the point feature map to obtain multi-directional feature data under safety constraints; The multi-directional feature data under the security constraints are subjected to attenuation pooling processing and integrated through a security enhancement feature fusion network to obtain an enhanced feature map.

5. The surface dirt image recognition method for a window cleaning robot according to claim 1, characterized in that: The enhanced feature map is subjected to edge safety perception dirt segmentation to obtain a dirt segmentation result and an edge safety area identifier, and a corresponding relationship between dirt characteristics and cleaning parameters of the window cleaning robot is established to obtain a cleaning strategy mapping table for dirt at different locations, including: Inputting the enhanced feature map into the encoder of the dirt segmentation network for multi-layer feature extraction to obtain a multi-scale encoding feature map, and performing transposed convolution upsampling processing on the multi-scale encoding feature map to obtain a multi-level decoding feature map; Based on the side wheel collision signal and the fall detection signal of the window cleaning robot, an edge response map is generated for the multi-scale coding feature map to obtain an edge response feature; Based on the edge wheel collision signal and the fall detection signal, a safety risk assessment is performed on the multi-scale coding feature map to obtain an edge safety constraint feature; Adaptively fusing the edge response feature and the edge safety constraint feature with the multi-level decoding feature map to obtain a decoding result, and generating a dirt segmentation result and an edge safety area identifier based on the decoding result; A correspondence between dirt characteristics and cleaning parameters of a window cleaning robot is established according to the dirt segmentation result and the edge safety area identifier, and a cleaning strategy mapping table for dirt at different locations is obtained.

6. The surface dirt image recognition method for a window cleaning robot according to claim 5, characterized in that: The corresponding relationship between the dirt characteristics and the cleaning parameters of the window cleaning robot is established according to the dirt segmentation result and the edge safety area identifier to obtain a cleaning strategy mapping table for dirt at different locations, including: Performing a connected domain analysis on the dirt segmentation result to obtain a dirt instance set, and performing a morphological analysis on each dirt instance in the dirt instance set to obtain basic geometric feature data; Performing statistical feature extraction on each dirt instance in the dirt instance set to obtain material feature data; Calculating the distance to the window edge, the distance to the fall risk area, and the edge risk value based on the spatial relationship between each dirt instance in the dirt instance set and the edge safety area identifier to obtain safety feature data; Performing deep feature fusion on the basic geometric feature data, the material feature data, and the security feature data to obtain a target dirt feature vector; The target dirt feature vector is input into the edge safety perception fuzzy decision tree for multi-level reasoning analysis to obtain a cleaning strategy mapping table for dirt at different locations.

7. The surface dirt image recognition method for a window cleaning robot according to claim 1, characterized in that: The step of generating a dirt density heat map based on the dirt segmentation result and the cleaning strategy mapping table, and executing a path planning algorithm to generate a safe cleaning trajectory includes: A weighted fusion calculation is performed based on the different dirt types, dirt severity, and regional importance in the dirt segmentation results, and the side wheel collision area and fall risk area of the window cleaning robot are set as safety constraint areas to obtain a dirt density heat map; Initialize the RRT* algorithm based on the dirt density heat map, and set a heat map-guided sampling strategy in which the dirt density is proportional to the sampling probability; The expansion step size of the RRT* algorithm is dynamically adjusted according to the current node's dirt density value and the distance to the safety constraint area to obtain the node expansion sequence; Performing local replanning optimization based on the heat map guided sampling strategy and the node expansion sequence to obtain an initial cleaning path; Assigning cleaning parameters to each node on the initial cleaning path according to the cleaning strategy mapping table, and inserting transition nodes where the difference in parameters of adjacent nodes exceeds a set threshold to obtain a parameterized cleaning trajectory; The parameterized cleaning trajectory is matched and verified with the physical constraint parameters of the window cleaning robot, and safety avoidance nodes are added in the collision risk area and the fall risk area to obtain a safe cleaning trajectory.

8. A surface dirt image recognition device used in a window cleaning robot, characterized in that: The method for recognizing a surface dirt image applied to a window cleaning robot according to any one of claims 1 to 7 is used, wherein the surface dirt image recognition device applied to the window cleaning robot comprises: The preprocessing module is used to preprocess the multispectral images collected by the window cleaning robot to obtain a preprocessed image, and at the same time integrate the window cleaning robot's side wheel collision signal and fall detection signal to obtain a window boundary safety map; a convolution processing module, configured to perform edge-constrained dynamic weighted hybrid attention processing and convolutional multi-scale strip pooling processing on the pre-processed image and the window boundary safety map to obtain an enhanced feature map; Establish a module for performing edge safety perception dirt segmentation on the enhanced feature map, obtaining dirt segmentation results and edge safety area identification, and establishing a correspondence between dirt characteristics and cleaning parameters of the window cleaning robot to obtain a cleaning strategy mapping table for dirt at different locations; A generation module is used to generate a dirt density heat map based on the dirt segmentation result and the cleaning strategy mapping table, and execute a path planning algorithm to generate a safe cleaning trajectory.

9. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the electronic device to execute the surface dirt image recognition method for a window cleaning robot according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the surface dirt image recognition method for a window cleaning robot according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Weight identification method and device based on large model and medium

    CN120869320A

  • High-voltage switch shell stain detection method and system based on machine vision

    CN120876477A

  • High-voltage switch shell stain detection method and system based on machine vision

    CN120876477B

  • Self-adaptive path planning control system based on environment characteristic perception

    CN121254857A