Window cleaning machine movement control method and related device based on machine vision

Through multispectral image acquisition and superpixel segmentation technology, combined with a bilinear system dynamics model and a soft-constrained visual iterative regulator, the problem of dirt identification and control of traditional window cleaning machines under complex lighting and external interference is solved, and smooth and stable cleaning path planning is achieved.

CN120491657BActive Publication Date: 2025-09-26SHENZHEN YIJIE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510992107.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-26
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Traditional window cleaning machine control methods lack the ability to intelligently perceive the dirt on the window surface, have difficulty identifying slight dirt, and have unstable motion trajectories under complex lighting and external interference.

Method used

Multispectral image acquisition and superpixel segmentation technology are used, combined with a bilinear system dynamics model and a soft-constrained visual iterative regulator to achieve accurate identification and smooth control of dirty areas.

Benefits of technology

It accurately identifies various degrees of dirt under different lighting conditions and maintains the smoothness and stability of the movement trajectory under external interference, improving cleaning efficiency and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491657B_ABST
    Figure CN120491657B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of machine vision technology and discloses a machine vision-based window cleaning machine movement control method and related devices. The method comprises: performing multispectral image acquisition and superpixel segmentation on the window surface to obtain a probability map of dirty areas; performing multispectral feature fusion and cleaning path planning on the probability map of dirty areas to obtain a movement trajectory of the window cleaning machine; establishing a bilinear system dynamics model based on the movement trajectory of the window cleaning machine and designing a visual feedback state observer to obtain real-time state information of the window cleaning machine; inputting the movement trajectory of the window cleaning machine and the real-time state information of the window cleaning machine into a soft-constrained visual iterative regulator for control calculation to obtain a smooth movement control instruction. The present invention solves the problem of detection difficulty caused by the similarity of features between slight dirt and clean surfaces, enables the window cleaning machine to accurately identify various degrees of dirt, and effectively avoids the control jitter problem when constraints are activated in traditional control methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and in particular to a machine vision-based window cleaning machine movement control method and related devices. Background Art

[0002] Traditional window cleaning machine control methods rely primarily on simple preset paths or manual remote control. These methods lack intelligent sensing of surface contamination, often leading to incomplete cleaning and wasted resources. Traditional vision systems face significant challenges when dealing with minor contamination, as it mimics the characteristics of clean surfaces and can be easily overlooked. This prevents the detector from providing sufficiently accurate information about contaminated areas, hindering the window cleaning machine's cleaning path planning.

[0003] The uneven lighting on window surfaces also poses a significant challenge to the vision systems of window cleaning robots. Window surfaces often exhibit strong light reflections, shadowed areas, and light incident at varying angles, causing the same dirt to present distinct visual characteristics under different lighting conditions. Traditional single-spectral imaging methods perform poorly in such complex lighting environments, making it difficult to accurately identify and locate dirty areas, especially on surfaces with high reflectivity, such as glass curtain walls. Furthermore, the control systems of existing window cleaning robots lack robustness to external disturbances (such as wind and changes in surface friction) during movement, resulting in unstable motion trajectories. Summary of the Invention

[0004] The present invention provides a machine vision-based motion control method and related devices for a window cleaning machine. The present invention solves the problem of difficulty in detecting slight dirt due to the similar surface features of clean surfaces, enabling the window cleaning machine to accurately identify various degrees of dirt, while effectively avoiding the control jitter problem when constraints are activated in traditional control methods.

[0005] In a first aspect, the present invention provides a method for controlling the movement of a window cleaning machine based on machine vision, the method comprising:

[0006] Perform multispectral image acquisition and superpixel segmentation on the window surface to obtain a probability map of dirty areas;

[0007] Performing multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain the moving trajectory of the window cleaning machine;

[0008] Based on the movement trajectory of the window cleaning machine, a bilinear system dynamics model is established and a visual feedback state observer is designed to obtain the real-time state information of the window cleaning machine;

[0009] The movement trajectory of the window cleaning machine and the real-time state information of the window cleaning machine are input into a soft-constraint visual iterative regulator for control calculation to obtain a smooth movement control instruction.

[0010] In a second aspect, the present invention provides a machine vision-based mobile control device for a window cleaning machine, the machine vision-based mobile control device for a window cleaning machine comprising:

[0011] The acquisition module is used to collect multispectral images and perform superpixel segmentation on the window surface to obtain a probability map of dirty areas;

[0012] A planning module, configured to perform multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain a movement trajectory of the window cleaning machine;

[0013] An observation module is used to establish a bilinear system dynamics model based on the movement trajectory of the window cleaning machine and design a visual feedback state observer to obtain real-time state information of the window cleaning machine;

[0014] The control calculation module is used to input the movement trajectory of the window cleaning machine and the real-time status information of the window cleaning machine into the soft constraint visual iterative regulator to perform control calculation and obtain a smooth movement control instruction.

[0015] A third aspect of the present invention provides a computer device comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned machine vision-based window cleaning machine movement control method.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned machine vision-based window cleaning machine movement control method.

[0017] In the technical solution provided by the present invention, through multispectral image acquisition and regional feature extraction technology based on superpixel networks, it is possible to effectively capture the subtle differences between dirty areas and clean areas, solving the problem of detection difficulties caused by the similarity of features between slight dirt and clean surfaces, and enabling the window cleaning machine to accurately identify various degrees of dirt. A multispectral imaging mechanism based on regional attention fusion is adopted, combining information from the two spectral domains of visible light and near-infrared, to perform adaptive feature fusion on the window surface under different lighting conditions, effectively solving the problem of dirt detection under complex lighting conditions such as strong light reflection and shadow areas. Based on the enhanced regional feature map, dirty area target segmentation and cleaning path planning are performed, and Z-shaped or spiral-shaped refined paths are automatically generated for large dirty areas. At the same time, the path continuity is ensured through B-spline curve smoothing, achieving efficient cleaning coverage of various complex dirt distributions. A bilinear system dynamics model is used to accurately capture the nonlinear dynamic behavior of the window cleaning machine, and a high-order sliding mode disturbance observer is combined to compensate for external disturbances, significantly improving the accuracy of state estimation under interference conditions such as wind force and surface friction changes. The soft-constrained visual iterative regulator of the present invention converts hard constraints into soft constraints by introducing slack variables and obstacle functions to process constraints, effectively avoiding the control jitter problem when constraints are activated in traditional control methods, and generating smooth and continuous control inputs even under conditions of external interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 Schematic diagram of the steps of a method for controlling the movement of a window cleaning machine based on machine vision in an embodiment of the present invention;

[0020] Figure 2 Schematic diagram of the structure of a mobile control device for a window cleaning machine based on machine vision in an embodiment of the present invention;

[0021] Figure 3 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] Embodiments of the present invention provide a machine vision-based motion control method for a window cleaning machine and related devices. The terms "first," "second," "third," "fourth," and so forth (if any) in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "comprising," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or apparatus.

[0023] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 In one embodiment of the present invention, a method for controlling the movement of a window cleaning machine based on machine vision includes:

[0024] Step S1: performing multispectral image acquisition and superpixel segmentation on the window surface to obtain a probability map of dirty areas;

[0025] It is understandable that the execution subject of the present invention can be a machine vision-based window cleaning machine mobile control device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.

[0026] Specifically, based on a dual-spectrum camera array installed at the front end of the window cleaning machine, the window surface is synchronously captured with visible light and near-infrared images to obtain high-resolution raw image data. The raw image data is then interpolated and filtered to restore color information through Bayer interpolation. Image artifacts introduced by uneven sampling are also reduced using methods such as Gaussian interpolation. The lens distortion correction algorithm, based on the camera's internal and external parameter models, eliminates barrel distortion or pincushion distortion caused by the wide-angle lens, aligning the image space coordinates with the real-world coordinates to obtain spatially corrected image data. Adaptive histogram equalization is then performed on the spatially corrected image. By analyzing global and local pixel distributions, the grayscale mapping relationship is dynamically adjusted to achieve a balance between overall and local brightness, effectively alleviating the effects of uneven window illumination, strong reflections, or shadow areas, and obtaining light-balanced image data. Gaussian blur filtering with a fixed kernel size is applied to the illumination-balanced image to reduce noise, filter out high-frequency noise, and improve the signal-to-noise ratio of dirt details, generating de-noised image data. Feature point matching and sub-pixel registration algorithms are then used to spatially align the multispectral images, enabling precise fusion of visible and near-infrared information at the pixel level. This ensures that each image block reflects both the surface material and infrared reflectance characteristics, resulting in spatially aligned multispectral image data. Simultaneously, temporal difference comparison is performed on multiple time-series images. By analyzing pixel changes between adjacent frames, spurious changes caused by robot vibration, slight drift, or external dynamics are removed, improving the stability and effective information density of the original image and generating a preprocessed image dataset. Based on the preprocessed image dataset, a superpixel segmentation algorithm is employed, focusing on texture, color, and infrared reflectance characteristics. Each image is divided into a large number of small regions with extremely high local homogeneity, enabling each superpixel region to more precisely capture the structural and material characteristics of the local dirt. On this basis, multi-dimensional regional features are extracted for each superpixel area, including multimodal feature vectors such as color histogram, texture filter response, and infrared reflectivity. The regional features are input into the deep neural network model. Combined with the spatial attention mechanism, the differences and similarities between superpixels are comprehensively analyzed, and the dirtiness probability value of each area is output. All superpixel probabilities are spliced ​​and reconstructed into a dirty area probability map.

[0027] In this example, the preprocessed image dataset is fed into the SLICO superpixel segmentation algorithm, which efficiently and adaptively segments the entire image into regions, producing a set of superpixel segmented image data with uniform spatial distribution and strong edge preservation. The SLICO algorithm dynamically adjusts the region scale parameters during the segmentation process, ensuring that each segmented superpixel region closely matches the actual dirt distribution characteristics of the window surface in terms of color, texture, and boundary shape. For each superpixel region in the segmented image data, core features are extracted from three dimensions: color features, which distinguish between dirty and clean areas by constructing an RGB histogram and calculating the color distribution of pixels within the region; texture features, which use Gabor filters to summarize the multi-directional and multi-scale responses of the superpixel region to characterize the regional surface structure and detail variations; and near-infrared reflectance intensity features, which calculate the mean and variance of the region in the near-infrared band to highlight dirt or special material responses that are difficult to identify using visible light. All features are concatenated into a high-dimensional regional feature vector, characterizing the multimodal composition and state of each superpixel region. At the same time, based on the spatial contact relationships between superpixels, a superpixel adjacency matrix is ​​automatically constructed to describe the spatial proximity and topological structure of each region in the image. The regional feature vectors of each pair of adjacent superpixels are combined, and a feature difference matrix is ​​calculated using methods such as cosine similarity. This matrix reflects subtle variations in color, texture, and infrared reflectance between adjacent regions on the window surface. The regional feature vectors and feature difference matrix are input into a pre-trained deep convolutional neural network. Convolution operations extract local spatial patterns and inter-regional correlations, enhancing the ability to recognize complex contamination morphologies and structures. A spatial attention mechanism is incorporated into the convolution process to amplify regions with significant feature mutations or high contamination suspicion, giving them a higher weight in the overall discrimination process. Using a probability mapping layer, the regional contamination feature data obtained through deep convolution is normalized using a Sigmoid or Softmax function, outputting a contamination probability value for each superpixel region. All probability values ​​are combined to reconstruct the final contamination region probability map.

[0028] Step S2: performing multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain the movement trajectory of the window cleaning machine;

[0029] Specifically, the probability map of dirty areas is fed into the encoder of a multispectral feature fusion network. Within the encoder, a multi-level convolutional feature extraction operation is used to abstract the original probability map layer by layer, fully exploiting the spatial relationships and multimodal details of dirt distribution at different scales to generate a set of feature maps containing contextual information. This multi-level feature map set is then fed into the decoder module of the fusion network. Within the decoder, spatial resolution is restored through layer-by-layer upsampling. Skip connections are implemented at each decoding level to fully integrate high-level semantic features with low-level detail features, ensuring that the fusion process preserves local texture and edge information while maintaining global structural integrity, ultimately forming an initial fused feature map. For each superpixel region, its texture complexity, brightness variation, and near-infrared reflectance characteristics are analyzed. A region-level attention weight matrix is ​​adaptively calculated to adjust the fusion strategy for each region. This matrix dynamically increases the weight of near-infrared features in complex scenes such as high light reflections, strong shadows, or unusual materials, while increasing the contribution of visible light features in areas with rich texture or prominent visible light, ultimately generating preliminary fused feature data. A global illumination assessment mechanism is introduced into the preliminary fused feature data. By analyzing the brightness histogram and gradient distribution of the overall image, the window surface is subdivided into different regions of strong illumination, weak illumination, and normal illumination. Differentiated feature fusion and enhancement strategies are then applied to each illumination category to ensure spatially balanced representation of the final feature data, resulting in illumination-balanced feature data. A nonlinear transformation using residual connections is then performed on the illumination-balanced feature data. After nonlinear mapping and feature enhancement, a highly discriminative enhanced regional feature map is generated. Based on the enhanced regional feature map, dirty region segmentation and clean path planning are performed. Fine segmentation algorithms such as conditional random fields are used to achieve high-precision segmentation of dirty regions, improving the accuracy of boundary fitting. Based on the segmentation results and combined with regional spatial connectivity and distribution density, a modified traveling salesman problem algorithm is used for path optimization and B-spline smoothing, achieving scientific planning of the window cleaning machine's trajectory.

[0030] In this embodiment, the enhanced region feature map is input into a conditional random field model for refined processing. By incorporating spatial dependencies between regions, the conditional random field model bases the labeling of each pixel or superpixel not only on its own features but also on the feature similarity and spatial coherence of its neighboring regions. This effectively mitigates the missegmentation issues caused by simple thresholding or single regional feature discrimination. After conditional random field energy optimization, the output target segmentation result can accurately depict the true boundaries of dirty regions and maintain high robustness in complex scenes such as high light and low contrast. A morphological opening operation is performed on the target segmentation result. By opening the segmented image using structuring elements, small isolated noise and pseudo-dirty regions are effectively removed while preserving the main contour structure of the dirty target, resulting in cleaner and more coherent segmentation data. Furthermore, the noise-filtered segmentation data is binarized, with all pixels calibrated as dirty or non-dirty according to a preset threshold. A dirty region mask is then output, reflecting the spatial structure of the dirt distribution on the window surface. Based on the dirty area mask map, a dirty area connectivity map is constructed through connected domain analysis. All independent dirty areas are numbered and feature extracted, and the minimum circumscribed rectangle or center point of each connected area is calculated as the target point data for path planning. Based on the path planning target points, an optimization algorithm such as the improved traveling salesman problem solver is used to generate a global clean path. This path aims to achieve the shortest coverage of all dirty target points, and the physical motion characteristics and obstacle avoidance of the window cleaning machine are fully considered in the path design. The global clean path is smoothed using B-spline curves to eliminate sharp corners in the path, and the continuity and curvature changes of the trajectory are controlled to improve the dynamic tracking capability and movement smoothness of the window cleaning machine. The smoothed global path is decomposed into a series of local path segments to obtain the movement trajectory of the window cleaning machine.

[0031] Step S3: establishing a bilinear system dynamics model based on the movement trajectory of the window cleaning machine and designing a visual feedback state observer to obtain real-time state information of the window cleaning machine;

[0032] Specifically, the motion of the window cleaning machine is parameterized based on the cleaning path planning results. Physical quantities such as the spatial position coordinates, heading angle, linear velocity, and angular velocity at each moment are collectively defined as a six-dimensional state vector. The independent driving forces of the left and right drive wheels are used as control input vectors, resulting in the window cleaning machine's motion state equation. This equation describes the dynamic characteristics and posture changes of the window cleaning machine within the two-dimensional plane of the window surface. Based on the motion state equation, a bilinear system dynamics model with state-dependent terms is constructed. This means that the system's state transitions depend not only on the existing state and inputs but also on the nonlinear coupling effects of the state variables themselves. This enhances the dynamics model's ability to represent complex phenomena such as friction, inertia, and changes in gravity distribution during actual operation. To ensure the accuracy of the model parameters, real-time state information is collected from a large number of discrete sampling points as the window cleaning machine moves along its trajectory. By analyzing the state transition relationships between these sampling points, the model parameters are globally calibrated using the least squares method, effectively reducing the influence of environmental noise and measurement errors to obtain the system dynamics parameters. Based on the system dynamics parameters, a visual feedback state observer is constructed. The position and attitude measurements of the BWM in the window coordinate system, acquired in real time by the machine vision module, are integrated with the bilinear dynamics model. State estimation is achieved using an extended Kalman filter (EKF). The EKF effectively handles the nonlinear relationship between the model and observations, continuously correcting state errors in the prediction-observation cycle and outputting preliminary estimates of the BWM's current position, velocity, and orientation. Based on these preliminary state estimates, the statistical characteristics of the observation and prediction errors are evaluated in real time, and the Kalman filter's covariance matrix parameters are dynamically adjusted. This allows the filter algorithm to adaptively address scenarios such as image clarity changes, external lighting disturbances, and occasional occlusions, improving robustness and accuracy. Furthermore, data from the IMU sensor on the BWM itself is incorporated into the state observation process. A complementary filtering strategy is used to supplement the Kalman filter results. The IMU's high-frequency, short-term, stable motion information compensates for the limitations of visual information under conditions of rapid motion or strong interference, achieving precise state fusion across a wide dynamic range and obtaining high-precision, low-latency, real-time state information about the BWM.

[0033] Step S4: Input the movement trajectory and real-time status information of the window cleaning machine into the soft-constraint visual iterative regulator for control calculation to obtain a smooth movement control instruction.

[0034] Specifically, the real-time state information obtained through the fusion of the trajectory and vision-dynamics is input into the soft-constrained visual iterative regulator module. Within this module, an overall cost function, including state tracking error, control input energy consumption, and soft constraint slack variables, is dynamically defined based on the current trajectory and state. This cost function emphasizes the precise tracking of the BMU along the planned trajectory while balancing the smoothness of the control input and the rationality of the physical constraints. Slack variables are introduced to quantify and adjust the flexibility of the constraint boundaries, thereby constructing an optimal control optimization problem model under soft constraints. The soft constraint model allows for the introduction of slack variables when the state or control variable approaches the constraint boundary, relaxing the originally strict non-violation constraints into a cost penalty term. This allows the optimization algorithm to maintain system stability and feasibility under extreme operating conditions or disturbances. For the constraint expression containing slack variables, the model is linearized along the BMU trajectory near the current state, approximating the nonlinear system as a series of local linear subproblems. The backpropagation algorithm is applied to these linear approximations, combined with the analytical recursion of the Riccati equation from optimal control theory, to gradually derive the optimal feedback gain and control correction at each step, forming an iteratively convergent control strategy. To enhance the rigor and flexibility of constraint handling, a barrier function is introduced into the control strategy design. The proximity of the state and control variables to the constraint boundaries is reflected in the overall cost via a logarithmic barrier term. The weights of the slack variables are adaptively adjusted based on real-time state information. When the BMU state approaches physical or safety boundaries, the relevant slack weights are automatically increased, enabling dynamic response and fine-tuning of constraint activity. Based on the control sequence derived from this optimization, an efficient quadratic programming solver is used to map the optimized instructions to the actual input range and system constraints. This not only ensures the smoothness and continuity of the motion trajectory within each control cycle, but also effectively suppresses unexpected anomalies caused by external disturbances and environmental changes, generating smooth motion control instructions.

[0035] In an embodiment of the present invention, multispectral image acquisition and superpixel network-based regional feature extraction technology can effectively capture subtle differences between dirty and clean areas, resolving the detection difficulties caused by the similarities between lightly dirty and clean surface features and enabling the window cleaning machine to accurately identify various degrees of dirt. A multispectral imaging mechanism based on regional attention fusion is employed, combining information from the visible and near-infrared spectral domains to adaptively fuse features of window surfaces under different lighting conditions, effectively addressing dirt detection issues under complex lighting conditions such as strong light reflections and shadowed areas. Dirty area target segmentation and cleaning path planning are performed based on enhanced regional feature maps. Zigzag or spiral-shaped refinement paths are automatically generated for large dirty areas, while B-spline curve smoothing ensures path continuity, achieving efficient cleaning coverage for various complex dirt distributions. A bilinear system dynamics model is employed to accurately capture the nonlinear dynamic behavior of the window cleaning machine. A high-order sliding mode disturbance observer is used to compensate for external disturbances, significantly improving the accuracy of state estimation under disturbance conditions such as wind and surface friction changes. The soft-constrained visual iterative regulator of the present invention converts hard constraints into soft constraints by introducing slack variables and obstacle functions to process constraints, effectively avoiding the control jitter problem when constraints are activated in traditional control methods, and generating smooth and continuous control inputs even under conditions of external interference.

[0036] In a specific embodiment, the process of executing step S1 may specifically include the following steps:

[0037] The dual-spectrum camera array installed at the front end of the window cleaning machine collects images of the window surface to obtain original image data;

[0038] Performing interpolation filtering and lens distortion correction on the original image data to obtain spatially corrected image data, and performing adaptive histogram equalization based on the spatially corrected image data to obtain illumination-balanced image data;

[0039] Performing Gaussian blur filtering on the light-balanced image data to reduce noise, obtaining reduced-noise image data, and spatially aligning the reduced-noise image data to obtain spatially aligned multispectral image data;

[0040] Perform temporal difference comparison on the spatially aligned multispectral image data to obtain a preprocessed image dataset;

[0041] Superpixel segmentation and regional feature extraction are performed based on the preprocessed image dataset to obtain a dirty area probability map.

[0042] Specifically, a dual-spectrum camera array is integrated at the front end of the window cleaning machine. This array consists of a visible light camera and a near-infrared camera, both of which are strictly aligned on the same visual axis. Precision mechanical structure and electronic synchronization control ensure that they capture the same area of ​​the window surface. The visible light channel can capture the color, shape, transparency, and general signs of dirt on the window surface, while the near-infrared channel can penetrate some surface reflections and film, and has higher sensitivity to hidden features such as faint dust, oil stains, fingerprints, and interfering reflections. The fusion of these two data enhances the overall system's adaptability to diverse dirt and special surface treatment processes. After the robot is powered on, the camera array synchronously captures raw image data at a fixed frame rate and resolution, forming a dual-channel, high-resolution, time-sequential raw visual sequence. Each raw image is then interpolated and filtered. Bayer interpolation restores the color image content, and spatial interpolation fills in small missing points caused by uneven pixel distribution, partial occlusion, or signal acquisition interference during the acquisition process. Simultaneously, the lens distortion correction module, based on the camera's intrinsic and extrinsic parameters obtained during factory calibration, employs a pinhole model and distortion polynomials to perform pixel-by-pixel correction for nonlinear distortions such as barrel and pincushion distortions in the original image. This maps the actual imaging space back to the standard orthogonal window plane, achieving a precise correspondence between spatial coordinates and the physical world, and outputs spatially corrected image data. An adaptive histogram equalization module dynamically adjusts the global and local brightness distribution of the spatially corrected image. Using block equalization and contrast-limited algorithms, it improves overall contrast and brightness balance, effectively suppressing overexposure in highlights and loss of detail in shadows, resulting in a light-balanced image. The light-balanced image data is then subjected to Gaussian blur filtering with a kernel size of 5×5 or 7×7. This filtering process significantly suppresses high-frequency noise points and extreme pixel abrupt changes while preserving the main image contours and texture structure, enhancing subsequent segmentation of subtle dirt structures. High-precision spatial alignment is performed on the multi-channel image data after denoising, as the denoised data can contain pixel-level spatial errors due to subtle lens parallax, mounting angle, or robot motion. Using a feature point matching-based image registration algorithm, key points are extracted between visible light and near-infrared images at the same time point using local feature operators such as SIFT and SURF. False matches are eliminated using the RANSAC algorithm. Combined with sub-pixel affine transformation and interpolation resampling, the multispectral images are precisely overlaid onto a unified spatial reference frame, forming spatially aligned multispectral image data. Temporal difference comparison is performed on the spatially aligned multispectral image data. By analyzing the pixel-level differences between adjacent time frames, non-static signals caused by robot micro-vibrations, short-term environmental disturbances, or occasional camera jitter are filtered out. Interference effects such as localized water reflections and dust disturbances are eliminated, retaining only the stable structure and true dirt distribution of the window surface, thus obtaining the preprocessed image dataset.Based on the preprocessed image dataset, the SLICO superpixel segmentation algorithm is used to perform fine-grained segmentation on each multispectral image. The SLICO algorithm can adaptively divide the entire image into several superpixel regions with extremely high local uniformity based on the color, brightness, texture, and spatial proximity of the pixels. Each region maintains clear boundaries of the main structure of the window surface and has good computational efficiency. The feature vectors and adjacency information of all superpixel regions are input into a deep convolutional neural network for global modeling. The convolution and spatial attention mechanisms are used to enhance the recognition sensitivity to complex dirt, fuzzy boundaries, and special material differences. The convolutional neural network outputs the dirt probability value corresponding to each superpixel region. Regions with high probabilities are judged as key cleaning objects, and regions with low probabilities are considered clean or background. The probabilities of all superpixels are spliced ​​and reconstructed into the final probability map of the dirty area.

[0043] In a specific embodiment, the step of performing superpixel segmentation and region feature extraction based on the preprocessed image dataset to obtain a dirty region probability map may specifically include the following steps:

[0044] The preprocessed image data set is input into the SLICO superpixel segmentation algorithm for region segmentation processing to obtain segmented image data;

[0045] Extract color features, texture features, and near-infrared reflection intensity features from each superpixel region in the segmented image data to obtain a regional feature vector, and construct a superpixel adjacency matrix based on the segmented image data;

[0046] Calculate the feature difference matrix between adjacent superpixel regions based on the regional feature vector and the superpixel adjacency matrix;

[0047] Deep convolution is performed on the regional feature vector and the feature difference matrix to obtain the regional dirt feature data, and probability mapping is performed on the regional dirt feature data to obtain a dirty area probability map.

[0048] Specifically, the preprocessed image dataset is subjected to regional segmentation and feature modeling. The SLICO superpixel segmentation algorithm is used as the core tool for region segmentation. The SLICO algorithm uses the color and spatial coordinates of pixels as its basis, combining their grayscale and texture information for adaptive iterative clustering. Its greatest advantage is its ability to autonomously determine the optimal clustering parameters and neighborhood scale based on the actual image content, ensuring that the segmented superpixel regions not only accurately fit the boundaries, but also have extremely high internal consistency and controllable regional scale. During the specific implementation of the algorithm, several superpixel centroids are uniformly initialized in the image space. Then, through multiple rounds of clustering iterations, pixels are each classified to the nearest superpixel center. In the clustering step, spatial distance, color distance, and gradient penalty terms are calculated in real time to ensure clear segmentation at window edges, strong reflective areas, and dirty boundaries. Ultimately, the entire image is efficiently divided into a large number of small and dense regional blocks, forming segmented image data. Multimodal regional feature extraction is performed on each superpixel region. Color features are extracted by statistically analyzing the RGB distribution histograms of all pixels within a superpixel, or by expanding them into a multidimensional feature space encompassing visible and near-infrared light under multispectral conditions. This allows for quantitative analysis of regional color principal components, saturation, and color difference. Texture feature extraction utilizes classic image processing tools such as Gabor filters or wavelet transforms to analyze pixels within a region at multiple directions and scales, capturing the periodicity, roughness, and directional characteristics of the local spatial structure, enhancing sensitivity to non-uniform contamination such as fingerprints, water stains, and atomized films. Near-infrared reflectance intensity features analyze the mean, standard deviation, and distribution pattern of superpixels in the near-infrared channel to capture hidden anomalies such as special materials, foreign matter beneath transparent films, and signs of aging that are difficult to detect using visible light. All extracted features are concatenated into a high-dimensional regional feature vector of uniform length according to predefined rules. Furthermore, a superpixel adjacency matrix is ​​constructed based on the spatial contact relationships within the superpixel segmented image. The adjacency matrix is ​​essentially a sparse binary relational table describing the direct adjacency between each pair of spatially adjacent superpixels. Based on the regional feature vectors and the superpixel adjacency matrix, a feature difference matrix is ​​calculated between adjacent superpixel regions. This difference matrix traverses all adjacent superpixel pairs and uses metrics such as cosine similarity, Euclidean distance, and Manhattan distance to measure the relative differences in color, texture, infrared, and other features between each pair of regions. This matrix quickly captures the boundary information and degree of blur between dirty and clean areas under complex phenomena such as edges, gradients, and blurred transitions. This difference matrix can effectively compensate for the limitations of single-region feature descriptions, especially when there are severe reflections or uneven lighting on window surfaces. The regional feature vectors and feature difference matrix are input into a deep convolutional neural network for end-to-end feature learning and discrimination.The network architecture comprises multiple convolutional layers, pooling layers, and a spatial attention mechanism. The convolution operation aggregates regional and neighborhood features to identify complex cross-regional dirt patterns and distribution patterns. The spatial attention module automatically assigns higher discriminant weights to superpixel regions with significant feature mutations and high probability, further enhancing detection sensitivity in complex or fuzzy areas. Through multiple rounds of forward and backpropagation training, combined with extensive sample data and manual annotation, the convolutional neural network gradually optimizes parameters, achieving adaptive discrimination and robust detection of dirt of different types, shapes, and materials. The network ultimately outputs a dirt feature value for each superpixel. After completing end-to-end learning of regional dirt feature data, the probability mapping layer converts the convolutional network output into a standardized dirt probability value. A sigmoid or softmax activation function is applied to the convolutional output data, mapping the feature values ​​to a range of 0 to 1, representing the probability of each superpixel being classified as dirt. These probability values ​​are then concatenated and reconstructed into a probability map of dirt regions in a spatially distributed manner.

[0049] In a specific embodiment, the process of executing step S2 may specifically include the following steps:

[0050] The dirty area probability map is input into the encoder of the multispectral feature fusion network for multi-level feature extraction to obtain a feature map set;

[0051] The decoder upsamples the feature map set and sets skip connections at each decoding level to obtain the initial fused feature map;

[0052] The regional attention weight is calculated based on the texture complexity, brightness change and near-infrared reflectance characteristics of each superpixel region to obtain the regional attention weight matrix;

[0053] Adaptively weight the visible light features and near-infrared features according to the regional attention weight matrix to obtain preliminary fused feature data;

[0054] Performing global illumination evaluation on the preliminary fused feature data and classifying it by illumination region to obtain illumination balance feature data, and performing nonlinear transformation processing of residual connection on the illumination balance feature data to obtain enhanced regional feature map;

[0055] Based on the enhanced region feature map, dirty area segmentation and cleaning path planning are performed to obtain the moving trajectory of the window cleaning machine.

[0056] Specifically, the probability map of dirty areas is input into the encoder of a multispectral feature fusion network. This encoder is designed as a multi-stage deep convolutional architecture, with each stage consisting of two or more convolutional layers and a max pooling layer. This architecture sequentially extracts multi-scale features, ranging from low-level edge details to high-level abstract semantics. During this process, the network analyzes the spatial distribution of the probability map itself and uses the original multispectral image (visible and near-infrared channels) spatially aligned with the probability map as auxiliary input. Through feature concatenation or parallel convolution of independent channels, the network fully exploits rich physical properties such as color, texture, and shape in visible light, and material and implicit reflectance in the near-infrared. The output of each convolutional stage is stored as a set of feature maps. This set of feature maps is then input into a symmetrically structured decoder module. The decoder continuously restores spatial resolution through progressive upsampling operations (such as deconvolution or transposed convolution), allowing high-level abstract features to be mapped back to the spatial domain of the original image size. At each decoding level, skip connections are designed to directly concatenate feature maps from the same encoder level into the current decoder input, fusing low-level spatial details with high-level semantic information. The U-Net-style architecture significantly improves the accuracy and stability of the network when segmenting complex shapes and fine-grained boundaries. After decoder upsampling and multi-level skip connections, the network outputs an initial fused feature map. For each superpixel region in the initial fused feature map, the original input and the encoder-decoder features are combined to calculate a region-level attention weight matrix. The region-level attention mechanism is centered on three dimensions: texture complexity measures the complexity of regional details and structure by statistically analyzing the Gabor response, multi-scale variance, and texture distribution uniformity of pixels within the region. Regions with high complexity are often more prone to heterogeneous dirt or reflective interference. Brightness variation quantifies regional lighting unevenness, shadows, overexposure, and other issues through local histogram, mean, and gradient analysis. Regions with drastic brightness changes are given higher weights to improve the robustness of dirt segmentation in both low-light and high-light areas. Near-infrared reflectivity analyzes the mean and variance of reflective intensity across multiple spectral channels to locate unique dirt or materials that are difficult to distinguish in visible light but stand out in the infrared band, thereby improving the detection sensitivity of hidden pollution. These three attributes are weighted and normalized to obtain the attention weight matrix of each superpixel. Based on the regional attention weight matrix, the visible light and near-infrared feature maps are adaptively weighted fused. For areas with high reflectivity, extreme lighting or complex structures, the network automatically increases the proportion of near-infrared features in the fusion, thereby suppressing misjudgment and information loss under visible light; while in areas with rich textures and obvious color differences, the contribution of visible light features is enhanced to achieve fine modeling of scenes that are difficult to distinguish with traditional vision. The weighted fusion mechanism is completed through weighted averaging, pixel-by-pixel fusion or feature splicing within the region to generate preliminary fusion feature data with strong spatial adaptability. Based on the preliminary fusion feature data, the brightness distribution in the entire image is evaluated globally, and the image is classified into illuminated areas.Through histogram analysis, block-by-block brightness statistics, and spatial gradient analysis, the window surface is automatically divided into three categories: strong, weak, and uniformly illuminated areas. Differentiated feature fusion parameters and subsequent processing strategies are applied to each area. For example, in strong-light areas, infrared wavelengths and contrast suppression are emphasized, while in weak-light areas, visible light brightness and texture enhancement are enhanced, ensuring that feature representation remains stable under varying lighting conditions. After illumination equalization, deep residual connections and nonlinear transformations are introduced into the fused feature map. By stacking residual blocks and introducing activation functions, high-order features are enhanced, anomalies are amplified, and background information is suppressed, resulting in an enhanced region feature map. Based on the enhanced region feature map, algorithms such as conditional random fields, deep discriminant networks, or adaptive threshold segmentation are used to achieve high-precision segmentation of dirty areas on the window surface. Segmentation results are presented as probabilistic or binary masks, improving boundary refinement and segmentation accuracy in heterogeneous materials. Through connected domain analysis and minimum bounding rectangle or region centroid extraction, all independent dirty areas are converted into a set of target nodes for path planning. A modified traveling salesman problem algorithm is used to optimize the global path to the target node. B-spline curves or cubic splines are used for path smoothing, addressing the mechanical losses and motion redundancy issues associated with traditional linear paths. The path is broken down into local segments based on the physical distance traveled, allowing for the planning of a trajectory that efficiently covers all dirty areas while satisfying the dynamic and actuator constraints of the window cleaning robot.

[0057] In a specific embodiment, the step of performing dirty area segmentation and cleaning path planning based on the enhanced region feature map to obtain the movement trajectory of the window cleaning machine may specifically include the following steps:

[0058] The enhanced region feature map is input into the conditional random field model for refinement to obtain the target segmentation result;

[0059] Perform morphological opening operation on the target segmentation result to obtain the noise-filtered segmentation data, and then perform binarization on the noise-filtered segmentation data to obtain the dirty area mask map;

[0060] Based on the dirty area mask map, a connected map of the dirty area is constructed to obtain the path planning target point data, and a global clean path is generated based on the path planning target point data;

[0061] The global cleaning path is smoothed using B-spline curves and decomposed into local path segments to obtain the moving trajectory of the window cleaning machine.

[0062] Specifically, high-precision dirt segmentation and target region identification are performed on the fused and enhanced region feature map. The enhanced region feature map is then fed into a conditional random field model for refined processing. Conditional random fields are probabilistic graphical models that model spatial continuity and contextual dependencies. Their advantage lies in their ability to fully leverage feature similarity and spatial adjacency between pixels (or superpixels), effectively suppressing common problems such as isolated misclassification and blurred boundaries. The conditional random field model employs a fully or sparsely connected structure, aiming to minimize energy. It jointly considers single-point features and neighborhood smoothness constraints at each pixel or region. This step not only allows for preliminary classification based on the region's own high-dimensional feature representation but also adaptively adjusts the segmentation boundary based on feature differences between adjacent regions, achieving refined optimization for scenarios such as high light reflections, complex dirt interface, and weak contrast boundaries. After iterative inference and energy optimization using the conditional random field, the target segmentation result is obtained. A morphological opening operation is then performed on the target segmentation result. This opening operation consists of a sequence of erosion and dilation operations, with a commonly used structuring element being a 3×3 or 5×5 circle, ellipse, or cross template. The introduction of the morphological opening operation effectively removes small, isolated noise, pseudo-dirt points, and sharp burrs at the boundaries of the segmentation mask, preserving the complete morphology and coherent structure of the primary dirty objects, resulting in a smoother and more connected segmentation result. After the opening operation, an adaptive threshold is set for the noise-filtered segmentation data. A binarization operation is performed to classify all pixels with a probability above the set threshold as dirty areas, and vice versa, as non-dirty areas, generating a high-resolution dirty area mask. Based on this dirty area mask, a connected domain analysis algorithm is used to identify and number all connected regions identified as dirty. Connected domain analysis automatically separates multiple, independent dirty blocks. For each connected domain, spatial feature information such as its bounding rectangle, centroid coordinates, principal axis direction, and boundary point set is calculated. The spatial distribution, area size, and relative position of all independent dirty areas are fully quantified into a target point dataset for path planning. Based on these target points, a global cleaning path covering all dirty target points is automatically generated for the window-cleaning robot using graph theory algorithms such as the modified traveling salesman problem or the shortest Hamiltonian path algorithm. During path generation, the robot's physical size, motion capabilities, cleaning coverage width, actual mechanical constraints, and window shape are comprehensively considered to ensure that the path achieves full coverage while minimizing motion redundancy and cleaning repetition. Based on the generated path, the global cleaning path is smoothed using B-spline curves. B-spline curves have excellent mathematical properties, and by setting a small number of control points, they achieve high-order continuity, smooth curvature, and controllable speed along the entire path.In specific implementation, the key nodes of the global path are used as the control points of the B-spline, and smooth cubic or higher-order spline curves are generated through interpolation and fitting algorithms to make the path reachable at all nodes, while eliminating sharp acceleration and deceleration, steering oscillations, and sudden changes in actuator loads during mechanical movement. The parameters of the spline curve are flexibly adjusted according to the robot's dynamic model, maximum acceleration, and actual window surface shape to ensure that the path can closely follow the distribution of dirt and fully comply with kinematic and dynamic safety constraints. In order to adapt to the requirements of the robot's actual execution stroke and state observation window, the smoothed global path is decomposed into multiple local path segments according to preset intervals. The length and complexity of each path segment are limited to the ability of the window cleaning robot to move in a single move, and switching, compensation, and replanning are carried out in real time according to the dynamic environment and sensor feedback to finally obtain the movement trajectory of the window cleaning machine.

[0063] In a specific embodiment, the process of executing step S3 may specifically include the following steps:

[0064] Based on the trajectory of the window cleaning machine, a six-dimensional state vector containing position coordinates, heading angle, and velocity components is defined, as well as a control input vector containing the driving forces of the left and right driving wheels. The state equation of motion of the window cleaning machine is obtained.

[0065] A bilinear system dynamics model including state-dependent terms is established based on the state equation of motion of the window cleaning machine.

[0066] By analyzing the state transition relationship between multiple sampling points on the moving trajectory of the window cleaning machine, the least squares method is used to calibrate the parameters and obtain the system dynamic parameters.

[0067] A visual feedback state observer is constructed based on the system dynamics parameters. The position and attitude measurement data obtained by machine vision are combined with the bilinear system dynamics model. The state fusion is performed through the extended Kalman filter algorithm to obtain the preliminary state estimation result.

[0068] The filtering parameters are dynamically adjusted based on the preliminary state estimation results, and complementary filtering is performed through IMU data to obtain the real-time state information of the window cleaning machine.

[0069] Specifically, based on the robot's planned path, a mathematical definition of the relationship between state variables and dynamics is performed. For the window cleaning robot's trajectory on the window surface, the spatial position coordinates, body orientation angle, and velocity components along the local coordinate system (including x- and y-direction velocity, and angular velocity) at each moment are combined to form a six-dimensional state vector. Furthermore, for the physical implementation of the window cleaning robot's drive unit, the system uses the actual controlled driving forces of the left and right drive wheels as control inputs. Based on these six-dimensional state vectors and control input vectors, the window cleaning robot's kinematic state equation is derived. This equation describes how the robot's state variables change over time under a given control input, expressed as a differential equation. During actual window cleaning robot movement, the robot exhibits strong nonlinearity and state dependence due to various perturbations, such as friction changes, center of gravity shift, structural elasticity, and external wind forces. Therefore, a bilinear system dynamics model is used to describe its core dynamic characteristics. This bilinear modeling approach effectively captures the complex coupling effects between state and input, and theoretically adapts to various high-order perturbations and nonlinear dynamics, improving the window cleaning robot's adaptability and predictive capabilities in dynamic environments. As the robot moves along its trajectory, it collects a large number of discrete state sampling points. By recording the state and control inputs at each time slice at high frequency and statistically analyzing the state transition relationships between adjacent sampling points, the actual discrete state transition sequence is constructed. Based on these sampling points, the least squares method is used to globally fit the unknown parameters of the dynamic model. This method iteratively adjusts the parameters in the state correlation matrix by minimizing the sum of squared errors between the observed data and the model predictions until the model output closely matches the actual measured data. A visual feedback state observer is constructed based on the dynamic parameters. The position and posture observations of the BMU in the window coordinate system, acquired in real time by the machine vision system, are deeply integrated with the established bilinear system dynamics model to achieve high-precision estimation of all state variables. In the actual algorithm implementation, the extended Kalman filter (EKF) is selected as the main state estimation algorithm. The EKF effectively handles nonlinear dynamics and observation relationships. Its principle is to perform first-order Taylor linearization on the system dynamics and observation model at each time step. Using a prediction-update loop structure, it continuously integrates the model predictions with the observed values ​​to output the current optimal state estimate. During the prediction step, the EKF uses the previous state estimate, known control inputs, and bilinear dynamic equations to infer the prior distribution of the current state. During the update step, the EKF combines position and attitude observations acquired by machine vision with the observation equations to correct the state prediction and output a posterior estimate of the filtered state. At each moment, the system dynamically adjusts the EKF filter's noise covariance matrix (Q and R parameters) based on the real-time error and residual distribution of the state estimate. This allows the system to adaptively respond to changes in observation error under typical operating conditions, such as varying lighting, strong reflections, and obstructed vision, thereby improving the stability and anti-interference capabilities of the filtering algorithm.To compensate for the shortcomings of visual measurement in situations involving short bursts of high-speed motion, strong vibration, or momentary frame loss, IMU (Inertial Measurement Unit) sensor data is introduced. A complementary filtering strategy is used to weightedly fuse high-frequency dynamic data such as acceleration and angular velocity provided by the IMU with the low-frequency, high-precision state estimate output by the EKF. Complementary filtering improves the response speed and continuity of state estimation over short timescales while maintaining the consistency and accuracy of global pose and trajectory over long timescales. This provides real-time status information for the window cleaning machine.

[0070] In a specific embodiment, the process of executing step S4 may specifically include the following steps:

[0071] Based on the BMW's movement trajectory and real-time status information, a cost function including state tracking error, control input, and slack variables is defined to obtain a soft-constrained optimization problem model.

[0072] Slack variables are introduced into the state and control constraints in the soft constraint optimization problem model, hard constraints are converted into soft constraints, and constraint expressions containing slack variables are obtained;

[0073] The constraint expression containing slack variables is linearized along the trajectory of the window cleaning machine to obtain a linear approximate model. The linear approximate model is then back-propagated and the Riccati equation is solved to obtain the control strategy.

[0074] Based on the control strategy, the obstacle function processing constraint is introduced, and the slack variable weights are adaptively adjusted according to the real-time status information of the window cleaning machine to obtain the optimized control sequence.

[0075] The optimized control sequence is constrained by a quadratic programming solver to obtain smooth movement control instructions.

[0076] Specifically, based on the BMU's trajectory and real-time state information, a dynamic cost function is constructed that comprehensively considers state error, control cost, and the ability to handle flexible constraints. This function aims to minimize control energy consumption, control the intensity of movement, and limit violations near physical limits, while ensuring the robot tracks the reference path as accurately as possible. Therefore, the cost function assigns different weights to trajectory error, actual output control commands, and the values ​​of slack variables, thereby balancing accuracy, stability, and safety. Based on this cost function, a soft-constrained optimization problem model is constructed. Slack variables are introduced during the modeling process, transforming the originally strict boundary conditions (such as position cannot exceed the bounds, speed cannot be excessive, acceleration cannot change suddenly, and driving force cannot exceed the maximum load) into soft constraints with tolerances. By introducing slack variables, appropriate "violations" are implemented when the constraints are not fully met, but these violations are quantified and penalized by incorporating them into the cost function. This approach effectively prevents the control strategy from becoming rigid in the face of complex environmental changes or model uncertainty. This is especially true in the face of brief disturbances, sensor anomalies, or when the path approaches obstacles. Soft constraints ensure that the robot maintains continuous and controllable operation, preventing system stalls or control command oscillations due to infeasible solutions. To enable the optimization model to quickly solve online during dynamic processes, the nonlinear constraint expressions containing slack variables are linearized. Based on the known trajectory of the window cleaning machine, a first-order expansion of the state evolution equation and constraint expressions is performed near the current state point, resulting in a simplified linear approximation model. In this linear model, the relationship between state variables, control variables, and slack variables is described as an efficient linear system, reducing computational complexity. A dynamic optimization method called backpropagation is used to solve the policy for this linear approximation model. By recursively working backward from the endpoint to the starting point, the control strategy is updated and modified at each stage, generating a control strategy sequence. During this reasoning process, the system uses a recursive structure to solve a value function that describes the optimal control strategy. During this process, the control strategy automatically adjusts based on the errors and constraint changes under different states, achieving a unified integration of state prediction and control feedback. An obstacle function is introduced based on the control strategy to handle constraints that are close to or about to trigger the constraint boundary. When the robot approaches a preset physical or safety boundary, such as reaching the edge of a window, the direction of movement deviates too much, or the control output approaches the motor limit, the obstacle function adds a rapidly growing nonlinear penalty term to the cost function, causing the system to pay close attention to and strongly suppress this state during the optimization process, thereby proactively staying away from risk areas and ensuring safe operation.The weight of each slack variable is dynamically adjusted based on the feedback results of the current real-time state of the window cleaning machine. For example, when the state is close to the boundary, the system automatically increases the penalty intensity of the slack variable to encourage the optimization result to converge within the boundary as much as possible. When the state is far away from the boundary, the system will appropriately reduce the weight so that the robot can focus more on optimizing tracking accuracy and control energy consumption. The control strategy sequence generated above is input into the quadratic programming solver for constraint processing and final optimization. While ensuring the minimization of cost, it ensures that all solution results meet the actual requirements of physical control constraints and actuator response boundaries. The quadratic programming solver fully considers the dynamic equations of the linearized model, the upper and lower limit expressions of soft constraints, the nonlinear penalty terms introduced by the barrier function, and the dynamic weights of the slack variables, and outputs the optimal control results within the specified time range to obtain smooth movement control instructions.

[0077] The above describes the window cleaning machine movement control method based on machine vision in the embodiment of the present invention. The following describes the window cleaning machine movement control device based on machine vision in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a window cleaning machine movement control device based on machine vision includes:

[0078] The acquisition module is used to collect multispectral images and perform superpixel segmentation on the window surface to obtain a probability map of dirty areas;

[0079] The planning module is used to perform multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain the movement trajectory of the window cleaning machine;

[0080] The observation module is used to establish a bilinear system dynamics model based on the movement trajectory of the window cleaning machine and design a visual feedback state observer to obtain the real-time status information of the window cleaning machine;

[0081] The control calculation module is used to input the movement trajectory and real-time status information of the window cleaning machine into the soft constraint visual iterative regulator for control calculation to obtain smooth movement control instructions.

[0082] Through the collaborative efforts of these components, multispectral image acquisition and superpixel network-based regional feature extraction techniques effectively capture subtle differences between dirty and clean areas, addressing the detection challenges caused by the similarities between lightly soiled and clean surfaces, and enabling the BMR to accurately identify various levels of soiling. A multispectral imaging mechanism based on region-level attention fusion combines information from both visible and near-infrared spectral domains to adaptively fuse features across window surfaces under varying lighting conditions, effectively addressing soiling detection in complex lighting conditions such as strong reflections and shadows. Enhanced regional feature maps are used for dirty area segmentation and cleaning path planning. Zigzag or spiral refinement paths are automatically generated for large dirty areas, while B-spline smoothing ensures path continuity, achieving efficient cleaning coverage for a wide range of complex soiling distributions. A bilinear system dynamics model accurately captures the nonlinear dynamic behavior of the BMR. A high-order sliding mode disturbance observer compensates for external disturbances, significantly improving state estimation accuracy under disturbances such as wind and surface friction. The soft-constrained visual iterative regulator of the present invention converts hard constraints into soft constraints by introducing slack variables and obstacle functions to process constraints, effectively avoiding the control jitter problem when constraints are activated in traditional control methods, and generating smooth and continuous control inputs even under conditions of external interference.

[0083] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0084] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0085] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0086] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.

[0087] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0089] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A machine vision-based window cleaning machine movement control method, characterized in that: include: Perform multispectral image acquisition and superpixel segmentation on the window surface to obtain a probability map of dirty areas; Performing multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain the moving trajectory of the window cleaning machine; Based on the moving trajectory of the window cleaning machine, a bilinear system dynamics model is established and a visual feedback state observer is designed to obtain the real-time state information of the window cleaning machine; specifically, the method includes: defining a six-dimensional state vector including position coordinates, orientation angle and velocity components according to the moving trajectory of the window cleaning machine, and a control input vector including the driving force of the left and right driving wheels to obtain the motion state equation of the window cleaning machine; establishing a bilinear system dynamics model including state dependencies based on the motion state equation of the window cleaning machine; analyzing the state transfer relationship between multiple sampling points on the moving trajectory of the window cleaning machine, using the least squares method to perform parameter calibration to obtain the system dynamics parameters; constructing a visual feedback state observer according to the system dynamics parameters, and combining the position and posture measurement data obtained by machine vision with the bilinear system dynamics model, performing state fusion through the extended Kalman filter algorithm to obtain a preliminary state estimation result; dynamically adjusting the filter parameters based on the preliminary state estimation result, and performing complementary filtering through the IMU data to obtain the real-time state information of the window cleaning machine; The moving trajectory of the window cleaning machine and the real-time status information of the window cleaning machine are input into a soft-constraint visual iterative regulator for control calculation to obtain a smooth movement control instruction; specifically, the method includes: defining a cost function including state tracking error, control input and slack variables based on the moving trajectory of the window cleaning machine and the real-time status information of the window cleaning machine to obtain a soft-constraint optimization problem model; introducing slack variables into the state and control constraints in the soft-constraint optimization problem model, converting hard constraints into soft constraint forms, and obtaining a constraint expression containing slack variables; linearizing the constraint expression containing slack variables along the moving trajectory of the window cleaning machine to obtain a linear approximation model, and performing backpropagation calculation on the linear approximation model to analyze the Riccati equation to obtain a control strategy; introducing an obstacle function processing constraint based on the control strategy, and adaptively adjusting the slack variable weight according to the real-time status information of the window cleaning machine to obtain an optimized control sequence; constraining the optimized control sequence through a quadratic programming solver to obtain a smooth movement control instruction.

2. The machine vision-based window cleaning machine movement control method according to claim 1, characterized in that: The multispectral image acquisition and superpixel segmentation of the window surface are performed to obtain a probability map of the dirty area, including: The dual-spectrum camera array installed at the front end of the window cleaning machine collects images of the window surface to obtain original image data; performing interpolation filtering and lens distortion correction on the original image data to obtain spatially corrected image data, and performing adaptive histogram equalization based on the spatially corrected image data to obtain illumination-balanced image data; Performing Gaussian blur filtering on the light-balanced image data to reduce noise, obtaining reduced-noise image data, and spatially aligning the reduced-noise image data to obtain spatially aligned multispectral image data; performing temporal difference comparison on the spatially aligned multispectral image data to obtain a preprocessed image data set; Superpixel segmentation and regional feature extraction are performed based on the preprocessed image data set to obtain a dirty region probability map.

3. The machine vision-based window cleaning machine movement control method according to claim 2, characterized in that: The superpixel segmentation and regional feature extraction based on the preprocessed image data set to obtain a dirty area probability map includes: Inputting the preprocessed image data set into the SLICO superpixel segmentation algorithm to perform region segmentation processing to obtain segmented image data; Extracting color features, texture features, and near-infrared reflection intensity features from each superpixel region in the segmented image data to obtain a regional feature vector, and constructing a superpixel adjacency matrix based on the segmented image data; Calculating a feature difference matrix between adjacent superpixel regions based on the region feature vector and the superpixel adjacency matrix; Performing depth convolution on the regional feature vector and the feature difference matrix to obtain regional dirt feature data, and performing probability mapping on the regional dirt feature data to obtain a dirt region probability map.

4. The machine vision-based window cleaning machine movement control method according to claim 1, characterized in that: The performing multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain the moving trajectory of the window cleaning machine includes: Inputting the dirty area probability map into the encoder of the multispectral feature fusion network to perform multi-level feature extraction to obtain a feature map set; Performing decoder upsampling on the feature map set and setting skip connections at each decoding level to obtain an initial fused feature map; The regional attention weight is calculated based on the texture complexity, brightness change and near-infrared reflectance characteristics of each superpixel region to obtain the regional attention weight matrix; Adaptively weighting the visible light features and the near infrared features according to the regional attention weight matrix to obtain preliminary fused feature data; Performing global illumination evaluation on the preliminary fused feature data and classifying them by illumination region to obtain illumination balance feature data, and performing nonlinear transformation processing of residual connection on the illumination balance feature data to obtain an enhanced region feature map; Dirty area segmentation and cleaning path planning are performed based on the enhanced area feature map to obtain the movement trajectory of the window cleaning machine.

5. The machine vision-based window cleaning machine movement control method according to claim 4, characterized in that: The dirty area segmentation and cleaning path planning based on the enhanced area feature map to obtain the movement trajectory of the window cleaning machine include: Inputting the enhanced region feature map into a conditional random field model for refinement processing to obtain a target segmentation result; Performing morphological opening processing on the target segmentation result to obtain noise-filtered segmentation data, and performing binarization processing on the noise-filtered segmentation data to obtain a dirty area mask map; Constructing a dirty area connectivity graph based on the dirty area mask graph to obtain path planning target point data, and generating a global cleaning path based on the path planning target point data; The global cleaning path is smoothed using a B-spline curve and decomposed into local path segments to obtain a movement trajectory of the window cleaning machine.

6. A mobile control device for a window cleaning machine based on machine vision, characterized in that: Used to execute the machine vision-based window cleaning machine movement control method according to any one of claims 1 to 5, the machine vision-based window cleaning machine movement control device comprising: The acquisition module is used to collect multispectral images and perform superpixel segmentation on the window surface to obtain a probability map of dirty areas; A planning module, configured to perform multispectral feature fusion and cleaning path planning on the dirty area probability map to obtain a movement trajectory of the window cleaning machine; An observation module is used to establish a bilinear system dynamics model based on the movement trajectory of the window cleaning machine and design a visual feedback state observer to obtain real-time state information of the window cleaning machine; The control calculation module is used to input the movement trajectory of the window cleaning machine and the real-time status information of the window cleaning machine into the soft constraint visual iterative regulator to perform control calculation and obtain a smooth movement control instruction.

7. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the machine vision-based window cleaning machine movement control method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the processor is enabled to execute the machine vision-based window cleaning machine movement control method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Window cleaning machine movement control method and system based on machine vision

    CN116758029A

  • Intelligent window cleaning robot capable of achieving negative-pressure adsorption, capable of easily crossing obstacles and flexible in operation

    CN117414072A