Underwater robot weak light environment target tracking method and system

By using multimodal fusion technology with visible light cameras and infrared sensors on an underwater robot, the problem of inaccurate target recognition in low-light environments has been solved. This enables accurate identification and stable tracking of debris such as fallen leaves, hair, and paper scraps, improving the automation level and efficiency of underwater cleaning operations.

CN121468601BActive Publication Date: 2026-04-14YITUO ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YITUO ELECTRIC CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Underwater robots struggle to accurately distinguish between real debris and light and shadow interference in low-light environments. Traditional enhanced light source solutions lead to increased energy consumption and decreased image quality, impacting user experience.

Method used

By simultaneously acquiring image data using an underwater robot equipped with a visible light camera and an infrared sensor, adaptive temporal fusion, contrast enhancement, and adaptive curve mapping are performed to extract surface texture and thermal distribution features of the object, generate a fused feature image, and perform target recognition and cross-frame continuous tracking.

Benefits of technology

It improves the target recognition accuracy and tracking stability of underwater robots in low-light environments, enhances the automation level and efficiency of cleaning operations, and avoids the problems of increased energy consumption and decreased image quality in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121468601B_ABST
    Figure CN121468601B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of robot control, and provides a weak light environment target tracking method and system of an underwater robot. The method comprises the following steps: synchronously collecting image data of an underwater environment; adaptively performing time sequence fusion, contrast enhancement and adaptive curve mapping processing on a plurality of frames of weak light visible light images collected continuously to obtain a low-illumination enhanced image; extracting object surface texture features from the low-illumination enhanced image and extracting thermal distribution features from infrared images collected continuously; adaptively weighting the object surface texture features and the thermal distribution features to perform fusion processing, and generating a fusion feature image; identifying and cross-frame continuously tracking a target impurity of a preset type in the fusion feature image to obtain a target tracking result; and according to the target tracking result, controlling a motion control system and a cleaning device of the underwater robot to track and capture the target impurity in real time, so that the target recognition accuracy and tracking stability in a weak light environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control technology, and in particular to a method and system for target tracking in low-light environments for underwater robots. Background Technology

[0002] To maintain water cleanliness and safety, underwater robots have become indispensable maintenance equipment in these locations. In the relevant technical field, underwater robots mainly rely on basic sensors (such as pressure sensors and collision sensors) and simple vision systems for autonomous navigation and cleaning operations.

[0003] The visual perception capabilities of underwater robots, especially their target tracking performance in low-light environments, directly impact the success rate and efficiency of task execution. Unlike open water, swimming pool environments are relatively enclosed and structurally regular, but they also suffer from uneven local lighting: shallow areas are usually well-lit, while deep water areas, corners, and areas affected by light and shadow interference from surface ripples can lead to areas of low light or drastic changes in light intensity. Furthermore, the targets for cleaning in swimming pools are mainly small, lightweight debris such as fallen leaves, hair, and paper scraps. These targets often exhibit irregular movement in the water and are difficult to distinguish from the texture and shadows of the pool bottom. Related technologies typically use fixed threshold segmentation or simple contour detection methods for target recognition, which perform reasonably well in well-lit areas, but the accuracy drops significantly in low-light environments or under conditions of strong reflection. Especially when sunlight is refracted by surface ripples to create shimmering spots, or in low-light conditions such as dusk or overcast days, traditional vision systems are prone to misinterpreting light and shadow as targets or missing actual debris.

[0004] To address the challenge of low-light environment perception, some pool cleaning robots employ enhanced lighting strategies, increasing underwater LED illumination to improve visibility. However, this approach has significant limitations: firstly, strong light sources accelerate energy consumption, shortening the duration of each operation; secondly, intense light exacerbates surface reflection and scattering effects from turbidity, reducing image contrast. Furthermore, strong light environments can interfere with the user's swimming experience, especially in night mode where excessive brightness can cause discomfort. Therefore, a technical solution is urgently needed to address at least one of these problems. Summary of the Invention

[0005] The main objective of this invention is to provide a target tracking method and system for underwater robots in low-light environments, aiming to solve the technical problems in related technologies where it is impossible to accurately distinguish between real objects and light and shadow interference in low-light environments, while traditional enhanced light source solutions lead to increased energy consumption, decreased image quality, and impaired user experience.

[0006] In a first aspect, embodiments of the present invention provide a target tracking method for an underwater robot in a low-light environment, comprising: synchronously acquiring image data of the underwater environment through a visible light camera and an infrared sensor mounted on the underwater robot, wherein the image data includes low-light visible light images and infrared images;

[0007] Adaptive temporal fusion, contrast enhancement, and adaptive curve mapping are performed on continuously acquired multi-frame low-light visible light images to obtain low-light enhanced images.

[0008] The surface texture features of the object are extracted from the low-light enhanced image, and the thermal distribution features are extracted from the continuously acquired infrared images. The surface texture features and thermal distribution features are fused using adaptive weights to generate a fused feature image. The adaptive weights are dynamically adjusted based on the edge morphology and thermal characteristics of the target debris in the pool environment.

[0009] The target debris of a preset type in the fused feature image is identified and continuously tracked across frames to obtain the target tracking result; the types of target debris include fallen leaves, hair and paper scraps.

[0010] Based on the target tracking results, the motion control system and cleaning device of the underwater robot are controlled to track and grab the target debris in real time.

[0011] Secondly, embodiments of the present invention provide a target tracking system for an underwater robot in a low-light environment, comprising: a data acquisition module, used to simultaneously acquire image data of the underwater environment through a visible light camera and an infrared sensor mounted on the underwater robot, wherein the image data includes low-light visible light images and infrared images;

[0012] The enhancement module is used to perform adaptive temporal fusion, contrast enhancement, and adaptive curve mapping processing on continuously acquired multi-frame low-light visible light images to obtain low-light enhanced images.

[0013] The fusion module is used to extract surface texture features of objects from the low-light enhanced image, extract thermal distribution features from continuously acquired infrared images, and fuse the surface texture features and thermal distribution features through adaptive weights to generate a fused feature image; the adaptive weights are dynamically adjusted according to the edge morphology characteristics and thermal characteristics differences of target debris in the pool environment.

[0014] The tracking module is used to identify and continuously track target debris of a preset type in the fused feature image to obtain target tracking results; the types of target debris include fallen leaves, hair and paper scraps; based on the target tracking results, the motion control system and cleaning device of the underwater robot are controlled to track and grasp the target debris in real time.

[0015] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device including a processor and a memory for storing a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the underwater robot target tracking method in a low-light environment as described in the first aspect or any embodiment of the present invention.

[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer software program, which, when executed by a processor, implements the underwater robot target tracking method in a low-light environment as described in the first aspect or any embodiment of the present invention.

[0017] This invention provides a method and system for target tracking in low-light environments using an underwater robot. In this embodiment, an underwater robot simultaneously acquires image data of the underwater environment using a visible light camera and an infrared sensor. The image data includes low-light visible light images and infrared images. Then, adaptive temporal fusion, contrast enhancement, and adaptive curve mapping are performed on the continuously acquired multi-frame low-light visible light images to obtain a low-light enhanced image. Next, surface texture features are extracted from the low-light enhanced image, and thermal distribution features are extracted from the continuously acquired infrared images. The surface texture features and thermal distribution features are fused using adaptive weights to generate a fused feature image. The adaptive weights are dynamically adjusted based on the edge morphology and thermal characteristics of target debris in the pool environment. Then, target debris of a preset type in the fused feature image is identified and continuously tracked across frames to obtain target tracking results. The types of target debris include fallen leaves, hair, and paper scraps. Finally, based on the target tracking results, the underwater robot's motion control system and cleaning device are controlled to track and grasp the target debris in real time. This invention, through the fusion of low-light visible and infrared images, and the identification and continuous cross-frame tracking of target debris, helps overcome visual perception barriers in low-light underwater environments. It improves the accuracy and tracking stability of underwater robots for small floating debris, and solves the technical problems of high false negative rates, tracking interruptions, and grasping failures caused by poor image quality and unclear target features in traditional underwater cleaning equipment under low-light conditions. Furthermore, through a multimodal perception fusion mechanism and an adaptive target tracking strategy, this invention achieves accurate identification and stable tracking of irregular targets such as fallen leaves, hair, and paper scraps in complex underwater environments, effectively improving the automation level and work efficiency of underwater cleaning operations. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a target tracking method for an underwater robot in a low-light environment, as provided in an embodiment of the present invention. Detailed Implementation

[0019] In the field of underwater robotics, ambient lighting conditions are one of the key factors affecting the performance of vision systems. Low-light environments refer to visual scenes with light intensity below 100 lux. In such environments, traditional vision systems struggle to acquire image data with sufficient signal-to-noise ratio and detail. In underwater swimming pool applications, low-light environments are primarily caused by a combination of factors: First, the absorption of light by water decreases exponentially with depth. Generally, visible light intensity decreases for every meter of descent, resulting in the light intensity at the pool bottom (typically 1.5-3 meters deep) being only 30%-50% of that at the surface. Second, suspended particulate matter in the water (including disinfection byproducts, algae, and human metabolites) causes light scattering, reducing image contrast and clarity. Third, the refraction and reflection of incident light by the water surface, as well as factors such as building shadows, weather conditions, and diurnal variations, further exacerbate the uneven distribution of underwater lighting. In particular, in the underwater environment of a swimming pool, the attenuation rate of different wavelengths of light varies. The red light band (600-700nm) attenuates the fastest, almost disappearing completely at a depth of 1.5 meters, while the blue-green light band (450-550nm) is relatively well preserved, causing severe color distortion. This makes it difficult for the vision system to accurately identify common debris such as fallen leaves, hair, and paper scraps. Images in such low-light environments generally suffer from insufficient brightness, high noise, blurred details, and color distortion, directly leading to a decrease in the target detection accuracy of traditional underwater robots, an increase in the frequency of tracking interruptions, and consequently, a serious impact on the efficiency and automation level of cleaning operations.

[0020] The following detailed description, with reference to the accompanying drawings, addresses some embodiments of the present invention and the technical problems encountered by underwater robots in low-light environments. Figure 1 This is a flowchart illustrating a target tracking method for an underwater robot in low-light environments, provided as an embodiment of the present invention. Please refer to... Figure 1 The method includes the following steps:

[0021] Step S101: Simultaneously collect image data of the underwater environment using the visible light camera and infrared sensor mounted on the underwater robot.

[0022] In one embodiment of the invention, the underwater robot refers to an autonomous or semi-autonomous underwater operating device specifically designed for pool cleaning tasks. This underwater robot has a streamlined, sealed body and integrates a buoyancy adjustment mechanism, a motion control system, a cleaning device, a central control unit, and an energy supply module. Its outer shell is made of corrosion-resistant and pressure-resistant engineering materials, ensuring long-term stable operation in underwater environments. This underwater robot can move flexibly on the bottom and sidewalls of the pool, possessing autonomous navigation, obstacle recognition, and target tracking capabilities, and is particularly optimized for operation in low-light conditions.

[0023] A visible light camera is an optical imaging device specifically optimized for underwater environments. It employs a high-sensitivity image sensor and a special optical lens to capture environmental images within the visible spectrum under extremely low light conditions. This visible light camera features a waterproof, sealed structure and is equipped with automatic white balance and color correction functions to effectively address underwater light attenuation and color shift issues.

[0024] Infrared sensors are imaging devices that can sense the distribution of thermal radiation from objects. Their working principle is based on the differences in infrared radiation emitted by objects at different temperatures. They are not limited by visible light illumination conditions and can acquire thermal distribution information of a scene in completely dark environments. The two sensors work in coordination through a synchronous triggering mechanism to ensure that the acquired image data are strictly correlated in time.

[0025] In this embodiment, the low-light visible light image refers to an optical image captured by a visible light camera in a dimly lit underwater swimming pool environment. Due to the absorption and scattering of light by water, especially with increasing depth, this image typically exhibits insufficient brightness, low contrast, color distortion, and significant noise. The infrared image, on the other hand, refers to an image obtained by an infrared sensor that reflects the surface temperature distribution of objects in a scene. It is unaffected by visible light illumination conditions and can present the thermal characteristics of objects, making it particularly suitable for target identification in completely dark or highly turbid underwater environments. Thus, the low-light visible light image preserves the texture and shape details of the target, while the infrared image provides the target's thermal distribution characteristics. The combination of the two can complement each other to improve the accuracy and robustness of target debris identification in complex underwater environments.

[0026] Step S102 involves adaptive temporal fusion, contrast enhancement, and adaptive curve mapping processing on multiple consecutively acquired low-light visible light images to obtain low-light enhanced images. Therefore, step S102 involves enhancing multiple consecutively acquired images in a low-light underwater swimming pool environment to obtain high-quality low-light enhanced images, optimizing brightness distribution, maintaining color consistency, and eliminating over-enhancement artifacts. This processing flow aims to overcome the image quality degradation caused by low-light underwater environments, particularly addressing challenges common in swimming pool environments such as uneven illumination, suspended particle interference, and color distortion.

[0027] As an optional embodiment, in step S102, for continuously acquired multi-frame low-light visible light images, the pixel difference and structural similarity between adjacent frames are calculated, and adaptive fusion weights are assigned to each frame of low-light visible light images to obtain a temporally fused low-light visible light image, thereby suppressing random noise and enhancing stable target features; local contrast enhancement is performed on the temporally fused low-light visible light image to obtain a contrast-enhanced image, thereby adaptively adjusting the contrast parameters of local regions in the low-light visible light image to enhance the distinction between target clutter and background; curve parameters are dynamically adjusted according to the brightness distribution characteristics of the contrast-enhanced image, and adaptive curve mapping processing is performed on the contrast-enhanced image using curve parameters and a no-reference learning method to optimize the overall brightness distribution of the image while maintaining the original color consistency, thereby obtaining the low-light enhanced image. Further optionally, the adaptive curve mapping processing uses a no-reference learning method, which does not rely on paired low / normal illumination training samples, effectively avoiding artifacts and color distortion caused by over-enhancing.

[0028] Specifically, in the temporal fusion stage, the system first performs inter-frame relationship analysis on multiple consecutively acquired low-light visible light images. For any two adjacent frames, a pixel-level difference metric is calculated, reflecting the degree of brightness and chromaticity changes of corresponding pixels between the two frames. Simultaneously, a structural similarity index is calculated, assessing the consistency of structural features such as edges and textures in the two frames. Based on these two metrics, the system assigns adaptive fusion weights to each frame. Regions with stable structural features but random fluctuations in pixel values ​​are assigned higher weights. Regions with significant structural feature changes (such as areas with fast-moving objects or water flow disturbances) are assigned lower weights. In a swimming pool environment, this weight allocation mechanism effectively enhances the feature representation of target debris such as fallen leaves and hair, while suppressing random noise caused by water ripples and suspended particles. For example, when the robot inspects the bottom of the pool, static features such as the texture of the pool bottom tiles and attached dirt are enhanced, while floating microbubbles and suspended matter are suppressed, improving the signal-to-noise ratio of the image.

[0029] Optionally, the pixel difference and structural similarity between adjacent frames are calculated. When assigning adaptive fusion weights to each frame of the low-light visible light image, the pixel difference is determined by comparing the numerical differences of corresponding pixels in adjacent frames. Specifically, for each group of adjacent frames (e.g., frame i and frame i+1), the average of the squared differences of all pixels is calculated to obtain the pixel difference of that group of adjacent frames. Structural similarity is quantitatively evaluated using the Structural Similarity Index (SSIM), which comprehensively considers the similarity between two frames in three dimensions: brightness, contrast, and structural features. When assigning adaptive fusion weights to each frame of the low-light visible light image, the weight of the first frame is directly set based on its structural similarity value with the second frame, the weight of the last frame is directly set based on its structural similarity value with the second-to-last frame, and the weight of the intermediate frames (i.e., frames 2 to N-1) is calculated based on the arithmetic mean of their structural similarity values ​​with the previous and next frames. Here, N is the total number of frames. After all weight values ​​are normalized (i.e., each weight value is divided by the sum of all weight values ​​to ensure that the sum of normalized weights is 1), a weight allocation scheme positively correlated with structural stability is formed: frames with higher structural similarity values ​​(indicating stable regional structures, such as stationary objects) receive higher fusion weights, while frames with lower structural similarity values ​​(indicating large changes in regional structures, such as fast-moving objects or water flow disturbances) receive lower fusion weights. This weight allocation mechanism enables the temporal fusion process to dynamically suppress random noise (such as pixel fluctuations) in low-light environments while enhancing stable target features (such as stationary objects), thereby generating high-quality temporally fused low-light visible light images.

[0030] In the local contrast enhancement stage, a spatially adaptive method is used to optimize contrast for the temporally fused image. Unlike traditional global contrast adjustment, this embodiment divides the image into multiple local regions and analyzes the grayscale distribution characteristics of each region independently. The optimal contrast enhancement parameters are dynamically determined based on the texture complexity, edge density, and contrast relationship with surrounding regions of each local region. In a swimming pool environment, this local adaptive processing is particularly effective in enhancing the distinction between different debris and the background. For example, for translucent paper scraps, it enhances their contrast with the water while preserving their edge details. For dark fallen leaves, it improves their distinction from the pool bottom without losing detail. This process ensures that the enhanced image has both high contrast and retains important detail information by protecting weak edges and suppressing noise amplification.

[0031] In the adaptive curve mapping stage, the overall brightness distribution characteristics of the contrast-enhanced image are analyzed, including the brightness histogram shape, dark / bright ratio, midtone distribution, and other features, and a nonlinear mapping curve is dynamically constructed. This curve mapping process employs a no-reference learning mechanism, not relying on pre-acquired low-light / normal-light image pairs as training samples, but optimizing based on the image's own visual perception quality indicators. A series of no-reference quality evaluation indicators are defined, including naturalness, information richness, and color harmony. By iteratively optimizing the mapping parameters, the output image achieves an optimal balance across these indicators. For example, naturalness is quantified by calculating the mean squared error (MSE) of pixel differences between the enhanced image and the original image in local regions. If the MSE value is lower than a preset threshold (e.g., 0.03), the image is considered to have high naturalness (indicating no obvious artifacts or over-enhancement). Information richness is calculated based on the local entropy value of the image. The image is divided into multiple small regions, and the pixel gray-level distribution entropy of each region is calculated (a higher entropy value indicates richer details). The entropy values ​​of all regions are then averaged to obtain the overall information richness indicator. Color harmony is evaluated by analyzing the uniformity of hue, saturation, and lightness distribution in the HSV color space. The weighted average of the standard deviations of hue, saturation, and lightness (weights of 0.4, 0.3, and 0.3 respectively) is calculated across the entire image. A smaller standard deviation indicates higher color harmony (i.e., smoother color transitions without abrupt distortion). All calculations are dynamically performed based on the image's own content, requiring no external reference data, ensuring that the evaluation results objectively reflect the image enhancement quality and effectively guide the optimization of the low-light enhancement process. In underwater swimming pool environments, this method effectively handles various complex lighting conditions, such as flickering light spots caused by water ripples, localized highlights formed by pool wall reflections, and shadows on the pool bottom, generating images with uniform brightness, rich detail, and natural colors. Especially for the common blue-green bias problem in underwater environments, this method can effectively correct color bias while enhancing brightness, restoring the true color characteristics of target debris, such as the brownish-yellow of fallen leaves and the dark brown of hair, which is crucial for subsequent target classification and recognition.

[0032] Optionally, in step S102, when dynamically adjusting the curve parameters based on the brightness distribution characteristics of the contrast-enhanced image, the statistical features of the image brightness histogram are first extracted, including the brightness mean, standard deviation, and brightness range distribution. Based on the brightness mean (if it is lower than a preset threshold, such as 50 / 255, it indicates that the image is generally dark), the Gamma parameter value is automatically increased (e.g., gradually increasing from 1.0 to 1.8) to enhance details in dark areas. Based on the brightness standard deviation (if it is higher than a preset threshold, such as 30 / 255, it indicates that the contrast is high), the Gamma parameter value is automatically decreased (e.g., gradually decreasing from 1.0 to 0.6) to avoid overexposure and artifacts. This parameter adjustment process adopts a no-reference learning method, that is, the Gamma parameter is dynamically calculated only by analyzing the brightness distribution characteristics of the current image itself (without relying on paired low / normal lighting training samples), ensuring that the parameter settings adaptively match the image content. Subsequently, the optimized Gamma parameter is used to perform adaptive curve mapping processing on the contrast-enhanced image. By adjusting pixel brightness through the Gamma correction function, the overall brightness distribution of the image converges to a range that is more in line with human visual perception (such as enhancing details in dark areas and preventing overexposure in bright areas). At the same time, by maintaining the relative proportions between RGB channels (such as fixing the relative differences between R / G / B channels), the consistency of the original colors is ensured, ultimately generating a low-light enhanced image that effectively suppresses brightness distortion and color deviation in low-light environments.

[0033] For example, in the adaptive curve mapping processing stage, a no-reference learning mechanism is used to dynamically adjust the curve parameters. This mechanism does not rely on paired low / normal lighting training samples, but rather optimizes based on visual quality metrics of the image's own content. Specifically, the brightness distribution characteristics of the contrast-enhanced image are first extracted, including key features such as the shape of the brightness histogram, the ratio of dark to bright areas, and the distribution of midtones. Subsequently, a nonlinear mapping function containing a Gamma parameter is constructed based on these features. This function adjusts the Gamma value through an iterative optimization process, enabling the output image to achieve an optimal balance across three key no-reference quality evaluation metrics.

[0034] The naturalness index is quantified by calculating the mean squared error (MSE) of pixel differences between the enhanced and original images in local regions. If the MSE value is lower than a preset threshold (e.g., 0.03), the image is considered to have high naturalness, indicating no obvious artifacts or over-enhancement. Information richness is calculated based on the local entropy value of the image. The image is divided into multiple small regions (e.g., 16×16 pixel blocks), and the pixel grayscale distribution entropy of each region is calculated (a higher entropy value indicates richer details). The entropy values ​​of all regions are then averaged to obtain the overall information richness index. Color harmony is evaluated by analyzing the uniformity of hue, saturation, and brightness distribution in the HSV color space. The weighted average of the standard deviations of hue, saturation, and brightness (with weights of 0.4, 0.3, and 0.3 respectively) is calculated for the entire image. A smaller standard deviation indicates higher color harmony, meaning smooth color transitions without abrupt distortion.

[0035] Optionally, gradient descent can be used for iterative optimization. The initial Gamma parameter is set to 1.0, and then the scores of the three quality metrics of the image at the current Gamma value are calculated. By calculating the gradient of each metric with respect to the Gamma parameter, the direction and magnitude of the Gamma parameter adjustment are determined to maximize the weighted sum of the three metrics: naturalness, information richness, and color harmony. For example, by calculating the gradients of naturalness, information richness, and color harmony with respect to the Gamma parameter, the Gamma value is iteratively adjusted along the upward direction of the gradient: if the gradient is positive, Gamma is increased; if it is negative, Gamma is decreased. The adjustment magnitude is controlled by an adaptive learning rate (e.g., an initial step size of 0.1, decaying by 10% each round) until the weighted sum of the three metrics reaches its maximum value. This process is completed within 5-10 iterations, ensuring that the output image achieves a dynamic balance in brightness detail, information richness, and color harmony without the need for external reference data. Furthermore, during the optimization process, the weights of each metric are dynamically balanced; the specific adjustment method is described below. When the image is generally dark, more attention is paid to improving information richness. When an image exhibits significant color distortion, optimization of color harmony is prioritized. When artifacts caused by over-enhancement appear, the weight of the naturalness metric is increased. Through multiple iterations (typically 5-10 times), the Gamma parameter value that maximizes the overall score of the three metrics is found. This ensures that the output image achieves the optimal balance in brightness distribution, detail preservation, and color consistency, effectively avoiding artifacts and color distortion problems caused by over-enhancement in traditional methods. This provides high-quality input images for subsequent feature extraction and object recognition.

[0036] Optionally, during the optimization process, when dynamically balancing the weights of various indicators, the weights of the three quality indicators—naturalness, information richness, and color harmony—are dynamically adjusted by analyzing the image content in real time. The specific implementation is as follows: First, calculate the mean luminance (μ) and naturalness index (MSE) of the current contrast-enhanced image. If the mean luminance μ is lower than a preset threshold of 50 (based on 255 gray levels), the weight of the information richness index is dynamically increased from the base value of 0.3 to 0.5. If μ is higher than 150, the weight of the color harmony index is increased from the base value of 0.3 to 0.5. Simultaneously, the naturalness index is detected: if the MSE value is higher than a preset threshold of 0.03 (indicating the presence of artifacts), the naturalness weight is increased from the base value of 0.4 to 0.6. If the MSE is lower than 0.03, the weight remains at 0.4. For the color harmony index, calculate the weighted standard deviation (Ci) of hue, saturation, and lightness in the HSV space. c If C c If the value exceeds the preset threshold of 5 (indicating severe color distortion), the color harmony weight is reduced from the base value of 0.3 to 0.2. If C... c If the value is below 5, the weight is increased to 0.4. The adjusted weights are then normalized (ensuring the sum is 1) to form a dynamic weight vector (e.g., when the image is dark and MSE > 0.03, the weights might be [0.5, 0.6, 0.2]). During optimization, the Gamma parameter is iteratively adjusted using gradient descent. In each iteration, the weighted score of the three indicators at the current Gamma value is calculated (score = w). natural ×(1 - MSE) + w info ×entropy avg + w color ×(1 / C c ), and update the Gamma value along the gradient ascent direction. Where, w natural This is the naturalness weight, with a base value of 0.4; w color This is the information richness weight, with a base value of 0.3; w info This is the color harmony weight, with a base value of 0.3. The sum of these three weights is 1. avg The information richness metric is the average information entropy of an image, reflecting the detail and texture complexity of the image; entropy avg A larger value indicates more information. For example, entropy avg The average is taken after calculation in the grayscale or gradient domain. c It is one of the indicators of color coordination. c It is the comprehensive color dispersion obtained by calculating the standard deviations of hue (H), saturation (S), and value (V) in the HSV color space and then summing them by weight. cThe higher the value, the more uncoordinated the color distribution (e.g., color cast, oversaturation). The above mechanism ensures that information richness is prioritized in low-light images (avoiding loss of detail), color harmony is prioritized in color-distorted images (avoiding color cast), and naturalness is prioritized in images with severe artifacts (avoiding over-enhancement). This achieves a globally optimal balance of the three metrics within 5-10 iterations, generating a low-light enhanced image that retains detail and is free of artifacts.

[0037] In practical pool cleaning tasks, underwater robots operating in low-light conditions can improve the input quality of the vision system. For example, in outdoor pools at dusk or in poorly lit indoor pools, traditional systems often struggle to identify small debris. However, with the processing flow of this embodiment, even in severely low-light environments, the vein structure of fallen leaves, the fine morphology of hairs, and the edge features of paper scraps can still be clearly presented. This improvement in image quality directly translates into higher target detection accuracy and more stable cross-frame tracking performance, effectively solving the technical challenge of underwater robots' inability to identify objects in low-light environments, and providing a solid foundation for the reliable implementation of automatic cleaning functions. Compared to traditional single-frame processing or simple brightness enhancement methods, this embodiment maintains image naturalness while reducing artifacts, halos, and color distortion caused by over-enhancement, significantly improving the robustness and accuracy of subsequent processing stages.

[0038] Through the coordinated processing of these three stages, step S102 can generate a low-light enhanced image that is optimized in terms of brightness distribution, contrast performance and color consistency, providing high-quality visual input for subsequent multimodal feature fusion and target tracking. It is particularly suitable for the accurate identification and tracking of small debris such as fallen leaves, hair and paper scraps in the underwater environment of a swimming pool.

[0039] Step S103: Extract the surface texture features of the object from the low-light enhanced image, extract the thermal distribution features from the continuously acquired infrared images, and perform fusion processing on the surface texture features and thermal distribution features of the object through adaptive weighting to generate a fused feature image.

[0040] This step fully utilizes the complementary characteristics of visible light and infrared information in the underwater environment, and specifically constructs an adaptive fusion mechanism to address the differences in the physical characteristics of common debris in swimming pool environments, thereby improving the robustness of target recognition under complex underwater lighting conditions.

[0041] In this embodiment of the invention, the adaptive weights are dynamically adjusted based on the differences in edge morphology and thermal properties of target debris in the swimming pool environment. That is, the fusion ratio of texture features and thermal distribution features is dynamically adjusted according to the differences in edge morphology and thermal properties of target debris in the swimming pool environment. Specifically, the extracted texture features and thermal distribution features are first evaluated for feature saliency. For texture features, evaluation indicators include edge sharpness, structural complexity, and local contrast. For thermal distribution features, evaluation indicators include thermal boundary sharpness, thermal contrast, and thermal distribution uniformity. Then, the system dynamically calculates the reliability weights of the two features based on the pool environment's lighting conditions, water quality, and prior knowledge of the target type. In shallow water areas with good lighting conditions, texture features are generally more reliable, and their fusion weight is automatically increased. In deep water areas with insufficient lighting or turbid water, the reliability of thermal distribution features is relatively increased, and their weight ratio is increased accordingly. Furthermore, the weight allocation strategy differs for different types of debris. For synthetic paper scraps with obvious thermal properties, thermal distribution features are given a higher weight. For fallen leaves with obvious morphological characteristics, the weight of texture features is increased accordingly. For long, thin hairs, the system employs a balanced weighting strategy, while also utilizing their characteristic behavior in both modalities.

[0042] In real-world pool cleaning scenarios, such as when an underwater robot is working in an outdoor pool at dusk, the quality of visible light imaging drops sharply, but floating leaves still maintain a slight temperature difference with the water. In this case, the weight of texture features is automatically reduced while the weight of thermal distribution features is increased, allowing the identification of leaves that are almost invisible in the visible light image. Conversely, in a clear indoor pool, when identifying paper tags attached to the pool bottom, because their thermal properties are similar to those of the pool bottom tiles but their textures are significantly different, the fusion weight of texture features can be automatically increased to locate these small targets. Similarly, when processing suspended hair, the elongated morphological characteristics of hair under visible light and its linear thermal distribution characteristics in infrared images can be comprehensively utilized. Through balanced weight allocation, stable identification of these deformable and difficult-to-track targets can be achieved.

[0043] As an optional embodiment, in step S103, the low-light enhancement image is subjected to multi-scale extraction using a multi-scale convolutional filter bank to extract edge, corner, and surface texture features at different scales, constructing multi-scale texture features; thermal contrast enhancement processing is performed on continuously acquired infrared images, and heat source distribution features are extracted through gradient domain analysis to generate a thermal distribution feature map; the feature saliency index of the target region in the multi-scale texture features and the thermal distribution feature map is calculated, the feature saliency index including edge sharpness, texture complexity, thermal contrast, and region consistency parameters; based on the feature saliency index, fusion weights are dynamically assigned to the texture features and thermal distribution features at different scales; according to the dynamically assigned fusion weights, the multi-scale texture features and the thermal distribution feature map are weighted and fused, and the fused feature image is generated through a feature reconstruction network. The fused feature image simultaneously retains the clear texture details and thermal distribution characteristics of the target debris, enhancing the robustness of identifying fallen leaves, hair, and paper scraps in low-light environments.

[0044] Optionally, based on the feature saliency index, when dynamically assigning fusion weights to texture features and thermal distribution features at different scales, the weight of thermal distribution features is increased in low-light areas, and the weight of texture features is increased in relatively well-lit areas. Specifically, if the thermal contrast index value is higher than a preset threshold (e.g., 0.7, indicating a significant difference in brightness between the heat source and the background), the area is determined to be a low-light area, and the fusion weight of the thermal distribution feature is dynamically increased to a higher level (e.g., from the base weight of 50% to over 70%). If the texture complexity index value is higher than a preset threshold (e.g., 0.5, indicating rich texture details and clear edges), the area is determined to be a relatively well-lit area, and the fusion weight of the texture features is dynamically increased to a higher level (e.g., from the base weight of 50% to over 70%). For areas that simultaneously meet both conditions, the relative magnitudes of thermal contrast and texture complexity are compared first. When thermal contrast > texture complexity, the weight of thermal distribution features is increased. Conversely, the weight of texture features is increased. All weights are normalized (ensuring the sum is 1) to ensure that thermal distribution characteristics are preferentially preserved in low-light areas (such as the thermal signal of fallen leaves in the dark underwater area) and texture features are preferentially preserved in well-lit areas (such as the texture details of paper scraps floating on the water surface), thereby improving the robustness of identifying target debris such as fallen leaves, hair and paper scraps in low-light environments.

[0045] In another embodiment, when dynamically assigning fusion weights to texture features and thermal distribution features at different scales based on the aforementioned feature saliency index, feature saliency index values ​​can also be calculated for the target region. Edge sharpness is determined by evaluating the sharpness and continuity of edges in the texture features (the sharper the edges, the higher the value). Texture complexity is determined by analyzing the richness and diversity of local details in the texture features (the more complex the texture, the higher the value). Thermal contrast is determined by quantifying the brightness difference between the target heat source and the background in the thermal distribution feature map (the higher the contrast, the higher the value). Region consistency is determined by detecting the degree of consistency of features within the target region (the higher the consistency, the higher the value). Subsequently, for thermal distribution features, their weight is jointly determined by the thermal contrast and region consistency parameters. When thermal contrast is high (e.g., significant heat source in low light) and region consistency is high (e.g., stable thermal distribution), the weight of the thermal distribution feature is automatically increased to a higher level. For texture features, their weight is jointly determined by the edge sharpness and texture complexity parameters. When edge sharpness is high (e.g., distinct texture edges in sufficient lighting) and texture complexity is high (e.g., rich target texture), the weight of the texture feature is automatically increased to a higher level. The weight allocation strategy is intelligently adjusted based on the current ambient lighting conditions. In low-light areas (where thermal contrast is dominant), the weight of thermal distribution features is significantly increased (e.g., from a base weight of 50% to over 70%). In relatively well-lit areas (where texture complexity is dominant), the weight of texture features is significantly increased (e.g., from a base weight of 50% to over 70%). Finally, all weights are normalized (ensuring that the sum of the weights of texture features and thermal distribution features is 1), achieving a weighted fusion of multi-scale texture features and thermal distribution features. This makes the fused feature image more prominent in low-light environments (such as the thermal signal of fallen leaves) and more prominent in well-lit areas (such as the fibrous structure of paper scraps), thereby improving the robustness of target object recognition.

[0046] First, the filters exhibit diversity in scale, orientation, and frequency response, enabling them to simultaneously capture both the microscopic details and macroscopic structure of target objects. Specifically, the filter bank comprises multiple scale levels, each focusing on feature extraction at a specific level of detail: smaller-scale filters capture fine-grained surface textures, such as the vein structure of fallen leaves or the fiber texture of paper scraps; medium-scale filters extract edge and contour information, identifying the outer boundaries of target objects; and larger-scale filters focus on overall shape and spatial layout features, aiding in understanding the positional relationships of targets within the scene. This multi-scale analysis is particularly important in underwater pool environments because different types of debris exhibit the most distinctive features at different scales. For example, hair appears as a thin, linear structure at small scales, but as a continuous curved shape at large scales; fallen leaves display unique irregular edge features at medium scales, while showing detailed leaf texture at small scales. Through this hierarchical feature extraction, a rich multi-scale texture feature pyramid is constructed, providing a comprehensive visual information foundation for subsequent adaptive fusion.

[0047] In the thermal distribution feature extraction stage, the continuously acquired infrared image sequence is first subjected to thermal contrast enhancement processing. Due to the high thermal conductivity of water, the temperature difference between underwater objects and their surrounding environment is usually small, resulting in low contrast in the original infrared images. Thermal contrast enhancement processing adaptively adjusts the dynamic range of the thermal images, enhancing the visual representation of subtle temperature differences and making the thermal characteristics of target debris more prominent. Subsequently, the spatial gradient of temperature changes in the thermal images is calculated to identify thermal boundaries and thermal anomaly regions. This gradient analysis effectively highlights objects with unique thermal distribution patterns, such as synthetic paper scraps with a large temperature difference from the water, or organic debris with specific thermal conductivity characteristics. In actual swimming pool environments, debris of different materials exhibits drastically different thermal characteristics. For example, plastic paper scraps typically maintain a large temperature difference with the water, appearing as distinct bright or dark areas in the thermal images; natural fallen leaves, due to water saturation, have a smaller temperature difference with the water, but may form specific thermal gradient distributions at the edges; hair, due to its slender shape and low heat capacity, usually appears as thin linear temperature anomalies in the thermal images. By combining thermal contrast enhancement with gradient analysis, this method can generate a thermal distribution feature map that clearly describes the thermal distribution characteristics of the target, providing key thermal information for multimodal fusion.

[0048] In the feature saliency assessment and weight allocation stage, various saliency indices for the target region in multi-scale texture and thermal distribution feature maps are calculated, and fusion weights are dynamically determined based on these indices. Edge sharpness assesses the clarity of the target boundary in both modalities, texture complexity measures the richness of surface details, thermal contrast reflects the temperature difference between the target and the background, and region consistency evaluates the spatial coherence of features. Based on a comprehensive analysis of these indices, optimal fusion weights are assigned to texture and thermal distribution features at different scales. In deep water areas of a swimming pool with poor lighting conditions or during dusk, thermal distribution features are generally more reliable, and their fusion weight is automatically increased; in shallow water areas with relatively sufficient lighting or during daytime, texture features are more reliable, and their weight ratio is correspondingly increased. Furthermore, the weight allocation strategy is dynamically adjusted according to water quality conditions (such as turbidity): visible light texture features have a higher weight under clear water conditions; thermal distribution features have a correspondingly higher weight under turbid water conditions. For example, when an underwater robot detects fallen leaves in murky water at night, it will significantly increase the weight of thermal distribution features because the texture information in the visible light image is severely degraded at this time, while the small temperature difference between the fallen leaves and the water can still be identified in the thermal image. Conversely, when identifying paper scraps at the bottom of a pool in clear water during the day, it will prioritize texture features because the visible light image can clearly show the shape and edge details of the paper scraps.

[0049] In the feature fusion and reconstruction stage, multi-scale texture features and thermal distribution features are fused layer by layer according to dynamically assigned fusion weights. This process is not a simple linear superposition, but rather employs feature space alignment and contextual information integration strategies to ensure that the features of the two modalities maintain consistency at both spatial and semantic levels. The fused features are then optimized through a lightweight feature reconstruction network, which is specially trained to retain the key discriminative features of the target clutter while suppressing intermodal noise interference. The final fused feature image retains both the clear texture details of the target clutter and integrates its thermal distribution characteristics, providing a high-quality feature representation for subsequent target detection and tracking.

[0050] In practical applications, for example, when an underwater robot is working in an outdoor swimming pool at dawn, the light on the water surface is weak and unevenly distributed, making it difficult for traditional vision systems to distinguish between floating leaves and water ripples. However, the multimodal fusion method used in this embodiment can simultaneously utilize the specific shape of leaves under visible light and their thermal distribution characteristics in infrared images to accurately identify and track these targets. Similarly, when dealing with hair tangled around a drain in a pool corner, due to lighting, shadows, and structural complexity, a single visible light modality can easily miss detections. However, by combining thermal distribution characteristics, the slender shape and thermal properties of the hair can be clearly identified, achieving precise positioning. Furthermore, for tiny pieces of paper floating near the water surface, they may be difficult to identify in visible light images due to water reflection. However, in thermal images, due to the difference in heat capacity between the paper and the water, adaptive weight allocation can effectively fuse the information from both modalities, achieving stable detection.

[0051] Step S104: Identify and continuously track target debris of a preset type in the fused feature image across frames to obtain target tracking results. In this embodiment of the invention, the types of target debris include fallen leaves, hair, and paper scraps. Thus, step S104 achieves accurate identification and stable tracking of specific types of target debris in the underwater swimming pool environment. This step fully utilizes the fused feature image generated in the preceding steps, combined with multi-scale feature analysis and intelligent tracking algorithms, effectively solving technical challenges such as unstable target detection and difficulty in cross-frame correlation in the underwater environment.

[0052] As an optional embodiment, in step S104, multi-scale feature extraction is performed on the fused feature image to generate feature maps of different resolutions; target detection is performed based on the feature maps of different resolutions to identify target debris of a preset type; a unique identifier is assigned to each detected target debris, and the target debris is continuously tracked across frames by combining the similarity of target appearance features and motion trajectory prediction; when the detected target debris is a fallen leaf, hair or paper scrap, a target tracking result containing the position, direction of motion and speed information of the target debris is generated.

[0053] In the above embodiments, during the multi-scale feature extraction stage, the fused feature image undergoes multi-level analysis to generate a set of feature maps with different spatial resolutions and semantic abstractions. This process employs a feature pyramid structure combining top-down and bottom-up approaches to ensure the simultaneous capture of both detailed target information and contextual semantics. The bottom-up path extracts high-semantic-level features through progressive downsampling, which helps in understanding the overall category attributes of the target; the top-down path, through upsampling and lateral connections, fuses high-level semantic information with low-level detailed features, enhancing spatial positioning accuracy. This multi-scale feature representation is particularly important in the underwater environment of a swimming pool because different types of debris exhibit the most distinctive visual patterns at different scales. For example, fallen leaves exhibit unique irregular contours and internal texture structures at a medium scale, hair appears as long, continuous lines in high-resolution feature maps, while paper scraps exhibit relatively regular geometric shapes at multiple scales. Through this hierarchical feature extraction, comprehensive and rich visual representations can be provided for subsequent target detection, effectively addressing the interference of complex factors such as changes in lighting, partial occlusion, and perspective changes in the underwater environment.

[0054] In the target detection phase, collaborative detection is performed based on feature maps of different resolutions to identify pre-defined types of debris targets. Specifically, high-resolution feature maps are mainly used to detect small debris (such as small pieces of paper and short hairs), medium-resolution feature maps are suitable for medium-sized targets (such as ordinary fallen leaves and clumps of hair), and low-resolution feature maps are responsible for capturing large or distant targets. Specially optimized target detectors are deployed on each resolution feature map. These detectors are custom-trained for the underwater pool environment and exhibit high specificity for the three target categories: fallen leaves, hair, and paper scraps. During detection, multi-scale detection results are comprehensively considered. A non-maximum suppression algorithm is used to eliminate duplicate detections, and reliable results are selected based on a confidence threshold. In particular, an environmental adaptive mechanism is introduced to dynamically adjust detection parameters according to current underwater lighting conditions, water turbidity, and other environmental parameters. For example, under low-light conditions, the confidence threshold is lowered, but the spatial continuity requirement is increased. Under high turbidity conditions, the detection results rely more heavily on thermal distribution characteristics. This adaptive mechanism improves the robustness of detection in various complex pool environments and effectively reduces false positives and false negatives.

[0055] In the target identity management and cross-frame tracking stage, each newly detected target object is assigned a globally unique identifier, and a complete target state profile is established. This profile records basic information such as the target's initial position, size, category, detection confidence, and first appearance time. Subsequently, the identified targets are continuously tracked, a process that integrates two mechanisms: appearance feature similarity matching and motion trajectory prediction. Appearance feature similarity calculation is based on the multi-dimensional feature vector of the target region, including shape descriptors, texture statistics, and heat distribution patterns. These features are normalized to reduce the impact of changes in viewpoint and lighting. Motion trajectory prediction uses an improved Kalman filter algorithm, combining the target's historical motion state (position, velocity, acceleration) with physical constraints in the pool environment (such as water flow direction and pool wall boundaries) to predict the target's possible position in the next frame. In each frame processing, the search range is first narrowed based on motion prediction, then the appearance similarity between the newly detected target and existing targets is calculated within this range, and finally, the optimal association is determined by weighted synthesis of the two matching results. When a target fails to match successfully within several consecutive frames, the tracking will not terminate immediately. Instead, it will maintain a brief hanging state based on the last known state and motion trend, allowing the target to resume tracking after a short period of occlusion or detection failure.

[0056] In real-world pool cleaning scenarios, when a fallen leaf drifts with the water flow over an uneven area of ​​the pool bottom, its appearance may change due to shadows and partial occlusion. However, by combining its stable motion trend (slowly moving with the water flow) with some visible appearance features, the target can be continuously tracked without being lost. Similarly, when a strand of hair is entangled near a pool wall outlet, its shape constantly changes due to water flow disturbances, and some areas may be temporarily obscured by air bubbles in the water. However, by analyzing the thermal distribution characteristics and motion constraints of its core area, the target can still be stably identified and tracked. Furthermore, when multiple small pieces of paper gather and float near the water surface, due to their proximity and perspective distortion, single-frame detection can easily confuse different individuals. However, by utilizing the uniqueness of cross-frame motion trajectories, each piece of paper can be accurately distinguished and tracked separately, providing accurate guidance for subsequent precise grasping.

[0057] During the target tracking result generation phase, not only are the target's category labels recorded, but rich motion state information is also comprehensively output. For each successfully tracked object, a comprehensive result is generated, including precise position coordinates (3D spatial position), motion direction vector, motion velocity magnitude, motion trajectory history, and predicted future position. This information is time-synchronized and spatially calibrated to ensure seamless integration with the underwater robot's control. Specifically, differentiated tracking strategy parameters are generated for different types of objects based on their type characteristics. For lightweight, floating fallen leaves, the focus is on their horizontal motion direction and velocity, predicting their drift path with the water flow. For easily entangled hair, the focus is on its morphological changes and attachment point location, predicting its possible extension direction. For small pieces of paper, the focus is on tracking their specific location and motion acceleration to cope with possible sudden changes in motion. This differentiated processing improves the accuracy and efficiency of subsequent grasping control.

[0058] Optionally, in step S104 above, target detection is performed based on feature maps of different resolutions to identify target debris of a preset type. This includes: performing bidirectional feature fusion of feature maps of different resolutions from top to bottom and bottom to top to enhance the semantic information in the high-resolution feature map and the spatial detail information in the low-resolution feature map, resulting in a fused multi-scale feature map; setting a size-adaptive detection window on the fused multi-scale feature map, and dynamically adjusting the density, size, and scale parameters of the detection window on each scale feature map based on historical detection data of fallen leaves, hair, and paper scraps in the pool environment, wherein small-sized detection is added to the high-resolution feature map. The distribution density of windows is used to obtain a set of local feature maps to improve the detection sensitivity of small debris. Each local feature map is classified and predicted using a lightweight convolutional network, which outputs the target category confidence and bounding box coordinates of each local feature map in the set. The lightweight convolutional network is optimized and trained for the morphological features of fallen leaves, hair, and paper scraps in a low-light underwater environment. The confidence threshold is dynamically adjusted by combining the ambient light intensity and water turbidity values ​​in the pool environment. Based on the target category confidence and bounding box coordinates of each local feature map, the set of local feature maps is adaptively filtered by the confidence threshold to identify the target debris set.

[0059] This embodiment employs a multi-scale feature fusion and adaptive detection strategy to optimize the identification of three types of debris in an underwater swimming pool environment: fallen leaves, hair, and paper scraps. Through refined feature processing and an environmental adaptation mechanism, this embodiment improves target detection performance under complex underwater conditions.

[0060] In the bidirectional feature fusion stage, feature maps of different resolutions are systematically integrated to establish a rich and hierarchical multi-scale feature representation. This process involves two complementary paths: a top-down path progressively transfers high-level semantic information to the high-resolution feature map, enhancing its ability to discriminate target categories; a bottom-up path aggregates low-level detailed features to the low-resolution feature map step by step, improving its spatial localization accuracy. This bidirectional fusion is particularly important in the underwater environment of a swimming pool because different types of debris require different levels of feature support: fallen leaves need rich semantic information to distinguish them from aquatic plants or shadows, hair needs precise spatial details to identify its slender shape, and paper scraps need to consider both morphological regularity and surface texture characteristics. Through this bidirectional information flow, the high-resolution feature map is endowed with stronger semantic understanding capabilities while retaining detailed information; at the same time, the low-resolution feature map gains more accurate boundary localization capabilities while maintaining a wide range of context awareness, laying a solid foundation for subsequent target detection.

[0061] In the adaptive detection window configuration phase, a set of detection windows with dynamically adjustable size and scale parameters were deployed on the fused multi-scale feature map. The design of these detection windows fully considered the typical size distribution and morphological characteristics of the three types of target debris in the pool environment. By analyzing historical detection data, optimal detection parameter models for different debris types on feature maps at different scales were established. For example, for tiny scraps of paper and fine hair fragments, small-sized detection windows were densely deployed on the high-resolution feature map to ensure these tiny targets were not overlooked due to excessively large windows; for medium-sized clumps of fallen leaves, medium-sized detection windows with specific scales were configured on the medium-resolution feature map to match their irregular shapes; for large clumps of fallen leaves or dense clusters of hair, large-sized detection windows were set on the low-resolution feature map to cover a wider area. This adaptive configuration not only considers the physical size of the target but also its performance characteristics in the underwater environment: under low light or high turbidity conditions, the effective visual size of the target changes, and the distribution density and size scale of the detection windows can be dynamically adjusted according to real-time environmental parameters. For example, when an increase in water turbidity is detected, the detection window size will be automatically enlarged to compensate for the positioning uncertainty caused by the blurring of the target edge; when the lighting conditions are improved, the density of small windows will be increased to improve the detection sensitivity of tiny impurities.

[0062] Optionally, a size-adaptive detection window can be set on the fused multi-scale feature map. When dynamically adjusting the density, size, and ratio parameters of the detection window on each scale feature map based on historical detection data of fallen leaves, hair, and paper scraps in the pool environment, the window parameters of each scale feature map can be dynamically optimized based on historical detection data of fallen leaves, hair, and paper scraps in the pool environment (e.g., size distribution, frequency of occurrence, and detection accuracy statistically analyzed from 1000 actual detections). For example, for high-resolution feature maps (corresponding to small target areas), based on the typical size of hair in historical data (e.g., average width 2-5 pixels, length 5-10 pixels) and dense distribution characteristics, the detection window size can be set to a small size (e.g., 4×8 pixels), and the window distribution density can be significantly increased (from a basic density of 10 / 100 pixels² to 25 / 100 pixels²), while setting the window ratio to 1:4 to match the slender shape of the hair. For medium-resolution feature maps (corresponding to medium-sized target areas), based on the historical size of the paper scraps (average width 8-15 pixels) and their common shapes, the window size is set to a medium value (e.g., 12×18 pixels), and the density is adjusted to 15 scraps / 100 pixels². For low-resolution feature maps (corresponding to larger target areas), based on the historical size of the fallen leaves (average width 10-20 pixels) and their sparser distribution, the window size is set to a larger value (e.g., 16×24 pixels), and the density is reduced to 8 scraps / 100 pixels². This dynamic adjustment mechanism ensures that the high-resolution feature map densely covers the detection area of ​​small debris (such as hair), improving the detection sensitivity for small targets while avoiding over-detection in large target areas, thus accurately identifying target debris such as fallen leaves, hair, and paper scraps in a swimming pool environment.

[0063] In another example, firstly, statistical features of the target type are extracted from historical data, such as the mean size distribution of hair (width 2.5 pixels ± 0.3 pixels, length 8.0 pixels ± 1.0 pixels) and frequency of occurrence (65% of all detected objects), the mean size distribution of fallen leaves (width 15.0 pixels ± 2.0 pixels, length 22.0 pixels ± 3.0 pixels) and frequency of occurrence (25% of all detected objects), and the mean size distribution of paper scraps (width 9.5 pixels ± 1.5 pixels, length 14.0 pixels ± 2.0 pixels) and frequency of occurrence (10% of all detected objects). Then, for the high-resolution feature map (corresponding to the small target region), the window size is set to 0.8 to 1.2 times the historical mean size (e.g., the hair window size is set to 2.0 × 6.4 pixels), and the window density is dynamically increased according to the frequency of occurrence (when the hair frequency is high, the density is increased from a base value of 10 per 100 pixels² to 25 per 100 pixels²). For medium-resolution feature maps, the window size is set to 1.0 times the historical average size (e.g., a paper scrap window size of 9.5 × 14.0 pixels), and the density is adjusted to 15 per 100 pixels². For low-resolution feature maps, the window size is set to 1.2 times the historical average size (e.g., a fallen leaf window size of 18.0 × 26.4 pixels), and the density is reduced to 8 per 100 pixels². This dynamic adjustment mechanism ensures that the window parameters closely match the actual size and distribution of targets in historical detections, significantly improving the detection sensitivity for small debris (such as hair) while avoiding redundant detection of large targets (such as fallen leaves), thus achieving accurate identification in swimming pool environments.

[0064] In the target classification and localization stage, a specially designed lightweight convolutional network is used to analyze local features within each detection window. This network is specifically optimized for the underwater pool environment, featuring high computational efficiency, low memory consumption, and fast inference speed, making it ideal for deployment on embedded platforms of underwater robots. The network training process fully utilizes samples of fallen leaves, hair, and paper scraps collected under various lighting conditions and water quality conditions, particularly enhancing the learning of morphological features in low-light underwater environments. The training dataset includes the performance of these debris under different poses, deformation states, and degrees of occlusion, enabling the network to identify targets even under partial occlusion or morphological changes. Furthermore, the network learns to distinguish common underwater interference factors (such as bubbles, water ripples, and pool bottom textures) from real targets, significantly reducing the false detection rate. During inference, the network outputs target category confidence and precise bounding box coordinates for each detection window; these results are post-processed to form a preliminary detection set.

[0065] During the adaptive threshold adjustment phase, the confidence threshold for target identification is dynamically adjusted based on the real-time status of the pool environment. This mechanism automatically balances detection sensitivity and accuracy according to environmental conditions. The system acquires ambient light intensity and water turbidity parameters in real time through built-in environmental sensors or image analysis modules. Under low light or high turbidity conditions, target features are often not obvious enough. In this case, the system automatically lowers the confidence threshold to increase detection sensitivity and avoid missing important targets. Under sufficient light or clear water conditions, the system raises the confidence threshold to strictly screen detection results and reduce false detection interference. This adaptive adjustment is not a simple global threshold change, but a fine-grained control for different types of debris and feature maps of different scales. For small hairs that are difficult to detect, the threshold adjustment range is larger to ensure detection even under harsh conditions. For large fallen leaves with obvious features, the threshold adjustment range is smaller to maintain detection stability. In addition, the continuity in the time dimension is also considered. For targets with large confidence fluctuations in a short period of time, smoothing filtering is used to avoid frequent jumps in detection results.

[0066] Optionally, when dynamically adjusting the confidence threshold based on the ambient light intensity and water turbidity values ​​in the swimming pool environment, the ambient light intensity (range 0-1000 Lux) is first acquired in real time using an underwater light sensor, and the water turbidity (range 0-100 NTU) is acquired using a turbidity sensor. A dynamic adjustment rule is established based on historical data: a base confidence threshold is set at 0.5; when the ambient light intensity is below 50 Lux (low light environment), the confidence threshold decreases dynamically in a linear relationship, specifically, for every 10 units below 50 Lux, the threshold decreases by 0.01 (e.g., when the light intensity is 30 Lux, the threshold decreases by 0.02, adjusting to 0.48). When the water turbidity is above 50 NTU (high turbidity environment), the confidence threshold further decreases linearly, specifically, for every 10 units above 50 NTU, the threshold decreases by 0.02 (e.g., when the turbidity is 65 NTU, the threshold decreases by 0.03, adjusting to 0.47). If both low light and high turbidity conditions are met simultaneously, the threshold is adjusted accordingly (e.g., with 30 Lux of light and 65 NTU of turbidity, the threshold = 0.5 - 0.02 - 0.03 = 0.45). This mechanism ensures that in low light or high turbidity environments (such as underwater dark areas or muddy waters), the confidence threshold is lowered to improve the detection sensitivity for target debris such as fallen leaves and hair, avoiding missed detections. When there is sufficient light (>50 Lux) and the water is clear (<50 NTU), the threshold is maintained at a higher level to reduce false detections, thereby achieving robust target recognition in swimming pool environments.

[0067] In real-world pool cleaning scenarios, such as when an underwater robot is working in an outdoor pool at dusk, lighting conditions deteriorate drastically, often rendering traditional detection methods ineffective. However, the adaptive detection mechanism in this embodiment, by lowering the confidence threshold and enhancing the detection density of the high-resolution feature map, successfully identifies tiny pieces of paper that are almost invisible under low light conditions. Similarly, when dealing with scattered hair in turbid water, the hair's varied morphology and low contrast with the background are addressed by enhancing its performance in multi-scale feature maps through bidirectional feature fusion and densely deploying elongated detection windows on the high-resolution feature map, successfully capturing these difficult-to-identify targets. Furthermore, when a semi-rotten leaf floats at the boundary between light and shadow on the pool bottom, its area is partially bright and partially dark. Traditional methods easily identify it as multiple fragments or miss it entirely. This embodiment, however, by fusing feature information from different scales and combining it with adaptive threshold adjustment, accurately identifies the complete target and precisely locates its boundaries.

[0068] Further optionally, in step S104 above, a unique identifier is assigned to each detected target object, and cross-frame continuous tracking of the target object is performed by combining the similarity of the target appearance features with the motion trajectory prediction. This includes: generating a unique identifier for each first-detected target object and recording the initial position, bounding box size, category information, and first detection timestamp of the target object; extracting a multi-dimensional appearance feature vector from the bounding box region of each target object, wherein the multi-dimensional appearance feature vector includes local binary pattern features, color histogram distribution, and depth appearance feature embedding; constructing a motion state vector based on the historical position data of the target object, wherein the motion state vector includes position coordinates, velocity components, acceleration components, and motion direction angle, and adopting an adaptive... The Kalman filter predicts the expected position and uncertainty range of the target in the next frame; it calculates the comprehensive matching score between each newly detected target clutter in the current frame and the historical target clutter with assigned identities. The comprehensive matching score is a weighted combination of appearance feature similarity score and motion prediction consistency score. The weights of appearance feature similarity score and motion prediction consistency score are adaptively adjusted according to ambient lighting conditions. When the target clutter is in a low-light area or partially occluded, the weight coefficient of motion prediction consistency score is increased to reduce dependence on unstable appearance features and maintain the continuity of cross-frame tracking. For the current detected target clutter and historical target clutter pairs whose comprehensive matching scores exceed a preset threshold, identity association is performed, and the state information of the historical target clutter is updated. Further optionally, for newly detected target clutter that does not match, it is marked as a first detection state; for historical target clutter that has not matched new detection results for multiple consecutive frames, the historical target clutter is marked as a lost state and the identity identifier resource is released.

[0069] In the further optimization of step S104, a unique identifier is assigned to each detected target debris, and continuous tracking is performed across frames by combining the similarity of the target's appearance features with motion trajectory prediction. This process not only improves the recognition accuracy of target debris such as fallen leaves, hair, and paper scraps in the pool environment, but also ensures the continuous tracking of these targets in the video sequence, even after occlusion or movement to different lighting conditions. First, when a target debris is detected for the first time, a unique identifier is generated for it, and the target's initial position, bounding box size, category information, and the timestamp of the first detection are recorded. This step provides basic data support for subsequent tracking, ensuring that each target can be independently identified and tracked.

[0070] Next, multi-dimensional appearance feature vectors are extracted from the bounding box region of each target debris, including local binary pattern features, color histogram distribution, and depth appearance feature embedding. These features effectively capture the texture, color, and morphological characteristics of the target, providing reliable discrimination capabilities even in complex environments. For example, in a swimming pool environment, the color of fallen leaves may change depending on their degree of decay, and the shape of hair may be distorted due to water flow. Through these multi-dimensional appearance feature vectors, the system can accurately identify and distinguish these target debris.

[0071] To predict the position of a target object in a future frame, a motion state vector is constructed based on its historical position data. This vector contains information such as position coordinates, velocity components, acceleration components, and motion direction angle. An adaptive Kalman filter is then used to predict the target's expected position and its uncertainty range in the next frame. The Kalman filter can estimate the optimal state prediction based on current observation information and previous states, making it particularly suitable for handling noisy dynamic systems. In the application scenario of pool cleaning robots, this prediction mechanism can help the robot plan its path in advance, approaching and cleaning target debris more efficiently.

[0072] To achieve cross-frame tracking, a comprehensive matching score is calculated between newly detected target clutter in the current frame and historical target clutter with assigned identities. This score consists of an appearance feature similarity score and a motion prediction consistency score, with the weights of these two scores adaptively adjusted based on ambient lighting conditions. When the target is in a low-light area or partially occluded, the weight coefficient of the motion prediction consistency score is increased to reduce the dependence on appearance feature instability, thereby maintaining the continuity of cross-frame tracking. For example, when a target clutter enters a shadow area or is temporarily occluded by other objects, although its appearance features may change, its position can still be tracked through motion trajectory prediction.

[0073] For currently detected target clutter pairs with a combined matching score exceeding a preset threshold, an identity association operation is performed, and the status information of the historical target clutter is updated. This ensures coherent target tracking throughout the entire video sequence, even in the event of temporary occlusion or changes in lighting. Conversely, newly detected target clutter that fails to match is marked as newly detected. Historical target clutter that fails to match new detection results for multiple consecutive frames is marked as lost, and the corresponding identity identifier resources are released, allowing the system to manage target tracking tasks more efficiently.

[0074] Step S105: Based on the target tracking results, control the underwater robot's motion control system and cleaning device to track and grab the target debris in real time.

[0075] As an optional embodiment, in step S105, based on the position coordinates, direction of motion, and velocity information of the target debris in the target tracking results, the real-time relative positional relationship between the underwater robot and the target debris is calculated, generating a smooth approach path that conforms to the robot's kinematic constraints. This smooth approach path avoids sharp turns to reduce water flow disturbance. Furthermore, when the relative distance between the underwater robot and the target debris is less than a preset distance threshold, the operating parameters of the cleaning device are adaptively adjusted according to the type of the target debris. Finally, during the grasping process, the target position change is monitored in real time, and the output parameters of the underwater robot's horizontal or vertical thrusters are dynamically adjusted, as are the suction strength and direction parameters of the cleaning device. When the target debris's position deviation is detected to exceed the deviation threshold, a repositioning mechanism is triggered to ensure that the target debris is completely captured and collected into the storage device.

[0076] Thus, step S105 achieves a complete control closed loop from target recognition to physical grasping, transforming visual perception results into precise mechanical actions. In this embodiment, a multi-layered adaptive control strategy dynamically adjusts the underwater robot's trajectory and cleaning parameters based on the type of target debris, its motion state, and environmental conditions, ensuring efficient and precise cleaning operations.

[0077] During the motion planning phase, the spatial relationship between the underwater robot and the target debris is calculated in real time based on the position coordinates, direction of motion, and velocity information from the target tracking results. This calculation considers not only the relative position at the current moment but also the target's motion trend to predict its future trajectory. Based on this, the system generates a smooth approach path that conforms to the robot's kinematic constraints. This path is carefully designed to avoid abrupt turns and acceleration changes, thereby minimizing water flow disturbances caused by the robot's motion. This smooth path planning is particularly important because severe water flow disturbances can cause lightweight debris (such as fallen leaves and paper scraps) to deviate from their intended position, increasing the difficulty of grasping. The path generation algorithm fully considers the dynamic characteristics of the underwater robot, including thruster response delay, fluid resistance, and inertial effects, ensuring that the generated path is physically feasible and energy efficient. Furthermore, the path planning integrates environmental constraint information, such as the positions of pool walls, steps, and other fixed obstacles, automatically generating obstacle avoidance trajectories to ensure the robot can navigate safely in complex pool environments.

[0078] During the adaptive cleaning parameter adjustment phase, the operating parameters of the cleaning device are dynamically configured based on the type and characteristics of the target debris. When the relative distance between the underwater robot and the target debris is less than a preset threshold, the corresponding cleaning mode is automatically activated: For large, lightweight targets such as fallen leaves, a wide-area suction mode is activated to expand the suction coverage and optimize the suction distribution, ensuring that the entire leaf is captured intact without being torn; for long, tangled targets such as hair, a rotating brush head combined with a directional suction mode is activated to loosen the attached hair through mechanical brushing, while the directional suction precisely guides the hair into the collection channel, preventing it from tangling on the robot parts; for small targets such as paper scraps, a precise positioning suction mode is activated to concentrate high-precision suction on the target area, avoiding disturbance of the surrounding water that could cause other debris to drift away. The parameter settings for each mode (such as suction intensity, brush head speed, and effective range) are optimized for the characteristics of the corresponding debris to ensure the highest capture efficiency with minimal energy consumption. This adaptive parameter adjustment not only improves the cleaning effect but also extends the service life of the equipment and reduces mechanical failures caused by improper operation.

[0079] During the real-time control and adjustment phase, the target's position changes are continuously monitored throughout the grasping process. The underwater robot's thruster output and cleaning device parameters are dynamically adjusted to form a closed-loop control system. The horizontal thrusters control the robot's position and orientation in the horizontal plane, while the vertical thrusters adjust the robot's depth and pitch angle. These two systems work together to ensure the robot maintains the optimal relative position to the target. Simultaneously, the suction strength and direction parameters of the cleaning device are adjusted based on real-time feedback. When the target approaches the suction inlet, the suction is appropriately reduced to avoid impact damage. When a target is detected entering the collection channel, the suction is increased to ensure complete capture. This dynamic adjustment mechanism is particularly suitable for handling targets drifting in water currents, such as fallen leaves carried by the current or paper scraps affected by eddies. It can sense these dynamic changes and compensate in real time, maintaining a stable capture process.

[0080] When a significant shift in the target debris's position is detected (e.g., due to sudden changes in water flow or external interference), the system automatically triggers a repositioning mechanism. This mechanism first suspends the current grasping action to avoid wasted energy or equipment damage from ineffective operations. Subsequently, based on the latest target position and motion state, it quickly replans the approach path and calculates the optimal repositioning trajectory. This process fully considers the underwater robot's dynamic limitations and environmental constraints, ensuring that the repositioning action is both rapid and smooth. After repositioning, the system resumes the grasping operation and adjusts the cleaning parameters accordingly to adapt to the new relative position. This self-healing capability improves the system's robustness in complex dynamic environments, enabling the underwater robot to complete cleaning tasks under various unpredictable conditions.

[0081] In real-world pool cleaning scenarios, for example, when a fallen leaf drifts with the water flow, a smooth approach path is first planned, allowing the robot to approach from downstream to avoid disturbing the leaf. When the distance approaches a threshold, the system automatically activates a wide-area suction mode, while dynamically adjusting the robot's horizontal thruster output based on the leaf's drift speed to maintain relative stillness. Under suction, the leaf gradually moves towards the collection port, with the system monitoring its position in real time and fine-tuning the suction intensity to ensure the leaf enters the storage device intact. Similarly, when dealing with tangled hair near the pool bottom drain, the robot first precisely navigates to the area above the hair, then activates a rotating brush head combined with directional suction. The brush head loosens the hair's attachment point to the pool bottom, and the directional suction then guides the loosened hair to the collection channel. Even if some hair shifts due to water flow disturbance during cleaning, the repositioning mechanism responds quickly, adjusting the robot's position and continuing the cleaning process. For example, for tiny pieces of paper floating near the water surface, a precise positioning suction mode is used. By controlling the vertical thruster to maintain a stable depth, the suction direction and intensity are adjusted to ensure that the paper scraps are completely captured and not blown to a farther location.

[0082] In the above embodiments, smooth approach path planning effectively reduces water flow disturbance caused by robot movement, improving the success rate of capturing lightweight floating debris. Adaptive cleaning parameter adjustment based on debris type significantly improves the processing efficiency for different targets, solving the technical challenge of traditional single-mode cleaning devices being unable to handle multiple debris types simultaneously. The real-time closed-loop control mechanism enables the system to adapt to dynamic changes in the underwater environment, improving operational stability under complex conditions. The repositioning mechanism enhances the system's fault tolerance, allowing it to automatically recover and complete the cleaning task even in the event of sudden changes in target position, significantly improving system reliability and user experience. This multi-level adaptive control strategy achieves efficient energy utilization, reducing energy consumption while ensuring cleaning effectiveness and extending the working time per charge. Overall, this embodiment solves the key control challenges of underwater cleaning robots in practical applications, making the automatic cleaning process more intelligent, efficient, and reliable, providing a revolutionary technical solution for pool maintenance, reducing the burden of manual cleaning, and improving pool hygiene and user experience.

[0083] Further, in step S105 above, a three-dimensional environment map is constructed, including the current position of the underwater robot, the predicted position of the target debris, and the motion trajectory, and the positions of known static obstacles are marked; an improved A** algorithm combined with a dynamic window method is used to generate an initial path point sequence, which avoids static obstacles and satisfies the robot's minimum turning radius constraint; B-spline curve fitting is applied to the initial path point sequence to generate a continuous and differentiable smooth trajectory, the rate of curvature change of which is controlled within the range that the robot's attitude control system can smoothly track; based on real-time water flow velocity and direction data in the underwater environment, feedforward compensation is performed on the generated smooth trajectory to predict and offset the interference effect of water flow on the robot's trajectory; the compensated smooth trajectory is decomposed into a series of time-position control points, and the interval between the time-position control points is dynamically adjusted according to the relative distance between the robot and the target debris, reducing the time interval when approaching the target to improve tracking accuracy. Therefore, the motion planning step S105 adopts advanced three-dimensional environment modeling and adaptive trajectory generation technology, which improves the navigation accuracy and motion stability of the underwater robot in complex swimming pool environments. It solves the technical problems of inaccurate positioning, energy waste and target loss caused by uneven path, water flow interference and obstacle avoidance in the process of traditional underwater robots approaching the target.

[0084] Specifically, in the 3D environment map construction phase, a dynamically updated environment model is established. This model integrates the underwater robot's precise current position, the predicted position of target debris and its trajectory, and marks the positions of known static obstacles. Static obstacles include fixed structures such as pool walls, steps, drain outlets, and light fixture supports. This information is partly derived from a pre-input pool CAD model and partly from the robot's real-time perception via sonar and visual sensors. The predicted position of target debris is calculated using a trajectory prediction algorithm based on motion parameters obtained from cross-frame tracking in step S104. This comprehensive environment model enables the robot to accurately understand its relative relationship with targets and obstacles in 3D space, providing a complete spatial information foundation for subsequent path planning. Furthermore, a hierarchical representation is adopted, dividing the environment into a macro-navigation layer and a fine-grained operation layer. The macro-navigation layer focuses on large-scale obstacle avoidance and overall path direction, while the fine-grained operation layer focuses on minor adjustments and precise alignment near the target. This hierarchical representation improves computational efficiency, enabling real-time path planning with limited embedded computing resources.

[0085] In the initial path planning phase, an improved A* algorithm combined with a dynamic window method is used to generate an initial path point sequence. The improved A* algorithm, based on traditional heuristic search, incorporates considerations of the underwater robot's kinematic constraints, particularly the minimum turning radius. The algorithm's cost function includes not only path length but also factors such as changes in turning angle, safe distances from obstacles, and estimated water resistance, ensuring the generated path is physically feasible and energy efficient. The dynamic window method, based on the coarse path generated by the A* algorithm, further considers the robot's speed and acceleration limitations and real-time environmental changes, generating local motion commands that conform to dynamic constraints. The combination of these two methods leverages their respective advantages: the A* algorithm provides a global optimality guarantee, while the dynamic window method ensures the executability and safety of local actions. In a pool environment, this method effectively avoids fixed obstacles (such as pool bottom decorations and underwater light fixtures) while generating an initial path that satisfies the underwater robot's motion characteristics. Specifically, for narrow areas commonly found in pools (such as under stairs or near drains), the path point density is automatically increased to improve navigation accuracy and avoid collision risks.

[0086] In the trajectory smoothing stage, B-spline curve fitting technology is applied to the initial path point sequence to generate a continuously differentiable smooth trajectory. This process transforms discrete path points into mathematically smooth curves, ensuring that the curvature and rate of change of the trajectory meet the range that allows the underwater robot's attitude control to track smoothly. The selection of B-spline curves is based on their local controllability and shape preservation characteristics; even if some control points are modified, it will not cause drastic changes in the overall trajectory, which is particularly important in real-time adjustment scenarios. During the smoothing process, special attention is paid to the continuity of the second derivative (jerk) of the trajectory to avoid abrupt changes in thruster load and reduce energy consumption and mechanical wear. At the same time, the smoothing trajectory also considers underwater hydrodynamic characteristics, such as avoiding large-angle turns perpendicular to the water flow direction as much as possible to reduce fluid resistance. This smooth trajectory not only improves the comfort and stability of movement but also reduces water flow disturbance caused by sharp turns, preventing target debris (especially lightweight fallen leaves and paper scraps) from drifting due to water flow interference, thus improving the success rate of subsequent grasping.

[0087] In the water flow disturbance compensation stage, feedforward compensation is applied to the generated smooth trajectory based on real-time water flow velocity and direction data in the underwater environment. The water flow data comes from a miniature flow velocity sensor and visual flow analysis algorithm mounted on the underwater robot, enabling real-time perception of changes in the local water flow field. The feedforward compensation mechanism predicts the interference effect of the water flow on the robot's trajectory and adjusts control commands in advance to counteract these disturbances. For example, when a lateral water flow is detected, the trajectory is pre-shifted in the opposite direction of the flow, ensuring the robot moves precisely along the expected path under the influence of the water flow. Compared to traditional feedback control, this feedforward compensation offers advantages such as faster response and smaller overshoot, making it particularly suitable for periodic water flows (such as those generated by circulating filters) and sudden disturbances (such as waves caused by swimmers) commonly found in swimming pools. The compensation algorithm employs an adaptive mechanism, dynamically adjusting the compensation intensity based on the stability of the water flow and the prediction accuracy. It increases the compensation amplitude in strongly turbulent environments and decreases the compensation in stable environments to avoid over-adjustment. This adaptive feedforward compensation improves the trajectory tracking accuracy of the underwater robot in dynamic water environments, enabling the robot to accurately approach floating or moving target debris.

[0088] During the trajectory execution optimization phase, the compensated smooth trajectory is decomposed into a series of time-position control points, which define the spatial position the robot should reach at a specific moment. The time interval between control points is not fixed but dynamically adjusted according to the relative distance between the robot and the target object: a larger time interval is used when the distance to the target is far to improve motion efficiency; as the distance to the target approaches, the time interval is gradually reduced to improve tracking accuracy and response speed. This dynamic adjustment strategy ensures a higher control frequency during critical operation phases (such as precise positioning before grasping) while maintaining computational efficiency during large-scale movement phases. Furthermore, the distribution of time-position control points also considers the robot's dynamic characteristics, such as increasing the density of control points in high-acceleration regions and reducing the number of control points in uniform linear motion regions to optimize computational resource allocation. A predictive control mechanism is also implemented, dynamically adjusting the parameters of future control points based on historical tracking errors to further improve trajectory tracking accuracy.

[0089] For example, when an underwater robot needs to approach fallen leaves located below a tiered platform in a pool, a 3D environmental map is first constructed, including the tiered structure, pool walls, and predicted leaf locations. An improved A* algorithm, combined with a dynamic window method, plans an initial path that avoids the edges of the tiered platform. B-spline curve fitting generates a smooth trajectory, ensuring the robot doesn't lose stability due to sharp turns. Simultaneously, a weak water flow generated by filtration is detected, and the trajectory is adjusted through feedforward compensation, allowing the robot to navigate accurately even under the influence of the water flow. As it approaches the fallen leaves, the control point time interval automatically decreases, and the robot performs fine-tuning of its positioning, preparing for subsequent grasping operations. Similarly, when tracking paper scraps drifting with the water flow, the robot not only predicts the scrap's position based on its current motion but also considers the impact of the water flow on its own motion, generating a composite trajectory that compensates for both motions, enabling the robot to efficiently intercept moving targets. Furthermore, when cleaning hair near a drain at the bottom of a pool, due to the narrow space and fixed obstacles, a fine trajectory with a high control point density is automatically generated, ensuring the robot can accurately position itself within the limited space and avoid collisions with the drain structure.

[0090] Overall, this multi-level, adaptive trajectory planning and control strategy enables underwater robots to achieve precise, stable, and efficient navigation and target approach in complex pool environments, laying a solid foundation for subsequent cleaning operations and improving overall cleaning efficiency and user experience.

[0091] Furthermore, in step S105 above, the operating parameters of the cleaning device are adaptively adjusted according to the type of target debris, including: for large, lightweight target debris such as fallen leaves, a wide-area suction mode is activated, adjusting the suction range to a circular area with a radius of 0.7-1.0 meters, and the suction intensity is set to 80-90 kPa. For example, when a large piece of fallen leaf is detected in a corner of a pool, the wide-area suction mode is automatically activated to expand the suction coverage area, ensuring that the entire leaf is completely sucked in without being torn. For long, tangled target debris such as hair, a rotating brush head is activated in conjunction with a directional suction mode, setting the rotating brush head speed to 35-45 rpm, the suction intensity to 55-65 kPa, and dynamically adjusting the suction direction to be consistent with the direction of hair extension. When hair tangled near the drain is detected, the rotating brush head is activated to first loosen the attachment point, and then the directional suction precisely guides the hair into the collection channel. For small pieces of debris like paper scraps, activate the fine-grained positioning suction mode, adjusting the suction range to a circular area with a radius of 0.2-0.3 meters, and setting the suction intensity to 35-45 kPa. When small pieces of paper are found floating on the water surface, use fine-grained positioning suction to avoid excessive suction causing the paper scraps to break or excessively disturbing the surrounding water. This adaptive parameter adjustment mechanism provides customized cleaning strategies for the physical characteristics of different debris, improving capture success rate and energy efficiency. By using dedicated modes that match the characteristics of the debris, it avoids the problems of incomplete capture, energy waste, or debris damage caused by traditional single-parameter settings. For lightweight fallen leaves, wide-area suction ensures complete capture; for easily tangled hair, a combination of mechanical loosening and directional suction prevents equipment blockage; for small pieces of paper scraps, precise positioning reduces water disturbance, comprehensively improving cleaning quality and equipment reliability.

[0092] Further, in step S105 above, during the grasping process, the target position change is monitored in real time, and the output parameters of the underwater robot's horizontal and / or vertical thrusters are dynamically adjusted. The suction strength and directional parameters of the cleaning device are also dynamically adjusted. This includes: establishing a time series for tracking the position of the target debris; obtaining the actual trajectory of the target debris relative to the underwater robot by calculating the difference between the position change vector of the target debris and the pose change of the underwater robot in consecutive frames; and performing time-series analysis of the target trajectory using a sliding window mechanism. When the rate of change of the target's motion direction exceeds a preset angle change rate threshold (e.g., direction change rate ≥ 15° / s) or the acceleration exceeds a preset acceleration threshold (e.g., acceleration ≥ 0.5), the analysis is performed accordingly. When the speed reaches (m / s²), a thruster compensation control signal is generated. The amplitude and phase of the compensation control signal are adaptively adjusted according to the motion characteristic parameters (including but not limited to at least one of the following: moment of inertia, direction of motion, velocity, and damping coefficient) corresponding to the type of target debris. A spatial mapping table of relative positions between the underwater robot and the target debris is constructed. The thruster thrust parameters are determined based on the polar coordinates of the relative position vector, where the radial thrust component is used to control the propulsion speed and the angular motion control component is used to control the turning rate. Based on the category identification result of the target debris and its current relative position, a suction field distribution model of the cleaning device is generated. The suction field distribution model dynamically calculates the optimal suction intensity distribution and direction of action using fluid dynamics simulation formulas based on the physical characteristic parameters of the target debris. The physical characteristic parameters include at least one of density, shape parameters, and material. Based on the optimal suction intensity distribution and direction of action, corresponding suction intensity and direction parameters are set. Among them, the shape parameters include but are not limited to particle size, area, and volume. Thus, the adaptability to dynamic aquatic environments is improved, and high-precision closed-loop control is achieved. By continuously monitoring the relative motion of the target and compensating in a timely manner, the system effectively overcomes the influence of complex factors such as water flow disturbance and target deformation, thus improving the success rate of capture. Simultaneously, suction field optimization based on category characteristics ensures the stability of different debris during the capture process, reducing capture failures caused by parameter mismatches. This mechanism also optimizes energy consumption through precise power allocation, extending the equipment's operating time while maintaining control accuracy.

[0093] For example, when a fallen leaf suddenly changes its drift direction with the water flow, this change is detected in real time, and the thrust of the right horizontal thruster is rapidly increased, while the suction direction is adjusted to align with the new direction of the leaf's movement. When hair begins to bend and deform under the suction and partially escapes the suction range, the suction intensity distribution is dynamically adjusted to enhance the suction in the edge areas and maintain overall control over the hair. When paper scraps rise to the surface due to air bubbles during their approach, the thrust of the vertical thruster is increased upwards, while the suction direction is adjusted to maintain stable tracking and capture of the paper scraps.

[0094] In the above embodiments, optionally, when constructing the spatial mapping table of relative positions between the underwater robot and the target debris, the relative position vector of the target debris relative to the robot (Cartesian coordinates acquired in real time by the vision system) is first converted into polar coordinates to obtain a radial distance component (representing the straight-line distance between the target debris and the robot) and an angular component (representing the azimuth angle of the target debris relative to the robot). Based on this, the radial thrust component is generated by calculating the absolute value of the radial distance component and multiplying it by a preset scaling factor (for example, this factor is dynamically adjusted according to the underwater environment flow velocity, taking a larger value such as 1.0 when the flow velocity is high and a smaller value such as 0.5 when the flow velocity is low). This component is directly used to control the propulsion speed of the underwater robot: when the radial distance is large (the target is far away), the radial thrust component increases to increase the propulsion speed and quickly approach the target; when the radial distance is small (the target is close), the radial thrust component decreases to reduce the propulsion speed and avoid collision. Simultaneously, the angular motion control component is generated by calculating the absolute value of the angular component and multiplying it by another preset proportional coefficient (for example, this coefficient is set based on the robot's steering inertia; a smaller value, such as 0.3, is used when the inertia is large, and a larger value, such as 0.7, is used when the inertia is small). This component is directly used to control the steering rate: when the angular deviation is large (the target orientation deviates significantly), the angular motion control component increases to improve the steering rate and quickly align with the target; when the angular deviation is small (the target orientation is close to directly in front), the angular motion control component decreases to reduce the steering rate and achieve smooth adjustment. The calculation of the above components is based on real-time relative position data, ensuring that the thruster output parameters (such as the thrust magnitude of the horizontal thruster and the thrust magnitude of the vertical thruster) dynamically match the target motion state, achieving precise tracking.

[0095] In the above embodiments, optionally, when the suction field distribution model dynamically calculates the optimal suction intensity distribution and direction based on the physical characteristics of the target debris using fluid dynamics simulation formulas, control can be achieved through computational fluid dynamics (CFD). The fluid dynamics simulation formulas can be based on the fundamental equations of fluid dynamics, particularly the momentum conservation equation (Navier-Stokes equations) and the continuity equation. In an underwater environment, the suction effect of the fluid on the target debris can be approximately expressed as: ;in, Suction intensity (unit: N). Let ρ be the density of water (approximately 1000 kg / m³), and v be the velocity of the fluid relative to the target debris (m / s). Let A be the suction coefficient (related to the shape and material of the target debris), and A be the effective area (m²), i.e., the particle size. Based on the above principle and combined with the physical characteristics of the target debris, the dynamic calculation formula for suction intensity is as follows: ;in, This is a proportionality coefficient (adjusted according to the underwater environment, typically ranging from 0.5 to 1.5). Furthermore, the optimal suction intensity distribution can be obtained using the gradient descent method.

[0096] Furthermore, the calculation process for the direction of suction (i.e., the direction mentioned earlier) is as follows: the streamline direction from fluid dynamics is multiplied by a first shape factor to obtain the first product. The first shape factor is determined based on the shape of the debris; 0.5 is used for spherical debris, and 0.8 for elongated debris. The pressure gradient is then calculated using CFD simulation, and an inverse trigonometric function is applied to the pressure gradient to obtain an intermediate value. This intermediate value is multiplied by a second direction adjustment factor (ranging from 0.3 to 0.8) to obtain the second product. The sum of the first and second products is taken as the direction of suction (i.e., the direction mentioned earlier).

[0097] Further, in step S105 above, a position offset monitoring window is set to calculate the Euclidean distance between the current position and the expected position of the target debris in real time. When the Euclidean distance exceeds a dynamic threshold set based on the type of target debris, it is determined to be a large position offset. When a large position offset is detected, the current grasping action is immediately paused, and the next position of the target debris is predicted based on its movement speed and direction. The shortest relocation path from the current robot position to the predicted position is calculated, taking into account robot posture change constraints and water flow resistance effects. The output power of the horizontal / vertical thrusters is adjusted to make the robot move quickly along the shortest relocation path to the vicinity of the predicted position, and the suction strength and direction of the cleaning device are adjusted simultaneously to ensure that the grasping action is resumed immediately after relocation is completed. During the relocation process, continuous tracking of the target debris is maintained, and the predicted position is dynamically updated based on the latest tracking results to achieve closed-loop control until the target debris is captured and stored in the storage device. Thus, this relocation mechanism gives the system excellent fault tolerance and robustness, effectively solving the technical problem of traditional underwater robots being unable to recover after target loss. A closed-loop control strategy involving intelligent pause, relocation, and recovery improves task completion rates in complex dynamic environments. The calculation of the shortest relocation path considers robot posture constraints and water flow effects, optimizing motion efficiency and reducing relocation time. Simultaneously, dynamic updates to the predicted position ensure the system can adapt to the uncertainty of target motion, significantly improving operational reliability in turbulent environments. This self-recovery capability not only increases the success rate of single tasks but also reduces the need for human intervention, achieving true autonomous cleaning.

[0098] The embodiments of the present invention achieve accurate identification and stable tracking of irregular targets such as fallen leaves, hair and paper scraps in complex underwater environments, effectively improving the automation level and work efficiency of underwater cleaning operations.

[0099] This invention provides a target tracking system for an underwater robot in low-light environments. The system includes: a data acquisition module for synchronously acquiring image data of the underwater environment using a visible light camera and an infrared sensor mounted on the underwater robot, the image data including low-light visible light images and infrared images; an enhancement module for adaptive temporal fusion, contrast enhancement, and adaptive curve mapping processing of continuously acquired multi-frame low-light visible light images to obtain a low-light enhanced image; a fusion module for extracting surface texture features from the low-light enhanced image and thermal distribution features from the continuously acquired infrared images, and fusing the surface texture features and thermal distribution features using adaptive weights to generate a fused feature image; the adaptive weights are dynamically adjusted based on the edge morphology and thermal characteristics differences of target debris in the pool environment; and a tracking module for identifying and continuously tracking target debris of a preset type in the fused feature image across frames to obtain target tracking results; the types of target debris include fallen leaves, hair, and paper scraps; and controlling the underwater robot's motion control system and cleaning device based on the target tracking results to track and grasp the target debris in real time. In some implementations, the underwater robot's low-light environment target tracking system can be applied to terminal devices. It should be noted that, for the sake of convenience and brevity, the specific working process of the underwater robot's low-light environment target tracking system described above can be referred to the corresponding process in the aforementioned embodiments of the underwater robot's low-light environment target tracking method, and will not be repeated here.

[0100] This invention provides a terminal device. The terminal device includes a processor and a memory, which are connected via a bus, such as an I / O bus. 2C-bus. Specifically, the processor provides computing and control capabilities to support the operation of the entire terminal device. The processor can be a central processing unit, or it can be other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Those skilled in the art will understand that the structures shown in the above embodiments are merely block diagrams of some structures related to the embodiments of the present invention, and do not constitute a limitation on the terminal device to which the embodiments of the present invention are applied. A specific server may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. The processor is used to run a computer program stored in the memory, and when executing the computer program, implements any of the underwater robot target tracking methods in low-light environments provided by the embodiments of the present invention. It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the aforementioned embodiments of the underwater robot target tracking method in low-light environments, and will not be repeated here. This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the steps of any of the underwater robot target tracking methods in low-light environments provided in the specification of this invention.

Claims

1. A target tracking method for underwater robots in low-light environments, characterized in that, The method includes: The underwater robot simultaneously acquires image data of the underwater environment using a visible light camera and an infrared sensor, including weak-light visible light images and infrared images. Adaptive temporal fusion, contrast enhancement, and adaptive curve mapping are performed on continuously acquired multi-frame low-light visible light images to obtain low-light enhanced images. The process involves extracting surface texture features from the low-light enhanced image and thermal distribution features from continuously acquired infrared images. The surface texture features and thermal distribution features are then fused using adaptive weights to generate a fused feature image. This includes: extracting edge, corner, and surface texture features at different scales from the low-light enhanced image using a multi-scale convolutional filter bank to construct a multi-scale texture feature map; enhancing the thermal contrast of the continuously acquired infrared images and extracting heat source distribution features through gradient domain analysis to generate a thermal distribution feature map; calculating the feature saliency index of the target region in the multi-scale texture features and the thermal distribution feature map, where the feature saliency index includes edge sharpness, texture complexity, thermal contrast, and region consistency parameters; dynamically assigning fusion weights to the texture features and thermal distribution features at different scales based on the feature saliency index; weighting the multi-scale texture features and the thermal distribution feature map according to the dynamically assigned fusion weights, and generating the fused feature image through a feature reconstruction network; the adaptive weights are dynamically adjusted based on the edge morphology and thermal characteristics differences of the target debris in the pool environment. The method involves identifying and continuously tracking target debris of a preset type in the fused feature image to obtain target tracking results. This includes: extracting multi-scale features from the fused feature image to generate feature maps of different resolutions; detecting targets based on these feature maps to identify target debris of a preset type; assigning a unique identifier to each detected target debris and continuously tracking the target debris across frames by combining target appearance feature similarity with motion trajectory prediction; and generating target tracking results containing the target debris's position, direction of motion, and speed when the detected target debris is a fallen leaf, hair, or paper scrap. The types of target debris include fallen leaves, hair, and paper scraps. Based on the target tracking results, the motion control system and cleaning device of the underwater robot are controlled to track and grab the target debris in real time.

2. The underwater robot target tracking method in low-light environments according to claim 1, characterized in that, The process of adaptive temporal fusion, contrast enhancement, and adaptive curve mapping of continuously acquired multi-frame low-light visible light images to obtain low-light enhanced images includes: For multiple frames of continuously acquired low-light visible light images, the pixel difference and structural similarity between adjacent frames are calculated, and adaptive fusion weights are assigned to each frame of low-light visible light images to obtain the temporally fused low-light visible light images. Local contrast enhancement is performed on the temporally fused low-light visible light image to obtain a contrast-enhanced image, so as to adaptively adjust the contrast parameters of local regions in the low-light visible light image. The curve parameters are dynamically adjusted based on the brightness distribution characteristics of the contrast-enhanced image, and the contrast-enhanced image is subjected to adaptive curve mapping processing using curve parameters and a no-reference learning method to obtain the low-light enhancement image.

3. The underwater robot target tracking method in low-light environments according to claim 1, characterized in that, The target detection based on feature maps of different resolutions identifies target clutter of a preset type, including: Bidirectional feature fusion, both top-down and bottom-up, is performed on feature maps of different resolutions to enhance the semantic information in high-resolution feature maps and the spatial detail information in low-resolution feature maps, resulting in fused multi-scale feature maps. Size-adaptive detection windows are set on the fused multi-scale feature maps. Based on historical detection data of fallen leaves, hair and paper scraps in the pool environment, the density, size and scale parameters of the detection windows on each scale feature map are dynamically adjusted. Among them, the distribution density of small-sized detection windows is increased on the high-resolution feature map to obtain a set of local feature maps. A lightweight convolutional network is used to classify and predict each local feature map, and outputs the target class confidence and bounding box coordinates of each local feature map in the local feature map set. The lightweight convolutional network is trained and optimized for the morphological features of fallen leaves, hair and paper scraps in a low-light underwater environment. The confidence threshold is dynamically adjusted by combining the ambient light intensity and water turbidity values ​​in the swimming pool environment. Based on the target category confidence and bounding box coordinates of each local feature map, an adaptive confidence threshold is used to filter the set of local feature maps to identify the target clutter set.

4. The underwater robot target tracking method in low-light environments according to claim 3, characterized in that, The process of assigning a unique identifier to each detected target object and continuously tracking the target object across frames by combining the similarity of the target's appearance features with motion trajectory prediction includes: A unique identifier is generated for each target object detected for the first time, and the initial position, bounding box size, category information and first detection timestamp of the target object are recorded; A multidimensional appearance feature vector is extracted from the bounding box region of each target clutter, and the multidimensional appearance feature vector includes local binary pattern features, color histogram distribution and depth appearance feature embedding. A motion state vector is constructed based on the historical position data of the target debris. The motion state vector includes position coordinates, velocity components, acceleration components, and motion direction angle. An adaptive Kalman filter is used to predict the expected position and uncertainty range of the target in the next frame. Calculate the comprehensive matching score between each newly detected target clutter in the current frame and the historical target clutter with assigned identity identifiers. The comprehensive matching score is a weighted combination of appearance feature similarity score and motion prediction consistency score. For the current detected target clutter whose comprehensive matching score exceeds a preset threshold, the identity is associated with the historical target clutter, and the status information of the historical target clutter is updated; for the newly detected target clutter that has not been matched, it is marked as the first detection state; for the historical target clutter that has not been matched with new detection results for multiple consecutive frames, the historical target clutter is marked as lost and the identity identifier resource is released.

5. The underwater robot target tracking method in low-light environments according to claim 1, characterized in that, The step of controlling the underwater robot's motion control system and cleaning device based on the target tracking results to track and grasp target debris in real time includes: Based on the position coordinates, direction of motion, and velocity information of the target debris in the target tracking results, the real-time relative positional relationship between the underwater robot and the target debris is calculated, and a smooth approach path that conforms to the robot's kinematic constraints is generated. When the relative distance between the underwater robot and the target debris is less than a preset distance threshold, the operating parameters of the cleaning device are adaptively adjusted according to the type of the target debris. During the grasping process, the target position changes are monitored in real time, and the output parameters of the underwater robot's horizontal or vertical thrusters are dynamically adjusted, as well as the suction strength and directional parameters of the cleaning device are dynamically adjusted. When the detected target debris position deviation exceeds the offset threshold, a relocation mechanism is triggered to ensure that the target debris is completely captured and collected into the storage device.

6. The underwater robot target tracking method in low-light environments according to claim 5, characterized in that, The adaptive adjustment of the cleaning device's operating parameters based on the type of target debris includes: For large areas of lightweight debris such as fallen leaves, activate the wide-area suction mode, adjust the suction range to a circular area with a radius of 0.7-1.0 meters, and set the suction intensity to 80-90 kPa. For fine, tangled hair and other debris, activate the rotating brush head and use the directional suction mode. Set the rotating brush head speed to 35-45 rpm and the suction intensity to 55-65 kPa, and dynamically adjust the suction direction to match the direction of hair extension. For small pieces of debris such as paper scraps, activate the fine-grained positioning suction mode, adjust the suction range to a circular area with a radius of 0.2-0.3 meters, and set the suction intensity to 35-45 kPa.

7. The underwater robot target tracking method in low-light environments according to claim 6, characterized in that, The process of real-time monitoring of target position changes during grasping, dynamically adjusting the output parameters of the underwater robot's horizontal and / or vertical thrusters, and dynamically adjusting the suction strength and direction parameters of the cleaning device includes: Establish a time series for tracking the position of the target debris. By calculating the difference between the position change vector of the target debris and the pose change of the underwater robot in consecutive frames, the actual motion trajectory of the target debris relative to the underwater robot can be obtained. A sliding window mechanism is used to perform time-series analysis on the target's motion trajectory. When the rate of change of the target's motion direction exceeds a preset angle change rate threshold or the acceleration exceeds a preset acceleration threshold, a thruster compensation control signal is generated. The amplitude and phase of the compensation control signal are adaptively adjusted according to the motion characteristic parameters corresponding to the type of target debris. Construct a spatial mapping table of relative positions between the underwater robot and the target debris, and determine the thruster parameters based on the polar coordinates of the relative position vector. The thruster parameters include radial thrust components and angular motion control components. The radial thrust components are used to control the propulsion speed, and the angular motion control components are used to control the turning rate. Based on the category identification results and current relative position of the target debris, a suction field distribution model of the cleaning device is generated. The suction field distribution model dynamically calculates the optimal suction intensity distribution and direction of action using fluid dynamics simulation formulas based on the physical characteristic parameters of the target debris. The physical characteristic parameters include at least one of density, shape parameters, and material. Based on the optimal suction intensity distribution and direction of action, corresponding suction intensity and direction parameters are set.

8. A target tracking system for underwater robots in low-light environments, characterized in that, The system is used to implement the underwater robot target tracking method in low-light environments as described in claim 1, and the system includes: The acquisition module is used to simultaneously acquire image data of the underwater environment through a visible light camera and an infrared sensor mounted on the underwater robot. The image data includes weak light visible light images and infrared images. The enhancement module is used to perform adaptive temporal fusion, contrast enhancement, and adaptive curve mapping processing on continuously acquired multi-frame low-light visible light images to obtain low-light enhanced images. The fusion module is used to extract surface texture features of objects from the low-light enhanced image, extract thermal distribution features from continuously acquired infrared images, and fuse the surface texture features and thermal distribution features through adaptive weights to generate a fused feature image; the adaptive weights are dynamically adjusted according to the edge morphology characteristics and thermal characteristics differences of target debris in the pool environment. The tracking module is used to identify and continuously track target debris of a preset type in the fused feature image to obtain target tracking results; the types of target debris include fallen leaves, hair and paper scraps; based on the target tracking results, the motion control system and cleaning device of the underwater robot are controlled to track and grasp the target debris in real time.

Citation Information

Patent Citations

  • Underwater target tracking and identifying method

    CN109961012A

  • Mechanical arm autonomous mobile grabbing method under complex illumination conditions based on visual-tactile fusion

    WO2023056670A1