Swimming pool robot positioning and trajectory prediction method and system fusing sonar and vision

By combining terahertz sonar with vision sensors, a panoramic model is built, which solves the problems of inaccurate distance measurement and difficulty in maintenance of swimming pool robots, and realizes high-precision positioning and intelligent management.

CN120489100AActive Publication Date: 2025-08-15YITUO ELECTRIC CO LTD

Patent Information

Application Number
CN202510700015.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-15
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The distance measurement accuracy of existing swimming pool robots is not high, and it is difficult to maintain and manage. It is likely to cause failure due to factors such as water quality, water temperature, and impurities in the water.

Method used

The terahertz sonar device is used to collect distance information, combine the multi-spectral and event camera collaborative acquisition strategy of vision sensors, and through deep learning and quantum machine learning algorithms, a panoramic model is built and updated in real time to realize the positioning and trajectory prediction of the swimming pool robot.

Benefits of technology

It improves the accuracy of ranging accuracy and trajectory prediction, reduces maintenance costs, and enhances the robot's independent operation ability and management efficiency in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120489100A_ABST
    Figure CN120489100A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a swimming pool robot positioning and trajectory prediction method and system fusing sonar and vision. The method comprises the following steps: collecting terahertz sonar data for a target area; based on terahertz sonar data, constructing a motion track model of the swimming pool robot; acquiring multi-source image data by adopting a multi-spectrum and event camera collaborative acquisition strategy; a dynamic SLAM algorithm based on deep learning is adopted, dynamic objects in the swimming pool environment are recognized and processed in real time based on the multi-source image data, and the environment image data with the dynamic objects removed are adopted to obtain a swimming pool environment model; performing deep fusion on the motion track model and the swimming pool environment model through a knowledge graph network, and predicting the operation state of the swimming pool robot by adopting a quantum machine learning algorithm to obtain a panoramic model; and creating and updating a digital twinborn model corresponding to the panoramic model in real time. The real-time performance and accuracy of positioning and track prediction of the swimming pool robot are improved, and a user can conveniently maintain and manage the swimming pool robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data prediction, and in particular to a method and system for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision. Background Art

[0002] Currently, a pool robot is an intelligent device specially designed for use in a swimming pool environment that can autonomously complete a series of tasks.

[0003] In the related art, although ultrasonic sensors are commonly used distance measuring sensors for swimming pool robots, their accuracy is easily affected by factors such as water quality, water temperature, and impurities in the water. For example, when the water temperature is high or the water quality is turbid, the propagation speed of ultrasonic waves will change, resulting in an increase in ranging errors. Moreover, the beam angle of ultrasonic sensors is large and easily affected by interference from the surrounding environment. For some smaller or closer objects, it may not be possible to accurately measure their distance. In addition, the swimming pool environment is complex, and impurities and microorganisms in the water can easily adhere to the surface of the robot, affecting its appearance and heat dissipation performance, and may even clog the sensor and water inlet, causing robot failure and making robot maintenance more difficult. Therefore, it is urgent to design a technical solution to overcome at least one technical problem existing in the related art. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to provide a method and system for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision, aiming to solve the technical problems of low ranging accuracy of swimming pool robots and high difficulty in maintenance and management of swimming pool robots.

[0005] In a first aspect, embodiments of the present application provide a method for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision, including:

[0006] The terahertz sonar device deployed in the robot collects terahertz sonar data for the target area to detect the distance information between the pool robot and surrounding objects in the target area; wherein the target area is the water environment where the pool robot is located;

[0007] constructing an activity trajectory model of the swimming pool robot based on the terahertz sonar data;

[0008] By deploying visual sensors around the pool and in the water environment, and using a multispectral and event camera collaborative acquisition strategy, we can acquire multi-source image data covering underwater, surface, and pool environment changes in the target area.

[0009] A deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, and a swimming pool environment model is constructed by removing the environmental image data of the dynamic objects; the dynamic objects include swimmers or moving objects in the environment;

[0010] Through the knowledge graph network, the activity trajectory model is deeply integrated with the swimming pool environment model, and the operation status of the swimming pool robot is predicted using a quantum machine learning algorithm to obtain a panoramic model including the predicted activity trajectory and predicted operation status of the swimming pool robot;

[0011] Create and update the digital twin model corresponding to the panoramic model in real time to achieve real-time simulation and visualization of the swimming pool robot's operating status and environmental changes, facilitating maintenance and management decisions for the swimming pool robot.

[0012] In a second aspect, embodiments of the present application provide a method and system for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision, including:

[0013] The first acquisition module, a terahertz sonar device deployed in the robot, is used to collect terahertz sonar data for a target area to detect the distance information between the pool robot and surrounding objects in the target area; wherein the target area is the water environment in which the pool robot is located;

[0014] A first building module is used to build an activity trajectory model of the swimming pool robot based on the terahertz sonar data;

[0015] The second acquisition module, a visual sensor deployed around the pool and in the water environment, is used to acquire multi-source image data including underwater, surface, and pool environment changes in the target area using a multispectral and event camera collaborative acquisition strategy;

[0016] A second construction module is configured to employ a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, and to construct a swimming pool environment model by using the environmental image data without the dynamic objects; the dynamic objects include swimmers or moving objects in the environment;

[0017] A fusion module is used to deeply fuse the activity trajectory model with the swimming pool environment model through a knowledge graph network, and use a quantum machine learning algorithm to predict the operating status of the swimming pool robot to obtain a panoramic model including the predicted activity trajectory and predicted operating status of the swimming pool robot;

[0018] The display module is used to create and update the digital twin model corresponding to the panoramic model in real time, realize real-time simulation and visualization of the operating status and environmental changes of the swimming pool robot, and facilitate maintenance and management decisions of the swimming pool robot.

[0019] In a third aspect, an embodiment of the present application further provides a terminal device, which includes a processor and a memory for storing a computer program; the processor is used to execute the computer program and implement the swimming pool robot positioning and trajectory prediction method that integrates sonar and vision as described in the first aspect or any embodiment of the present application when executing the computer program.

[0020] The embodiment of the present application provides a method and system for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision. In this method, terahertz sonar data for a target area is collected by deploying a terahertz sonar device in the robot to detect the distance information between the swimming pool robot and surrounding objects in the target area; wherein the target area is the water environment where the swimming pool robot is located; based on the terahertz sonar data, an activity trajectory model of the swimming pool robot is constructed; through the visual sensors deployed around the swimming pool and in the water environment, a multi-spectral and event camera collaborative acquisition strategy is adopted to obtain multi-source image data containing changes in the underwater, water surface and swimming pool environment in the target area; a dynamic SLAM algorithm based on deep learning is used to identify and process objects in real time based on the multi-source image data. The system processes dynamic objects in the swimming pool environment and constructs a swimming pool environment model by removing environmental image data of the dynamic objects; the dynamic objects include swimmers or active objects in the environment; through the knowledge graph network, the activity trajectory model is deeply integrated with the swimming pool environment model, and the operation status of the swimming pool robot is predicted by the quantum machine learning algorithm to obtain a panoramic model including the predicted activity trajectory and predicted operation status of the swimming pool robot; the digital twin model corresponding to the panoramic model is created and updated in real time to realize real-time simulation and visualization of the operation status of the swimming pool robot and environmental changes, which is convenient for maintenance and management decision-making of the swimming pool robot.

[0021] The embodiments of the present application can achieve an intelligent upgrade of the ranging and positioning process of the swimming pool robot from aspects such as environmental perception, data processing, model building to operation prediction and visual management. In terms of positioning accuracy, the fusion of terahertz sonar and multi-spectral visual data, combined with advanced algorithms, improves the accurate determination of the robot's position in complex environments. Secondly, in trajectory prediction, quantum machine learning and multi-model fusion technology significantly improve the accuracy and real-time performance of the prediction. In terms of environmental modeling capabilities, the three-dimensional map constructed by dynamic SLAM technology can reflect changes in the swimming pool environment in real time. Overall, the embodiments of the present application effectively improve the autonomous operation capability, environmental adaptability and intelligent management level of the swimming pool robot, and provide reliable and efficient technical support for tasks such as swimming pool cleaning and monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flow chart of a method for positioning and trajectory prediction of a swimming pool robot integrating sonar and vision provided in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of the module structure of a swimming pool robot positioning and trajectory prediction system that integrates sonar and vision, provided in an embodiment of the present application;

[0024] Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] In response to the technical problems existing in the related art, the embodiments of the present application propose a swimming pool robot positioning and trajectory prediction method and system that integrates sonar and vision.

[0026] Specifically, data is collected through the terahertz sonar device deployed in the robot. Terahertz waves have broadband and low photon energy characteristics, and can penetrate interfering substances such as foam on the water surface, achieving super-resolution distance detection in a low signal-to-noise ratio environment. Compared with traditional sonar, this technology can more accurately obtain the distance information between the swimming pool robot and surrounding objects, especially the detection ability of tiny foreign objects or underwater structural details is significantly improved, providing a high-precision data foundation for the subsequent trajectory model construction. Secondly, an activity trajectory model is constructed based on terahertz sonar data, combined with the robot's built-in inertial measurement unit (IMU) data, and a spatiotemporal Transformer network is used to mine the correlation between data in the time and space dimensions. At the same time, the model parameters are dynamically optimized through the reinforcement learning algorithm, so that the model can adapt to different swimming pool environments and robot motion modes, accurately depict the robot's motion trajectory in the swimming pool, and provide reliable motion information for operation status prediction. Thirdly, a multi-spectral and event camera collaborative acquisition strategy is adopted to obtain multi-source image data. Multispectral cameras capture images in different wavelengths, enhancing the recognition of complex environmental entities such as transparent floating objects and underwater pipelines. Event cameras are sensitive to dynamic changes. Together, these two systems comprehensively capture changes in the underwater, surface, and pool environments. Furthermore, an intelligent triggering system and dynamic sampling strategy based on environmental awareness reduce redundant data collection, improving data acquisition efficiency and quality. Next, a deep learning-based dynamic SLAM algorithm, utilizing a hybrid architecture of a spatiotemporal graph convolutional network (ST-GCN) and a Transformer, combines multi-source image data to achieve accurate recognition and real-time tracking of dynamic objects. Assisted by a conditional random field (CRF) algorithm and event camera data, dynamic objects are precisely segmented and removed, resulting in clean static environmental data for map construction. Parameters are then optimized using the sparse bundle adjustment (SBA) algorithm and deep reinforcement learning, along with incremental point cloud fusion technology, to update the map in real time. This generates an accurate and dynamic 3D pool environment model, providing a reliable environmental reference for path planning and trajectory prediction. Furthermore, knowledge graph technology is used to construct a knowledge graph network containing semantic information about the pool environment and the robot's motion rules, integrating knowledge from multiple sources. The graph attention mechanism (GAT) is used to calculate the information weights of different models, realize the intelligent fusion of the activity trajectory model and the pool environment model, and fully explore the correlation information between the models. Based on the quantum machine learning algorithm, the data is mapped to the quantum feature space, and the superposition and entanglement characteristics of quantum states are used to efficiently process high-dimensional complex data. The quantum support vector machine (QSVM) or quantum neural network (QNN) is selected for training and prediction. This can more accurately predict the operating status of the pool robot and generate a panoramic model that includes predicted activity trajectories and operating conditions, thereby improving the accuracy and foresight of the prediction. Finally, a digital twin model corresponding to the panoramic model is created and updated in real time. Combined with augmented reality (AR) and mixed reality (MR) technologies, real-time interaction between the virtual model and the physical robot is achieved through edge computing and 5G communication.By mapping the robot's operational and environmental perception data to a virtual model in real time, the system simulates and visualizes the pool robot's operational status and environmental changes. Reinforcement learning simulations based on the virtual model provide optimal solutions for robot maintenance and management decisions, enabling intelligent operation and maintenance, reducing maintenance costs, and improving management efficiency.

[0027] The present application provides a method and system for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision. This method can be applied to terminal devices, such as mobile phones, virtual reality devices, tablet computers, laptop computers, desktop computers, wearable devices, and other electronic devices. The terminal device can be a server connected to a tower crane or a server cluster. This connection can be achieved through hardware circuitry or a communication module.

[0028] The following is a detailed description of some embodiments of the present application in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Figure 1 , Figure 1 A flow chart of a method for positioning and trajectory prediction of a swimming pool robot that integrates sonar and vision, provided in an embodiment of the present application.

[0029] like Figure 1 As shown, the swimming pool robot positioning and trajectory prediction method integrating sonar and vision includes the following steps S101 to S106.

[0030] Step S101: Using a terahertz sonar device deployed in the robot, terahertz sonar data for a target area is collected to detect distance information between the swimming pool robot and surrounding objects in the target area.

[0031] Step S102: constructing an activity trajectory model of the swimming pool robot based on the terahertz sonar data.

[0032] In the embodiment of the present application, the target area is the water environment where the swimming pool robot is located.

[0033] In the embodiment of the present application, terahertz waves have the characteristics of broadband and low photon energy. Combining terahertz technology with sonar can achieve higher-resolution distance measurement, penetrate interfering substances such as foam on the water surface, and accurately detect tiny foreign objects in the swimming pool or underwater structure details.

[0034] Terahertz sonar devices primarily use the characteristics of terahertz waves to detect distances. The principle is as follows: Terahertz waves are electromagnetic waves between microwaves and infrared, with broadband and low photon energy. The device's transmitting module generates and transmits terahertz waves toward the target area. When the terahertz waves encounter objects in the pool (such as pool walls, floating objects, underwater pipes, etc.), they are reflected. The receiving module captures these reflected terahertz wave signals and converts them into electrical or digital signals. By calculating the time delay from terahertz wave transmission to reception, combined with the propagation speed of terahertz waves in water, the distance between the pool robot and surrounding objects can be accurately calculated.

[0035] Terahertz sonar data primarily includes the following: Distance information is the most core data, clearly recording the straight-line distance between the pool robot and each detected object. This is the key basis for determining the robot's relative position to its surroundings. Signal strength reflects the strength of the reflected terahertz wave signal. Signal strength is related to factors such as the object's material, surface roughness, and distance. For example, objects with smooth surfaces reflect stronger signals, and the signal strength decreases with distance. This information can help determine the object's material and size. Timestamps record the specific moments of terahertz wave transmission and reception, which are used to calculate time delays. They also provide a time dimension reference for subsequent data processing and analysis, making it easier to understand the robot's detection status at different times. Waveform features include characteristic information such as the waveform shape and frequency of the reflected signal. The waveforms of terahertz waves reflected by different objects vary. By analyzing these features, the type and properties of the object can be further identified.

[0036] For example, in step S101, when the pool robot performs underwater cleaning tasks, it needs to detect in real time whether there are obstacles in front of it (such as the pool wall, floating toys, underwater pipes, etc.) to avoid collisions and plan the cleaning path. The terahertz sonar device transmitting module at the bottom of the pool robot transmits a beam of terahertz waves (for example, electromagnetic waves with a frequency of 0.1 THz) to the front. The wave has a frequency of about 2.25×10 8 m / s in water (slightly slower than the speed of light in air). When the terahertz wave encounters the pool wall 2 meters ahead, part of the wave is reflected back to the robot. The receiving module captures the reflected signal and records the signal transmission time (t1) and reception time (t2). For example, the time delay Δt = t2 - t1 = approximately 1.78 × 10 -8 seconds (calculated based on a distance of 2 meters: Δt=2×2 / 2.25×10 8 ≈1.78×10 -8 seconds, the round-trip distance needs to be multiplied by 2). According to the formula, distance = (propagation speed × time delay) / 2, the distance between the robot and the pool wall is calculated to be 2 meters.

[0037] At the same time, the device records the reflected signal's strength (e.g., -30dBm, reflecting the smoothness of the pool wall's surface) and waveform characteristics (used to subsequently distinguish between object materials, for example, the waveform of a concrete pool wall differs from that of a floating plastic object). Based on the detected 2-meter distance, the robot determines that it does not need to turn immediately and continues cleaning along the original path. If the distance falls below a safety threshold (e.g., 0.5 meters), a turn command is triggered. Multiple acquisitions of distance data (e.g., the continuous position of the pool wall) can be used to construct a pool contour map to assist in path planning.

[0038] Optionally, in step S101, after collecting terahertz sonar data for the target area using a terahertz sonar device deployed in the robot, swimming pool environmental data may also be acquired. This swimming pool environmental data includes at least: venue type, water turbidity, time-varying swimming pool lighting conditions, and weather information in the swimming pool area. Next, dynamic changes in the swimming pool environment are predicted based on this pool environmental data, and sonar noise reduction parameters are adjusted based on the prediction results. Dynamic changes in the swimming pool environment include: trends in water visibility, lighting conditions, and physical environmental disturbances caused by weather changes. Finally, the optimized and adjusted sonar noise reduction parameters are used to perform real-time noise reduction on the terahertz sonar data to improve sonar data quality.

[0039] Specifically, when terahertz sonar equipment collects data from a target area, environmental factors in the pool can interfere with data quality. Different venue types have different spatial structures and echo characteristics. For example, the enclosed space of an indoor swimming pool differs from the open environment of an outdoor swimming pool, resulting in different propagation paths and interference sources for sonar reflection signals. Water turbidity affects sound wave propagation; higher turbidity increases sound wave scattering and absorption. Fluctuating pool lighting conditions and weather factors can alter water surface fluctuations and ambient temperature, indirectly affecting sonar data. First, various pool environmental data, such as venue type and water turbidity, are collected. Data analysis and prediction algorithms are then used to analyze correlations between these environmental data and predict dynamic trends in factors such as water visibility, lighting, and physical disturbances. For example, weather information can be used to predict strong winds causing significant water surface fluctuations, and lighting trends can be used to estimate periods of increased reflective interference. Based on these predictions, sonar noise reduction parameters, such as filter intensity and frequency range, can be adjusted according to pre-set rules or machine learning models. Finally, the optimized parameters are used to reduce the noise of terahertz sonar data in real time, remove the noise caused by environmental interference, improve data quality, and provide reliable data support for subsequent robot positioning, target detection and other tasks.

[0040] Understandably, in real-world applications, sudden weather changes in outdoor swimming pools, such as strong winds, can cause dramatic surface fluctuations, generating a large amount of interfering sound waves. At the same time, cloud cover can alter lighting conditions. Based on weather information and changing lighting conditions, the system predicts increased physical disturbances and the potential for reflection interference caused by these changes. It proactively adjusts sonar noise reduction parameters to enhance filtering of high-frequency noise. When collecting data, the terahertz sonar device effectively suppresses noise generated by surface fluctuations, providing a clear view of the pool floor and the robot's surroundings. This prevents positioning errors and misidentification of targets caused by noise interference, enabling the robot to accurately perform cleaning tasks and efficiently avoid obstacles. For example, in turbid swimming pools, the system, based on water turbidity data, predicts reduced water visibility and increased sound wave scattering, adjusting sonar noise reduction parameters to optimize signal reception and processing. Even after noise reduction, the terahertz sonar data can still accurately identify objects such as the pool wall and underwater obstacles, providing the robot with precise environmental perception and ensuring stable operation in complex water conditions. This significantly improves the robot's adaptability and reliability in diverse swimming pool environments.

[0041] Further optionally, in the above steps, predicting the dynamic changes of the swimming pool environment according to the swimming pool environment data, and adjusting the sonar noise reduction parameters based on the prediction results include:

[0042] The various environmental data including venue type, water turbidity, swimming pool lighting conditions, and weather information are input into the adaptive deep learning model in the form of spatiotemporal graph nodes combined with sequence data; the dynamic capture module in the adaptive deep learning model is used to capture the correlation between the various environmental data in the spatial dimension to obtain spatiotemporal environmental correlation features, and the spatiotemporal correlation module is used to establish the long-range dependency relationship between the various environmental data in the time dimension to obtain long-range dependency features; the multi-scale feature fusion module is used to combine the high-frequency detail features and low-frequency structural features in the terahertz sonar data with the features used to represent the swimming pool. The spatiotemporal environmental correlation features and the long-range dependency features of the dynamically changing environment are cross-dimensionally fused to obtain environmental fusion features. The signal-to-noise ratio improvement value and positioning accuracy improvement rate of the denoised sonar data are used as reward functions, and the sonar noise reduction parameter adjustment strategy is optimized through a reinforcement learning module. Combined with the sonar noise reduction parameter adjustment strategy obtained by reinforcement learning, the convolution kernel size and the hidden layer parameters of the recursive neural network used in the noise reduction process are dynamically optimized and adjusted, and the optimized sonar noise reduction parameters are used to perform real-time noise reduction on the environmental fusion features to filter out noise data in the terahertz sonar data.

[0043] Specifically, environmental data such as venue type, water turbidity, pool lighting conditions, and weather information are converted into a combination of spatiotemporal graph nodes and sequence data as input to the adaptive deep learning model. The dynamic capture module analyzes the spatiotemporal graph nodes through a graph neural network (GNN) to explore the spatial correlation between different environmental data, such as how the enclosed structure of indoor venues affects sound wave reflection, and the correlation between water turbidity and venue ventilation conditions, thereby extracting spatiotemporal environmental correlation features. The spatiotemporal correlation module uses sequence modeling structures such as long short-term memory networks or Transformers to process sequence data and establish long-range dependencies between environmental data in the temporal dimension. For example, it predicts the subsequent impact of weather changes on pool lighting and water surface fluctuations, thereby obtaining long-range dependency features.

[0044] The multi-scale feature fusion module decomposes the terahertz sonar data to obtain high-frequency detail features (such as reflection signals from tiny obstacles) and low-frequency structural features (such as the overall outline signal of the swimming pool wall). It then integrates these features with spatiotemporal environmental correlation features and long-range dependency features across dimensions to form environmental fusion features that incorporate both dynamic environmental changes and sonar data characteristics. The reinforcement learning module uses the improvement in the signal-to-noise ratio and positioning accuracy of the denoised sonar data as reward functions. Through trial and error and learning, it explores the optimal sonar noise reduction parameter adjustment strategy. For example, it tries different combinations of parameters such as filter strength and frequency range, and optimizes the strategy based on reward feedback. Finally, combined with the strategy learned from reinforcement learning, it dynamically adjusts the convolution kernel size and the hidden layer parameters of the recurrent neural network during the noise reduction process, performing real-time noise reduction on the environmental fusion features, specifically filtering out noise in the terahertz sonar data and retaining valid signals.

[0045] For example, in an outdoor swimming pool scenario, weather information for the day indicates strong winds in the afternoon, high water turbidity, and strong midday sunlight. This environmental data is converted into spatiotemporal graph nodes and sequence data, which are then fed into an adaptive deep learning model. The motion capture module analyzes and finds that strong winds exacerbate water surface fluctuations, while strong sunlight enhances surface reflections. These two interact spatially, increasing interference with sonar data. The spatiotemporal correlation module predicts that with the onset of strong winds, water surface fluctuations will continue to intensify, and changes in light angle will also cause changes in the reflective area.

[0046] The multi-scale feature fusion module processes terahertz sonar data, integrating high-frequency reflection details of water ripples and low-frequency contours of the pool bottom with the aforementioned environmental features. The reinforcement learning module, through continuous experimentation with different noise reduction parameters, such as adjusting the filter convolution kernel size and the number of neurons in the recurrent neural network's hidden layer, discovered that increasing the filter intensity in specific frequency bands and adjusting the convolution kernel size can effectively reduce noise interference caused by water surface fluctuations and reflections, improving the signal-to-noise ratio and positioning accuracy. This led to the determination of the optimal noise reduction parameter adjustment strategy. Ultimately, the optimized parameters were used to reduce the noise of the sonar data in real time, enabling the robot to clearly identify the pool bottom and surrounding obstacles, and accurately plan its movement path even in harsh environments.

[0047] Thus, the above steps significantly improve the accuracy and adaptability of terahertz sonar data processing. In the complex and ever-changing swimming pool environment, through in-depth analysis and feature fusion of environmental data, it is possible to predict the impact of environmental changes on sonar data in advance and optimize noise reduction parameters in a targeted manner. Compared with traditional fixed parameter noise reduction methods, this can improve the signal-to-noise ratio of sonar data, effectively reduce positioning errors caused by noise, improve positioning accuracy, reduce the risk of misjudgment and collision of robots caused by data noise, and improve their working stability and reliability in different venue types, weather conditions, and water quality conditions, ensuring that swimming pool robots can efficiently complete cleaning, inspection and other tasks.

[0048] For example, when water visibility is predicted to decrease, the reinforcement learning module prioritizes adjusting the filter intensity parameters of the denoising algorithm. If lighting conditions change dramatically, the focus is on optimizing the attention mechanism parameters of the deep learning-based denoising model, making the denoising process more targeted.

[0049] Further optionally, in step S102, constructing an activity trajectory model of the swimming pool robot based on the terahertz sonar data includes:

[0050] A spatiotemporal Transformer model based on a hybrid architecture of a spatiotemporal graph convolutional network (STGCN) and a Transformer network is constructed; the inertial measurement module (IMU) built into the swimming pool robot is obtained to obtain inertial measurement data; the spatiotemporal Transformer model is used to simulate the activity trajectory of the swimming pool robot based on the inertial measurement data to construct an initial activity trajectory model of the swimming pool robot; the distance relationship between the swimming pool robot and surrounding objects is simulated based on the terahertz sonar data to optimize the initial activity trajectory model, so that the initial activity trajectory model is adaptive to different swimming pool environments and robot motion modes, and the final output activity trajectory model is obtained.

[0051] In the embodiment of the present application, the activity trajectory model is a mathematical model used to describe the motion trajectory of the pool robot in the pool environment, and plays a key role in the entire positioning and trajectory prediction method. The activity trajectory model construction method realizes accurate trajectory modeling through the fusion of network architecture and multi-source data. Specifically, it is constructed based on terahertz sonar data and the robot's built-in inertial measurement unit (IMU) data. Terahertz sonar data provides distance information between the robot and surrounding objects, and clarifies the positional relationship of the robot in the environment; IMU data reflects the robot's own acceleration, angular velocity and other motion state changes. A spatiotemporal Transformer network is adopted, that is, a hybrid architecture of a spatiotemporal graph convolutional network (STGCN) and a Transformer network. STGCN mines the spatial correlation and temporal dynamic change rules of data, and the Transformer network uses the attention mechanism to globally model long sequence data and capture complex dependencies. The combination of the two effectively processes the spatiotemporal characteristics and long-term information of the robot's motion data. First, the robot's activity trajectory is simulated based on inertial measurement data through the spatiotemporal Transformer model to obtain an initial activity trajectory model. Then, the distance relationship between the robot and surrounding objects is simulated with terahertz sonar data. The initial model is optimized and the robot's position parameters and movement direction are adjusted to make the model adapt to different swimming pool environments and robot motion modes, and finally an accurate activity trajectory model is output.

[0052] Obviously, the model accurately depicts the robot's motion trajectory in the swimming pool, providing a reliable motion information basis for subsequent operation status prediction, and also helps the robot to plan the path. At the same time, it provides key data for the digital twin model, helping to realize intelligent operation and maintenance.

[0053] Among them, the Spatiotemporal Graph Convolutional Network (STGCN) excels at processing data with spatial topological structures and time series characteristics. It can mine spatial correlations (such as the relative positional relationship between the robot and surrounding objects) and temporal dynamics of swimming pool robot motion data. The Transformer network, with its powerful attention mechanism, can globally model long-sequence data and capture complex dependencies within the data. The hybrid architecture of the two can not only process the spatiotemporal characteristics of robot motion data, but also effectively analyze long-term motion information, providing a powerful tool for trajectory modeling.

[0054] The inertial measurement unit (IMU) acquires the robot's acceleration, angular velocity, and other inertial measurement data in real time, reflecting changes in the robot's motion state. Terahertz sonar data provides information on the distance between the robot and surrounding objects, clarifying the robot's positional relationship within the pool environment. Using the spatiotemporal Transformer model, inertial measurement data is first used to initially simulate the robot's trajectory and construct an initial model. This initial model is then optimized using terahertz sonar data, allowing the model to comprehensively consider the robot's motion state and external environmental factors, allowing it to adapt to different pool environments and motion patterns.

[0055] When optimizing the initial trajectory model based on terahertz sonar data, the robot's position parameters, movement direction, and other parameters are adjusted within the model by calculating the distance between the robot and surrounding objects. For example, if the terahertz sonar detects an obstacle ahead, the model will modify the robot's original trajectory based on this distance information, circumventing the obstacle. This continuously optimizes the model and improves the accuracy of trajectory modeling.

[0056] For example, assume that a pool robot starts cleaning a standard swimming pool. In the initial stage, the inertial measurement module (IMU) collects the robot's acceleration data in real time at 0.5 m / s. 2 , with an angular velocity of 10° / s. These inertial measurement data are input into the space-time Transformer model to preliminarily simulate the robot's motion trajectory in the swimming pool and obtain the initial activity trajectory model, which shows that the robot is moving in a straight line in a certain direction of the swimming pool in its current state. As the robot moves, the terahertz sonar device continuously collects data and detects a floating object 2 meters ahead. At this time, based on the distance relationship between the robot and the floating object in the terahertz sonar data, the initial activity trajectory model is optimized. The model adjusts the robot's movement direction so that the trajectory bypasses the floating object and replans a new trajectory, thereby obtaining the final output activity trajectory model to ensure that the robot can perform tasks safely and efficiently in the complex swimming pool environment.

[0057] Therefore, by combining inertial measurement data and terahertz sonar data, a comprehensive model is built from two levels: the robot's own motion state and the external environment. Compared with a single data source, it can more accurately depict the robot's activity trajectory, providing a reliable basis for subsequent operation state prediction, path planning, etc. By optimizing the model with terahertz sonar data, the activity trajectory model can automatically adjust the trajectory according to changes in the swimming pool environment (such as the appearance of new obstacles, different pool shapes) and changes in the robot's motion mode (such as acceleration and turning), thereby enhancing the model's adaptability in different scenarios and improving the robot's autonomous operation capability in complex swimming pool environments. The hybrid architecture of the spatiotemporal Transformer model fully leverages the advantages of STGCN and Transformer networks, can quickly process large amounts of spatiotemporal data, and realize real-time modeling and updating of the robot's activity trajectory, meeting the swimming pool robot's requirements for timely trajectory modeling in a dynamic environment and improving the overall system's operating efficiency.

[0058] Step S103: By deploying visual sensors around the swimming pool and in the water environment, a multispectral and event camera collaborative acquisition strategy is adopted to obtain multi-source image data including underwater, water surface and swimming pool environment changes in the target area.

[0059] In the embodiments of the present application, the visual sensors surrounding the pool can be cameras installed at locations such as the pool's edge, ceiling, or walls. These sensors are used to monitor the overall pool condition from various angles, including the water surface, the pool's surroundings, and any potential anomalies. For example, they can detect whether someone has accidentally fallen into the water or whether objects have been dropped around the pool. These sensors typically have a wide field of view, capable of covering a large area of the pool, but their ability to monitor underwater conditions is relatively limited.

[0060] Visual sensors in aquatic environments primarily refer to cameras or other optical sensors installed underwater. These sensors can directly capture underwater image information, including the condition of the pool's bottom and walls, the status of underwater facilities, and objects within the water. For example, they can be used to detect stains and cracks on the pool's bottom or identify the movements of swimmers underwater. To adapt to underwater environments, these sensors typically need to be waterproof and pressure-resistant.

[0061] For example, in a collaborative acquisition strategy involving multispectral and event cameras, a multispectral camera can capture light across different wavelengths, thereby acquiring image information from multiple spectral channels. In a swimming pool environment, different substances have varying absorption and reflection characteristics for different wavelengths of light. Multispectral imaging can leverage these characteristics to distinguish between different objects or scene features. For example, using light of specific wavelengths can enhance the ability to identify water quality, underwater organisms, and stains. Differences in reflection across different spectra can also be used to better detect surface fluctuations, changes in light and shadow, and the outlines of underwater objects.

[0062] An event camera is a new type of visual sensor. Unlike traditional cameras, which capture images at a fixed frame rate, it instead records brightness changes in a scene based on pixel-level event triggering. In a swimming pool environment, an event camera can quickly capture dynamic changes in the scene, such as a swimmer's rapid movements, momentary fluctuations in the water surface, and sudden movements of objects, with extremely high temporal resolution. Because it only records brightness change events, it significantly reduces the amount of data required for processing dynamic scenes while also more accurately capturing rapidly occurring events.

[0063] The collaborative acquisition strategy combines multispectral cameras and event cameras to leverage their respective strengths. Multispectral cameras provide rich spectral information, facilitating detailed scene analysis and identification, but may be subject to frame rate limitations when capturing rapid dynamics. Event cameras, on the other hand, are well-suited to capturing rapidly changing events, but lack spectral information and a complete description of static scenes. Through collaborative acquisition, when the event camera detects a dynamic event in the scene, it triggers the multispectral camera to acquire hyperspectral images at the corresponding moment, or uses the information from the event camera to guide the adjustment of the multispectral camera's shooting parameters to better capture the multispectral information of the dynamic scene. At the same time, the images captured by the multispectral camera can provide the event camera with a static background and richer scene information, assisting the event camera in understanding and locating the event. In this way, the two complement each other to obtain more comprehensive and accurate multi-source image data encompassing changes in the underwater, surface, and swimming pool environments in the target area.

[0064] Optionally, the visual sensor further comprises a multispectral camera for acquiring image information in different wavelength bands. Based on the aforementioned assumption, by deploying visual sensors around the swimming pool and in the water environment, a multispectral and event camera collaborative acquisition strategy is employed to acquire multi-source image data encompassing underwater, surface, and swimming pool environmental changes in the target area, including:

[0065] Light sensors and water velocity sensors deployed in the swimming pool monitor environmental changes in real time to assess the complexity of the swimming pool environment. A dynamic sampling strategy is trained using a reinforcement learning algorithm to dynamically adjust the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor collaboration mode according to the complexity of the swimming pool environment. When a sudden change in light or an intensified water disturbance is detected, the multispectral camera and the event camera are automatically triggered to enter a high-frame-rate collaborative acquisition mode. The multispectral camera collects visual image data in different bands for the target area based on the event trigger threshold and the sensor collaboration mode. Based on the visual image data in different bands, complex environmental entities in the swimming pool environment are identified, and the identified complex environmental entities are subjected to image enhancement processing by a dynamic filter group to obtain optimized visual image data. The complex environmental entities include at least transparent floating objects and underwater pipes.

[0066] Specifically, light sensors can monitor real-time changes in light intensity and color, while water velocity sensors can monitor changes in water velocity and direction. This information reflects the dynamic nature of the pool environment and can be used to assess environmental complexity. For example, a sudden change in light may indicate the presence of an object or a change in lighting, while increased water disturbances may be caused by swimmer activity or equipment. A dynamic sampling strategy is trained using a reinforcement learning algorithm to adjust the multispectral camera's band combination, the event camera's event trigger threshold, and the sensor collaboration mode based on environmental complexity. Through trial and error, reinforcement learning learns the optimal sampling strategy for different environmental conditions, adapting to environmental changes and improving data collection efficiency. When a sudden change in light or increased water disturbance is detected, indicating rapid environmental changes, the multispectral camera and event camera are automatically triggered into high-frame-rate collaborative acquisition mode to more accurately capture these dynamic changes and avoid missing important information. Based on the set event trigger threshold and sensor collaboration mode, the multispectral camera collects visual image data in different bands. Different bands reflect different object and scene characteristics, facilitating the identification of entities in complex environments. Then, the image enhancement processing of the identified complex environment entities is performed through a dynamic filter bank to highlight the target features, suppress noise and interference, and improve image quality.

[0067] For example, imagine a large swimming pool filled with transparent plastic floats and a complex underwater piping system. When a swimmer passes quickly, the water velocity suddenly changes. This change is detected by the water velocity sensor, while the light sensor also detects the change in light reflection caused by the swimmer's movements. Based on these monitored environmental changes, a reinforcement learning algorithm adjusts the multispectral camera to select specific band combinations sensitive to transparent objects and underwater pipes for acquisition, while also lowering the event trigger threshold of the event camera, making it more sensitive to events related to these objects. After collecting image data in different bands, the multispectral camera uses an image recognition algorithm to identify complex environmental entities such as transparent floats and underwater pipes. It then uses a dynamic filter bank to enhance these entities, for example, by increasing the edge contrast of transparent floats and highlighting the outlines of underwater pipes, making them more clearly visible in the image.

[0068] For example, a recognition system based on a large multimodal model (e.g., a CLIP-like architecture) can be constructed, combining the spectral features of multispectral images with the spatiotemporal features of event cameras. Through comparative learning, data from different modalities can be aligned within the same semantic space. The Transformer's global attention mechanism can be used to capture subtle multimodal differences in complex environmental entities such as transparent floating objects and underwater pipes. Combined with a meta-learning algorithm, the model can quickly adapt to the changing characteristics of entities in different swimming pool environments, improving recognition accuracy.

[0069] For example, for identified complex environmental entities, an image enhancement method based on a conditional generative adversarial network (cGAN) and a dynamic filter bank is employed. The cGAN generates enhancement conditions based on the entity type (transparent floating objects, underwater pipes, etc.), while the dynamic filter bank adaptively adjusts filtering parameters based on the band characteristics of the multispectral image to perform targeted enhancement on the entity. Simultaneously, a reinforcement learning module is introduced, using the edge clarity and semantic consistency of the enhanced image as a reward function to optimize the enhancement process and obtain high-quality optimized visual image data.

[0070] In this way, by dynamically adjusting the acquisition strategy based on the complexity of the environment, useful data can be collected in a more targeted manner, avoiding over-acquisition in simple environments or data loss in complex environments, thereby improving the quality and effectiveness of image data. The multispectral camera's acquisition of different bands and image enhancement processing help to more accurately identify various complex environmental entities in the swimming pool environment, such as transparent floating objects and underwater pipes. Even if these objects are difficult to clearly distinguish under conventional vision, they can be clearly presented in the processed image, providing a good foundation for subsequent analysis and processing. It can adapt to dynamic changes in the swimming pool environment in real time, such as sudden changes in light and water flow, and promptly adjust the acquisition mode and parameters to ensure that the collected data can reflect the actual situation of the current environment, thereby improving the robustness and adaptability of the system.

[0071] Further optionally, in the above steps, a dynamic sampling strategy is trained using a reinforcement learning algorithm to dynamically adjust the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor collaboration mode according to the complexity of the swimming pool environment, including:

[0072] The multispectral camera adopts an adaptive band switching method to obtain the spatial and spectral correlation between images of each band according to the real-time swimming pool light, water turbidity, and weather conditions in the complexity of the swimming pool environment. Based on the correlation, the camera automatically adjusts the acquisition band combination, event trigger threshold, and sensor collaboration mode to ensure that visual image data adapted to different swimming pool environments is collected under different working conditions, reducing the amount of redundant data collection and improving data validity.

[0073] In principle, multispectral cameras and related sensors monitor the swimming pool environment in real time, obtaining information such as the pool's real-time lighting (such as light intensity and direction), water turbidity (reflecting water clarity), and weather conditions (such as sunny, cloudy, and rainy days). Multispectral cameras capture images in different wavelengths, which reflect objects of different materials and conditions differently. By analyzing the spatial and spectral correlations between images in each wavelength, the characteristics and distribution of objects in the swimming pool environment can be understood. For example, some wavelengths are more sensitive to transparent objects, while others better depict the details of underwater structures. Correlation analysis reveals the relationship between these wavelengths and their relationship to swimming pool environmental factors.

[0074] The reinforcement learning algorithm uses the complexity of the swimming pool environment as input, and its goal is to find the optimal dynamic sampling strategy through continuous trial and error and learning. The algorithm uses the acquisition band combination, event trigger threshold, and sensor collaboration mode as adjustable actions, and the effectiveness of the collected visual image data (such as the useful information contained in the data, the accuracy of identifying objects in the swimming pool environment, etc.) as a reward. During the training process, the algorithm selects actions based on the current environmental state (adjusting the sampling strategy), and then evaluates the quality of the actions based on the ability of the collected data to describe the swimming pool environment (reward), and continuously adjusts the strategy to maximize the reward. For example, when the water turbidity is high, the algorithm tries different band combinations and adjusts the subsequent sampling strategy based on the recognizability of underwater objects in the collected images to improve the effectiveness of the data.

[0075] Based on strategies derived from a reinforcement learning algorithm, the multispectral camera adaptively switches between acquisition band combinations, adjusts the event camera's event trigger threshold, and adjusts the sensor collaboration mode. For example, in low-light environments, the event camera's event trigger threshold is raised to reduce false triggers due to light interference. Simultaneously, the multispectral camera selects a band combination more suitable for low-light environments to obtain clearer image data. This adaptive adjustment ensures that visual image data tailored to the pool environment is captured under various operating conditions, reduces the amount of redundant data collected, and improves data validity and the ability to describe the pool environment.

[0076] For example, the real-time light in the swimming pool is strong and the water turbidity is low. The multispectral camera analyzes the spatial and spectral correlation between the images of each band and finds that certain bands (such as the visible light band) have good imaging effects on floating objects on the water surface and swimming pool facilities. The reinforcement learning algorithm selects the appropriate band combination (such as enhancing the acquisition of the visible light band) according to the current environmental state, and adjusts the event trigger threshold of the event camera (appropriately increased because the environment is relatively stable), as well as the sensor collaboration mode (such as increasing the synchronous acquisition frequency of the multispectral camera and the event camera). At this time, the collected image data is clear and can accurately identify objects in the swimming pool, with less redundant data and high data validity.

[0077] For example, the light is dim and the water turbidity is high. After the multispectral camera acquires images of each band, it analyzes the correlation and finds that the infrared band has a strong ability to penetrate underwater objects and can better capture underwater structures. The reinforcement learning algorithm adjusts the sampling strategy, switches to an acquisition band combination based on the infrared band, lowers the event trigger threshold of the event camera (because light changes and turbid water may make the dynamic changes of objects difficult to capture), and adjusts the sensor collaboration mode (such as changing the acquisition time interval of the multispectral camera and the event camera). Through these adjustments, the collected image data can present underwater objects more clearly. Although the environment is complex, the adaptive sampling strategy reduces redundant data and improves the data's ability to describe the swimming pool environment. For example, it can accurately identify the position and shape of underwater pipes.

[0078] Based on the above principles, using reinforcement learning algorithms to train dynamic sampling strategies can enable multispectral cameras and event cameras to better adapt to different swimming pool environments and improve the quality and efficiency of data collection.

[0079] Step S104: Using a deep learning-based dynamic SLAM algorithm, dynamic objects in the swimming pool environment are identified and processed in real time based on the multi-source image data, and the swimming pool environment model is constructed by using the environmental image data without the dynamic objects.

[0080] In an embodiment of the present application, a deep learning-based dynamic SLAM algorithm uses a deep learning-based dynamic SLAM (simultaneous localization and mapping) algorithm to identify and process dynamic objects (such as swimmers) in a swimming pool environment in real time, avoid mistakenly incorporating them into map construction, generate a more accurate and dynamic three-dimensional map, and provide a more reliable environmental model for path planning and trajectory prediction.

[0081] Image data is acquired through visual sensors (such as the multispectral camera and event camera mentioned above), and deep learning algorithms are used to extract and match features in the images. For example, a convolutional neural network (CNN) is used to extract feature points such as corners and edges in the image. A feature matching algorithm is then used to find correspondences between different frames, thereby calculating the camera's pose changes and achieving positioning.

[0082] For example, based on positioning, the environmental information observed by the camera is integrated to construct a map. For dynamic SLAM, the key is to distinguish between dynamic objects and static environments. Deep learning target detection algorithms such as YOLO (You Only Look Once) or SSD (Single Shot MultiBox Detector) are used to identify dynamic objects (such as swimmers) in the image. Then, when constructing the map, the information of these dynamic objects is excluded, and only the image data of the static environment is used to construct the map, which can obtain a more accurate model of the swimming pool environment.

[0083] Deep learning object detection is used to identify dynamic objects. Taking the YOLO algorithm as an example, the input image is divided into multiple grids, each responsible for detecting the center position of an object. Convolutional and fully connected layers are used to extract and classify features from the image, predicting the object's category, location, and confidence score. In a swimming pool environment, this algorithm can quickly detect dynamic objects such as swimmers. Feature extraction is performed using deep learning models, such as deep learning versions of traditional feature extraction methods like SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF). These models learn more robust features that are unaffected by factors such as lighting and viewpoint. Feature matching determines correspondences between different frames by calculating the similarity between features. Common methods include matching algorithms based on Euclidean distance or cosine similarity. Optionally, the camera's pose change can be calculated based on the feature matching results. Methods such as the PnP (Perspective-n-Point) algorithm can be used to determine the camera's rotation and translation parameters using known feature point coordinates and camera intrinsic parameters, thereby determining the camera's pose at different times.

[0084] In addition to using target detection algorithms to identify dynamic objects, other methods can also be used to track their movement. For example, filtering algorithms such as Kalman filters can be used to predict and update the positions of dynamic objects, effectively eliminating their influence when building maps. Furthermore, for static objects that may be mistakenly identified as dynamic (such as floating objects in water), further semantic analysis or contextual information can be used to identify and correct them.

[0085] In this way, deep learning can learn rich image features, making it more accurate for identifying dynamic objects and modeling static environments. Compared to traditional SLAM algorithms, it reduces map errors caused by interference from dynamic objects. It has good adaptability to complex environments such as lighting changes and turbid water. During training, deep learning models can learn features from different environments and can stably perform positioning and map construction under various working conditions. With the continuous improvement of hardware computing power, deep learning-based algorithms can complete calculations in a shorter time and meet real-time requirements. For example, using a high-performance GPU can accelerate the inference process of the deep learning model, allowing the dynamic SLAM system to process image data in real time and generate a three-dimensional map of the swimming pool environment.

[0086] In the embodiment of the present application, the swimming pool environment model is a digital representation of the swimming pool and its surrounding environment. It is constructed based on information such as collected multi-source image data, and is intended to accurately present the static structure and characteristics of the swimming pool in the form of a three-dimensional map, etc., including the shape, size, depth, underwater facilities (such as underwater pipes), surrounding buildings or facilities, etc. of the swimming pool, providing a basis for subsequent path planning, trajectory prediction and other analyses and applications related to the swimming pool environment.

[0087] Specifically, visual sensors deployed around the pool and in the water environment first collect multi-source image data, including information about underwater, surface, and changes in the pool environment. A deep learning-based dynamic SLAM algorithm is then used to identify and process dynamic objects in the pool environment in real time. Using the environmental image data after removing dynamic objects as a basis, a series of image processing, feature extraction, matching, and 3D reconstruction techniques are used to gradually construct a 3D model of the pool environment. During the construction process, image information in different bands obtained by multispectral cameras is combined to enhance the ability to identify and model different objects in the environment. For example, by analyzing images in different bands, the location and shape of facilities such as underwater pipelines can be more accurately determined.

[0088] This provides accurate environmental information for pool cleaning robots, life-saving equipment, and other equipment, helping them plan reasonable paths of action, avoid collisions with obstacles, and complete their tasks efficiently. For example, a pool cleaning robot can plan a cleaning path that covers the entire bottom of the pool based on the environmental model, while avoiding fixed facilities such as underwater pipes. This can be used as a reference model for comparison with real-time monitored images to promptly detect abnormal changes in the pool environment, such as whether new objects have entered the pool or whether the pool facilities are damaged, helping to ensure the safe operation of the pool. By combining the motion information of dynamic objects with the environmental model, the future trajectories of dynamic objects such as swimmers can be predicted, allowing corresponding preparations to be made in advance. For example, lifeguards can use the prediction results to promptly detect swimmers who may be in danger and take rescue measures.

[0089] In the embodiment of the present application, dynamic objects in a swimming pool environment mainly refer to objects whose position, posture or state changes over time, such as swimmers, or floating objects in the water (if their motion state is unstable), working swimming pool cleaning equipment, etc.

[0090] This is the most prominent characteristic of dynamic objects. In a swimming pool, they move at varying speeds and directions, and their motion trajectories can be complex curves, influenced by factors such as the swimmer's movements and water currents. For example, a swimmer may adopt various strokes during swimming, such as breaststroke, freestyle, and backstroke. The appearance of their body in the image constantly changes, posing certain challenges to recognition and tracking.

[0091] The movement of dynamic objects can cause changes in the surrounding environment. For example, a swimmer's strokes generate currents, affecting the ripples on the water surface. This interactive information can also serve as important clues for identifying and analyzing dynamic objects. A deep learning-based dynamic SLAM algorithm is used to identify dynamic objects. This algorithm typically first extracts features from multi-source image data, learning the characteristic patterns of dynamic objects under different postures and lighting conditions. It then uses target detection and tracking techniques to determine the position and range of dynamic objects in the image in real time. When constructing the pool environment model, dynamic objects are removed from the environmental image data to prevent them from interfering with the modeling of the static environment. For example, the swimmer's body outline is identified and subtracted from the image, retaining information such as the static background and facilities of the pool for constructing the environment model. Furthermore, the motion information of dynamic objects is separately analyzed and recorded for subsequent operations such as trajectory prediction.

[0092] As an optional embodiment, in step S104, a deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, including:

[0093] The method obtains visual image data in different bands collected by a multispectral camera and a spatiotemporal feature map collected by an event camera activated by an event trigger threshold; performs spatiotemporal alignment on the visual image data in different bands and the spatiotemporal feature map; establishes a correspondence between the visual image data in different bands and the spatiotemporal feature map based on the color and spectral information of the multispectral image and the characteristic that the event camera is sensitive to dynamic changes through a feature matching algorithm; enhances the visual image data in different bands using a generative adversarial network (GAN) to strengthen the edge and texture details of each object in the visual image data in different bands; adopts a multi-head attention mechanism to globally model the visual image data in different bands, the spatiotemporal feature map, and the correspondence to obtain high-dimensional fusion features; The high-dimensional fusion features are used to capture the intrinsic connections of the dynamic objects in the spectral, temporal and spatial dimensions; the local areas and / or feature points in the high-dimensional fusion features are constructed as corresponding graph nodes, and the spatial structural relationship between the graph nodes and the motion pattern in the time series are learned through graph convolution operations, and the dynamic objects in the swimming pool environment are preliminarily labeled based on the motion pattern; the conditional random field algorithm is used to perform pixel-level segmentation on the visual image data in different bands based on the preliminary labeling information, and the boundary features of the dynamic objects are obtained through the spatial position relationship and feature similarity between each pixel; the event stream data of the event camera is combined to assist in the judgment of the boundary features; the boundary features obtained through the auxiliary judgment are used to delete the dynamic objects from the multi-source image data to obtain the environmental image data.

[0094] Specifically, the system first acquires visual image data in different wavelength bands from a multispectral camera. This data contains rich spectral information, which helps identify objects of different materials and characteristics. Simultaneously, an event camera, activated by an event-triggered threshold, collects spatiotemporal feature maps. The event camera is highly sensitive to dynamic changes and can capture rapidly occurring events in a scene.

[0095] Furthermore, the visual image data and spatiotemporal feature maps from different bands are temporally aligned to ensure temporal and spatial consistency between the two data types. Based on the color and spectral information of multispectral images and the event camera's sensitivity to dynamic changes, a feature matching algorithm is used to establish a correspondence between the visual image data from different bands and the spatiotemporal feature maps. For example, the corresponding position of an object in the multispectral image in the event camera's spatiotemporal feature map can be found, and vice versa.

[0096] Next, a generative adversarial network (GAN) was used to enhance the visual image data from different wavelengths. A GAN consists of a generator and a discriminator: the generator attempts to produce images similar to real images, while the discriminator distinguishes between real and generated images. Through this adversarial process, the edges and texture details of each object in the visual image data from different wavelengths are enhanced, making it easier for subsequent recognition tasks to detect the object's features.

[0097] Then, a multi-head attention mechanism is used to globally model the visual image data, spatiotemporal feature maps, and established correspondences in different bands. The multi-head attention mechanism can simultaneously focus on information at different locations and comprehensively process data from different modalities to obtain high-dimensional fusion features. High-dimensional fusion features are used to capture the intrinsic connections between dynamic objects in the spectral, spatiotemporal and spatial dimensions. For example, through this fusion, the spectral features of objects in multispectral images (such as reflectivity in specific bands) can be combined with the spatiotemporal features of the object's motion recorded by the event camera (such as direction and speed of motion) to form a more comprehensive and representative feature vector for better identification and understanding of dynamic objects.

[0098] Local regions and / or feature points from the high-dimensional fused features are constructed as corresponding graph nodes. Graph convolution operations are used to learn the spatial structural relationships between graph nodes and the motion patterns over time. Based on the learned motion patterns, dynamic objects in the swimming pool environment are preliminarily annotated. For example, based on the connectivity and motion trends between graph nodes, regions likely to be dynamic objects are identified, along with their approximate locations and postures.

[0099] The conditional random field (CRF) algorithm is used to perform pixel-level segmentation of visual image data in different bands based on preliminary annotation information. CRF takes into account the spatial position relationship and feature similarity between pixels and can accurately outline the boundaries of dynamic objects. The boundary features of dynamic objects are obtained by analyzing the spatial position relationship and feature similarity between each pixel. For example, the similarity of adjacent pixels in features such as color and texture, as well as their relative position relationship in space, are used to determine the boundaries of dynamic objects.

[0100] Combined with the event stream data from the event camera, boundary features are used to assist in the identification of dynamic objects. If a region generates a large number of event responses from the event camera and matches the segmented object region in the multispectral image, it is considered a dynamic object. Otherwise, the misidentified region is removed. Using the boundary features obtained through auxiliary identification, dynamic objects are removed from the multi-source image data, resulting in environmental image data containing only static environmental information. This completes the identification and processing of dynamic objects, laying the foundation for the subsequent construction of an accurate pool environment model.

[0101] Through the above steps, the deep learning-based dynamic SLAM algorithm can effectively identify and process dynamic objects in the swimming pool environment, providing reliable data support for generating accurate three-dimensional maps and subsequent path planning, trajectory prediction and other tasks.

[0102] Continuing with the above embodiment, in step S104, the swimming pool environment model is constructed by using the environment image data with the dynamic objects removed, including:

[0103] A sparse bundle adjustment (SBA) algorithm is used to construct a three-dimensional map of the swimming pool environment based on feature point information in the environmental image data to obtain the swimming pool environment model. Based on an incremental point cloud fusion algorithm, multi-source image data collected in real time is fused into the swimming pool environment model, and newly added feature point information is added to the three-dimensional map through feature matching and point cloud registration, thereby achieving real-time updating of the swimming pool environment model.

[0104] The Sparse Bundle Adjustment (SBA) algorithm is an optimization method used in multi-view geometry that aims to simultaneously optimize the camera pose and the 3D coordinates of scene points by minimizing the reprojection error. In swimming pool environment modeling, it works based on feature point information in the environment image data.

[0105] Specifically, feature points are first extracted from the environmental image data. These feature points can be corners, edges, or other points with unique characteristics. For example, the edges of swimming pool walls and the corners of underwater facilities can all serve as feature points. Common feature point extraction algorithms include SIFT (Scale-Invariant Feature Transform) and ORB (Accelerated Robust Features). Based on the correspondence between feature points in different images, methods such as triangulation are used to preliminarily estimate the camera pose (including rotation and translation). The camera pose describes the camera's position and orientation in space and is crucial for constructing a 3D map. The estimated camera pose and the 3D coordinates of the scene points are projected back onto the image plane to obtain reprojected points. The error between the reprojected points and the observed feature points is calculated, which is the reprojection error. A nonlinear optimization algorithm (such as the Levenberg-Marquardt algorithm) is used to minimize the reprojection error. During the optimization process, the camera pose and the 3D coordinates of the scene points are simultaneously adjusted to minimize the reprojection error. After multiple iterations of optimization, more accurate camera poses and 3D coordinates of the scene points are obtained, thus constructing a 3D map of the swimming pool environment, thus obtaining a swimming pool environment model.

[0106] The core principle of the incremental point cloud fusion algorithm is to gradually fuse multi-source image data collected in real time into an existing pool environment model. The model is updated by continuously adding new feature point information, allowing it to reflect real-time changes in the pool environment. Multi-source image data is collected in real time by various sensors deployed around the pool and in the water (such as multispectral cameras and event cameras). This data contains information about the pool environment at different times, potentially including newly appeared feature points or changes to existing feature points. Feature points in the real-time multi-source image data are matched with those in the existing pool environment model. This is achieved by calculating the similarity between the feature point descriptors (such as SIFT descriptors). Once matching feature point pairs are found, the correspondence between the newly collected data and the model is determined. Based on the feature matching results, a point cloud registration algorithm (such as the Iterative Closest Point (ICP) algorithm) is used to register the newly collected point cloud data with the existing pool environment model. The goal of registration is to accurately align the new point cloud data with the model, ensuring that its position and pose in 3D space are consistent with the model. After point cloud registration, the newly added feature points are added to the 3D map, updating the pool environment model. This allows the model to continuously incorporate new information over time, enabling real-time updates and maintaining an accurate description of the pool environment.

[0107] Through the above two steps, the SBA algorithm is first used to build an initial swimming pool environment model. Then, based on the incremental point cloud fusion algorithm, the real-time collected data is integrated into the model for real-time updating, thus obtaining an accurate and dynamic swimming pool environment model, which provides a reliable foundation for subsequent tasks such as path planning and trajectory prediction.

[0108] Step S105: Through the knowledge graph network, the activity trajectory model is deeply integrated with the swimming pool environment model, and the operation status of the swimming pool robot is predicted using a quantum machine learning algorithm to obtain a panoramic model including the predicted activity trajectory and predicted operation status of the swimming pool robot.

[0109] In the embodiment of the present application, the knowledge graph network is a graph-based data structure composed of nodes (entities) and edges (relations) for representing various knowledge and the associations between knowledge. In the present application, the knowledge graph network is constructed based on various entities in the swimming pool environment (such as the boundaries, facilities, swimmers, etc. of the swimming pool) and the relationships between them (such as positional relationships, motion relationships, etc.). It can integrate and associate information obtained from different data sources (such as multispectral cameras, event cameras, etc.) to form a structured knowledge network to better understand and analyze the swimming pool environment.

[0110] This fusion of information from the activity trajectory model and the pool environment model—for example, combining a swimmer's activity trajectory with the pool's static environment—enables the system to gain a global understanding of both the dynamic and static conditions within the pool. Relationships within the knowledge graph network enable reasoning and prediction. For example, based on the swimmer's current position and movement trends, as well as information about obstacles in the pool environment, possible collision risks can be inferred, providing a more comprehensive basis for the pool robot's path planning and operational status prediction.

[0111] Representing knowledge with an intuitive graph structure makes it easier for computers to understand and process it, and also helps humans manage and maintain swimming pool environment knowledge.

[0112] In the embodiments of this application, quantum machine learning is a method that applies the principles and techniques of quantum computing to the field of machine learning. It utilizes properties such as superposition and entanglement of quantum states to process and analyze data. Compared to traditional machine learning algorithms, it may have higher efficiency and performance when dealing with certain complex problems. For example, in this application, quantum bits may be used to represent the state information or environmental characteristics of a swimming pool robot. These quantum states are transformed and processed through quantum gate operations to predict the operating status of the swimming pool robot.

[0113] Pool environments contain a vast amount of multi-source data, including images and motion trajectories. Quantum machine learning algorithms can more efficiently process this complex data, extracting key features and thus more accurately predicting the operating status of pool robots. Leveraging the unique advantages of quantum computing, the prediction model is optimized, for example, more quickly finding the global optimal solution when searching for optimal parameters, improving the accuracy and reliability of predictions. This enables the model to better generalize across diverse pool environments and operational scenarios, adapting to a variety of complex situations and providing more stable and intelligent operational guidance for pool robots.

[0114] As an optional embodiment, in step S105, knowledge graph technology is used to construct a knowledge graph network containing semantic information of the swimming pool environment and robot motion rules; the robot's historical motion trajectory data and robot motion state information in the activity trajectory model, as well as the static map data and real-time detected obstacle information in the swimming pool environment model, are converted into node feature vectors in the knowledge graph network; the weights of the activity trajectory model and the swimming pool environment model are calculated using a graph attention mechanism (GAT) to obtain a fusion result between the activity trajectory model and the swimming pool environment model; wherein the attention weights of the swimming pool robot and the neighboring nodes in the fusion result are used to reflect the degree of influence of the neighboring nodes on the operating state of the swimming pool robot; the neighboring nodes include: surrounding obstacle nodes and swimming pool wall nodes; the fusion result is mapped to a quantum feature space, and the operating state of the swimming pool robot in the future is predicted by using quantum state superposition and entanglement characteristics, using a quantum support vector machine or a quantum neural network as a quantum machine learning model; the predicted result is integrated with the activity trajectory model and the swimming pool environment model to generate the panoramic model containing the predicted activity trajectory and predicted operating status, which is used to provide a decision basis for path planning and task scheduling of the swimming pool robot.

[0115] Knowledge graph technology is used to construct a network that encompasses semantic information about the pool environment and the robot's motion rules. Like an intelligent knowledge database, it organizes various aspects of the pool environment, such as its shape, size, and facility location, along with the rules the robot must follow within the pool, such as speed limits and steering angle restrictions, in a structured manner. This knowledge graph network clearly represents the relationships between different pieces of information, providing a foundation for subsequent analysis and processing.

[0116] Specifically, in step S105, the robot's historical motion trajectory data, motion state information (such as current speed, direction, etc.) in the activity trajectory model, as well as the static map data (such as the layout of the pool, fixed obstacle positions), and real-time detected obstacle information in the swimming pool environment model are all converted into node feature vectors in the knowledge graph network. This is equivalent to converting various types of information into digital vector forms that can be understood and processed by computers, so that the knowledge graph network can perform unified analysis and operations on this information. For example, the robot's historical motion trajectory can be converted into a vector representing features such as trajectory shape and speed changes through some feature extraction methods, while static map data can be converted into vectors representing the positions and attributes of different areas.

[0117] Next, the graph attention mechanism (GAT) is used to calculate the respective weights of the activity trajectory model and the swimming pool environment model. GAT can automatically learn the importance of each node relative to other nodes, that is, the attention weight. In this scenario, GAT can determine the importance of each part of the activity trajectory model and the swimming pool environment model for describing the operating status of the swimming pool robot, thereby obtaining the fusion result between the two. For example, for the swimming pool robot, the attention weights of the obstacle nodes and swimming pool wall nodes around it are higher because these neighboring nodes have a greater impact on its operating status. In this way, the information of the two models can be more accurately fused, highlighting the factors that have an important impact on the robot's operating status.

[0118] The fusion results are then mapped to the quantum feature space, which uses the characteristics of the quantum field to further process information. Quantum states have superposition and entanglement properties, which enable quantum systems to process multiple pieces of information simultaneously, and there are correlations between different quantum states. Quantum support vector machines or quantum neural networks are used as quantum machine learning models, and these characteristics are used to predict the operating state of the pool robot in the future. Quantum machine learning models can handle complex nonlinear problems more efficiently. For complex motion systems such as pool robots, they can more accurately predict their future operation, such as predicting their position and speed changes over the next period of time.

[0119] Finally, the prediction results are integrated with the activity trajectory model and the pool environment model to generate a comprehensive model that includes the predicted activity trajectories and predicted operational status. This comprehensive model integrates all relevant information, including the robot's past motion trajectory, current environmental information, and predictions of future operational status. This provides a comprehensive and accurate basis for the pool robot's path planning and task scheduling. For example, based on the comprehensive model, the robot can plan a path in advance to avoid obstacles or arrange the order and timing of cleaning tasks based on predicted operational status.

[0120] Step S106: Create and update in real time the digital twin model corresponding to the panoramic model to achieve real-time simulation and visualization of the operating status and environmental changes of the swimming pool robot, thereby facilitating maintenance and management decisions of the swimming pool robot.

[0121] In the embodiment of the present application, the digital twin model is a virtual model corresponding to a real physical system. It accurately simulates and maps the state, behavior and performance of the physical entity through data connection and real-time interaction, so as to better understand, predict and control the physical entity.

[0122] Specifically, the digital twin model leverages technologies such as the Internet of Things, big data, and artificial intelligence to transfer data from the physical world to a virtual model in real time, modeling and simulating physical entities through a data-driven approach. Simultaneously, the virtual model can optimize and control the physical entity based on simulation results, enabling two-way interaction between the physical and virtual worlds. Sensors and other devices collect various data from the physical entity, such as position, speed, and temperature, and transmit this data to the digital twin model via a network. Mathematical models and computer simulation techniques are used to simulate and predict the behavior and performance of the physical entity. Artificial intelligence and machine learning algorithms are used to analyze and process large amounts of data, enabling automated model optimization and decision support. The digital twin model's simulation results are presented intuitively, such as 3D graphics and animations, to facilitate user understanding and analysis. Based on information provided by the panoramic model, the digital twin model can simulate the pool robot's operating status in real time, including its position, speed, and posture, as well as changes in the surrounding environment, such as the location of obstacles and water flow. The simulation results are presented to the user in a visual format, allowing them to intuitively view the pool robot's operation through a graphical interface, allowing them to identify and make adjustments promptly. By analyzing the digital twin model, we can predict potential faults and problems with the pool robot, allowing for proactive maintenance and upkeep. Simulation results can also be used to optimize the robot's task scheduling and path planning, improving work efficiency.

[0123] In the embodiments of the present application, it is possible to realize an intelligent upgrade of the ranging and positioning process of the swimming pool robot from aspects such as environmental perception, data processing, model building to operation prediction and visual management. In terms of positioning accuracy, the fusion of terahertz sonar and multi-spectral visual data, combined with advanced algorithms, improves the accurate determination of the robot's position in complex environments. Secondly, in trajectory prediction, quantum machine learning and multi-model fusion technology significantly improve the accuracy and real-time performance of the prediction. In terms of environmental modeling capabilities, the three-dimensional map constructed by dynamic SLAM technology can reflect changes in the swimming pool environment in real time. Overall, the embodiments of the present application effectively improve the autonomous operation capability, environmental adaptability and intelligent management level of the swimming pool robot, and provide reliable and efficient technical support for tasks such as swimming pool cleaning and monitoring.

[0124] See also Figure 2 , Figure 2 The present invention provides a method and system for positioning and trajectory prediction of a swimming pool robot by integrating sonar and vision. The method and system for positioning and trajectory prediction of a swimming pool robot by integrating sonar and vision include the following modules:

[0125] The first acquisition module, a terahertz sonar device deployed in the robot, is used to collect terahertz sonar data for a target area to detect the distance information between the pool robot and surrounding objects in the target area; wherein the target area is the water environment in which the pool robot is located;

[0126] A first building module is used to build an activity trajectory model of the swimming pool robot based on the terahertz sonar data;

[0127] The second acquisition module, a visual sensor deployed around the pool and in the water environment, is used to acquire multi-source image data including underwater, surface, and pool environment changes in the target area using a multispectral and event camera collaborative acquisition strategy;

[0128] A second construction module is configured to employ a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, and to construct a swimming pool environment model by using the environmental image data without the dynamic objects; the dynamic objects include swimmers or moving objects in the environment;

[0129] A fusion module is used to deeply fuse the activity trajectory model with the swimming pool environment model through a knowledge graph network, and use a quantum machine learning algorithm to predict the operating status of the swimming pool robot to obtain a panoramic model including the predicted activity trajectory and predicted operating status of the swimming pool robot;

[0130] The display module is used to create and update the digital twin model corresponding to the panoramic model in real time, realize real-time simulation and visualization of the operating status and environmental changes of the swimming pool robot, and facilitate maintenance and management decisions of the swimming pool robot.

[0131] In some embodiments, the sonar-vision-integrated pool robot positioning and trajectory prediction method and system can be applied to a terminal device. It should be noted that for ease of description and brevity, the specific operating process of the sonar-vision-integrated pool robot positioning and trajectory prediction system described above can be referenced to the corresponding process in the aforementioned sonar-vision-integrated pool robot positioning and trajectory prediction method embodiment, and will not be repeated here.

[0132] See also Figure 3 , Figure 3 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present application. Figure 3As shown, terminal device 300 includes a processor 301 and a memory 302, which are connected via a bus 303, such as an I2C bus. Specifically, processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. Processor 301 can be a central processing unit, or it can be other general-purpose processors, digital signal processors, application-specific integrated circuits, field programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Specifically, memory 302 can be a Flash chip, a read-only memory disk, an optical disk, a USB flash drive, or a mobile hard disk, etc.

[0133] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the embodiment of the present application, and does not constitute a limitation on the terminal device to which the embodiment of the present application is applied. The specific server may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components. The processor is used to run the computer program stored in the memory, and implement any one of the methods for positioning and trajectory prediction of swimming pool robots that integrate sonar and vision provided by the embodiments of the present application when executing the computer program. It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the terminal device described above can refer to the aforementioned embodiment of the method for positioning and trajectory prediction of swimming pool robots that integrate sonar and vision, and will not be repeated here.

Claims

1. A swimming pool robot positioning and trajectory prediction method integrating sonar and vision, characterized in that: include: The terahertz sonar device deployed in the robot collects terahertz sonar data for the target area to detect the distance information between the pool robot and surrounding objects in the target area; wherein the target area is the water environment where the pool robot is located; constructing an activity trajectory model of the swimming pool robot based on the terahertz sonar data; By deploying visual sensors around the pool and in the water environment, and using a multispectral and event camera collaborative acquisition strategy, we can acquire multi-source image data covering underwater, surface, and pool environment changes in the target area. A deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, and a swimming pool environment model is constructed by removing the environmental image data of the dynamic objects; the dynamic objects include swimmers or moving objects in the environment; Through the knowledge graph network, the activity trajectory model is deeply integrated with the swimming pool environment model, and the operation status of the swimming pool robot is predicted using a quantum machine learning algorithm to obtain a panoramic model including the predicted activity trajectory and predicted operation status of the swimming pool robot; Create and update the digital twin model corresponding to the panoramic model in real time to achieve real-time simulation and visualization of the swimming pool robot's operating status and environmental changes, facilitating maintenance and management decisions for the swimming pool robot.

2. The method according to claim 1, characterized in that After collecting terahertz sonar data for the target area by the terahertz sonar device deployed in the robot, the method further includes: Obtaining swimming pool environmental data; the swimming pool environmental data includes at least: venue type, water turbidity, swimming pool lighting conditions that change over time, and weather information in the area where the swimming pool is located; Predicting dynamic changes in the swimming pool environment based on the swimming pool environment data, and adjusting sonar noise reduction parameters based on the predicted results; wherein the dynamic changes in the swimming pool environment include: trends in changes in water visibility, trends in changes in lighting conditions, and physical environmental disturbances caused by weather changes; The optimized and adjusted sonar noise reduction parameters are used to perform real-time noise reduction on the terahertz sonar data to improve the quality of the sonar data.

3. The method according to claim 2, characterized in that The predicting of the dynamic change of the swimming pool environment according to the swimming pool environment data, and adjusting the sonar noise reduction parameters based on the prediction results, includes: Input various environmental data such as venue type, water turbidity, swimming pool lighting conditions, and weather information into the adaptive deep learning model in the form of spatiotemporal graph nodes combined with sequence data; The dynamic capture module in the adaptive deep learning model captures the correlation between the various environmental data in the spatial dimension to obtain the spatiotemporal environmental correlation feature, and the spatiotemporal correlation module establishes the long-range dependency relationship between the various environmental data in the time dimension to obtain the long-range dependency feature; A multi-scale feature fusion module is used to perform cross-dimensional fusion of the high-frequency detail features and low-frequency structural features in the terahertz sonar data with the spatiotemporal environment correlation features and the long-range dependency features used to represent the dynamic changes of the swimming pool environment, so as to obtain an environmental fusion feature; The sonar noise reduction parameter adjustment strategy is optimized through a reinforcement learning module, using the improvement in the signal-to-noise ratio and positioning accuracy of the de-noised sonar data as reward functions. Combined with the sonar noise reduction parameter adjustment strategy obtained by reinforcement learning, the convolution kernel size and the hidden layer parameters of the recursive neural network used in the noise reduction process are dynamically optimized and adjusted. The optimized sonar noise reduction parameters are used to perform real-time noise reduction on the environmental fusion features to filter the noise data in the terahertz sonar data.

4. The method according to claim 1, wherein The step of constructing an activity trajectory model of the swimming pool robot based on the terahertz sonar data includes: Build a spatiotemporal Transformer model based on a hybrid architecture of the spatiotemporal graph convolutional network (STGCN) and the Transformer network; Obtain the inertial measurement module (IMU) built into the pool robot to obtain inertial measurement data; Using a spatiotemporal Transformer model, simulating the activity trajectory of the swimming pool robot based on the inertial measurement data to construct an initial activity trajectory model of the swimming pool robot; The distance relationship between the swimming pool robot and surrounding objects is simulated based on the terahertz sonar data to optimize the initial activity trajectory model, so that the initial activity trajectory model is adaptive to different swimming pool environments and robot motion modes, and the final output activity trajectory model is obtained.

5. The method according to claim 1, wherein The visual sensor further comprises: a multispectral camera for acquiring image information of different bands; The system uses visual sensors deployed around the pool and in the water environment, and adopts a multispectral and event camera collaborative acquisition strategy to obtain multi-source image data containing underwater, water surface, and pool environment changes in the target area, including: Light sensors and water flow sensors deployed in the swimming pool monitor environmental changes in real time to assess the complexity of the swimming pool environment; A dynamic sampling strategy is trained using a reinforcement learning algorithm to dynamically adjust the multispectral camera's band acquisition combination, the event camera's event trigger threshold, and the sensor collaboration mode based on the complexity of the swimming pool environment. When a sudden change in light or an intensified water disturbance is detected, the multispectral camera and the event camera are automatically triggered to enter a high-frame-rate collaborative acquisition mode. Through multispectral cameras, based on event trigger thresholds and sensor collaboration modes, visual image data in different bands is collected for the target area; Based on visual image data in different bands, complex environmental entities in the swimming pool environment are identified, and a dynamic filter group is used to perform image enhancement processing on the identified complex environmental entities to obtain optimized visual image data; wherein, the complex environmental entities include at least: transparent floating objects and underwater pipes.

6. The method according to claim 5, characterized in that The method of using a reinforcement learning algorithm to train a dynamic sampling strategy and dynamically adjusting the band acquisition combination of the multispectral camera, the event triggering threshold of the event camera, and the sensor collaboration mode according to the complexity of the swimming pool environment includes: The multispectral camera adopts an adaptive band switching method to obtain the spatial and spectral correlation between images of each band according to the real-time swimming pool light, water turbidity, and weather conditions in the complexity of the swimming pool environment. Based on the correlation, the camera automatically adjusts the acquisition band combination, event trigger threshold, and sensor collaboration mode to ensure that visual image data adapted to different swimming pool environments is collected under different working conditions, reducing the amount of redundant data collection and improving data validity.

7. The method according to claim 1, characterized in that The method adopts a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in a swimming pool environment in real time based on the multi-source image data, including: Obtain visual image data in different bands collected by a multispectral camera, and spatiotemporal feature maps collected by an event camera activated by an event trigger threshold; Performing spatiotemporal alignment on the visual image data in different bands and the spatiotemporal feature map; Based on the color and spectral information of the multispectral image and the characteristic of the event camera being sensitive to dynamic changes, a feature matching algorithm is used to establish a correspondence between the visual image data in different bands and the spatiotemporal feature map; Generative adversarial networks (GANs) are used to enhance visual image data in different bands to strengthen the edges and texture details of each object in the visual image data in different bands. A multi-head attention mechanism is used to globally model the visual image data in different bands, the spatiotemporal feature map, and the corresponding relationship to obtain high-dimensional fusion features; the high-dimensional fusion features are used to capture the intrinsic connections of the dynamic objects in the spectral, spatiotemporal and spatial dimensions; The local regions and / or feature points in the high-dimensional fusion features are constructed as corresponding graph nodes, and the spatial structural relationship between the graph nodes and the motion patterns in the time series are learned through graph convolution operations. The dynamic objects in the swimming pool environment are preliminarily annotated based on the motion patterns. A conditional random field algorithm is used to perform pixel-level segmentation on visual image data in different bands based on preliminary annotation information. The boundary features of the dynamic object are obtained by analyzing the spatial position relationship and feature similarity between each pixel. The event stream data from the event camera is combined to assist in the judgment of the boundary features. The boundary features determined through auxiliary judgment are used to delete the dynamic objects from the multi-source image data to obtain the environmental image data.

8. The method according to claim 7, characterized in that The method of constructing a swimming pool environment model by using the environment image data without the dynamic objects includes: Using a sparse bundle adjustment (SBA) algorithm, based on feature point information in the environmental image data, a three-dimensional map of the swimming pool environment is constructed to obtain the swimming pool environment model; Based on an incremental point cloud fusion algorithm, multi-source image data collected in real time is fused into the swimming pool environment model, so that newly added feature point information is added to the three-dimensional map through feature matching and point cloud registration, thereby achieving real-time updating of the swimming pool environment model.

9. The method according to claim 1, characterized in that The activity trajectory model is deeply integrated with the swimming pool environment model through the knowledge graph network, and the operation status of the swimming pool robot is predicted using a quantum machine learning algorithm to obtain a panoramic model containing the predicted activity trajectory and predicted operation status of the swimming pool robot, including: Using knowledge graph technology, we build a knowledge graph network that includes semantic information about the swimming pool environment and robot motion rules. Converting the robot's historical motion trajectory data and robot motion state information in the activity trajectory model, as well as the static map data and real-time detected obstacle information in the swimming pool environment model, into node feature vectors in the knowledge graph network; The weights of the activity trajectory model and the swimming pool environment model are calculated using the graph attention mechanism (GAT) to obtain a fusion result between the activity trajectory model and the swimming pool environment model. The attention weights of the swimming pool robot and the neighboring nodes in the fusion result are used to reflect the degree of influence of the neighboring nodes on the operating state of the swimming pool robot. The neighboring nodes include surrounding obstacle nodes and swimming pool wall nodes. The fusion results are mapped to the quantum feature space, and the quantum state superposition and entanglement characteristics are utilized. A quantum support vector machine or quantum neural network is used as a quantum machine learning model to predict the operating state of the swimming pool robot in the future. The prediction results are integrated with the activity trajectory model and the swimming pool environment model to generate the panoramic model including the predicted activity trajectory and the predicted operation status, which is used to provide a decision basis for the path planning and task scheduling of the swimming pool robot.

10. A swimming pool robot positioning and trajectory prediction system integrating sonar and vision, characterized in that: The system comprises: The first acquisition module, a terahertz sonar device deployed in the robot, is used to collect terahertz sonar data for a target area to detect the distance information between the pool robot and surrounding objects in the target area; wherein the target area is the water environment in which the pool robot is located; A first building module is used to build an activity trajectory model of the swimming pool robot based on the terahertz sonar data; The second acquisition module, a visual sensor deployed around the pool and in the water environment, is used to acquire multi-source image data including underwater, surface, and pool environment changes in the target area using a multispectral and event camera collaborative acquisition strategy; A second construction module is configured to employ a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the swimming pool environment in real time based on the multi-source image data, and to construct a swimming pool environment model by using the environmental image data without the dynamic objects; the dynamic objects include swimmers or moving objects in the environment; A fusion module is used to deeply fuse the activity trajectory model with the swimming pool environment model through a knowledge graph network, and use a quantum machine learning algorithm to predict the operating status of the swimming pool robot to obtain a panoramic model including the predicted activity trajectory and predicted operating status of the swimming pool robot; The display module is used to create and update the digital twin model corresponding to the panoramic model in real time, realize real-time simulation and visualization of the operating status and environmental changes of the swimming pool robot, and facilitate maintenance and management decisions of the swimming pool robot.

Citation Information

Patent Citations

  • Robot obstacle avoidance method based on deep learning panorama camera

    CN115946111A

  • Multi-sensor fusion underwater robot positioning method and system

    CN119178430A

  • Movement control method and device for swimming pool cleaning robot and swimming pool map construction method and device

    CN119937553A

  • Robot control method and apparatus

    WO2024221634A1

Cited By

  • Underwater acoustic-magnetic multi-modal fusion processing method

    CN120688019A

  • Method and system for automatically collecting fallen leaves and sediments in swimming pool based on visual identification

    CN120848560A

  • Method and system for automatic collection of fallen leaves and sediments in a swimming pool based on visual recognition

    CN120848560B

  • Swimming pool environment drowning risk identification management and control system based on VTN model

    CN120913158A

  • Underwater positioning system of underwater robot based on man-machine interaction

    CN120949243A